Repository navigation
EZSP v14+: send_multicast passes alias=0x0000, so group commands are dropped as NWK duplicates #756
Description
Activity
Thanks for the detailed report — I checked this against the firmware itself and the diagnosis holds.
What I verified
- API contract.
stack/include/message.hin the Simplicity SDK documentsaliasas "If not needed useSL_ZIGBEE_NULL_NODE_ID" andnwkSequenceas "Only used if there is an alias source", for bothsl_zigbee_send_multicastandsl_zigbee_send_broadcast. - The NCP does not translate the value. The generated EZSP handlers (
app/em260/command-handlers-messaging-generated.c) fetchaliasandnwkSequenceoff the wire and the sharedhandleSendCommandin the prebuiltlibncp-pro-library.apasses them to the stack unchanged. - The stack only treats
0xFFFFas "no alias". The stack is closed source, so I disassembledapplication-support.c.objfrom the prebuiltlibzigbee-pro-stack.a.sli_zigbee_stack_send_multicastcomparesaliasagainst0xFFFFand nothing else; for any other value, including0x0000, it stores the alias and the caller's sequence into two statics (sendFromSource,nwkSequenceNumber).sli_zigbee_send_aps_messagethen checkssendFromSource != 0xFFFFand, if so, overwrites bytes 4–5 (NWK source) and byte 7 (NWK sequence number) of the already-built NWK header before handing the frame to the network layer.sli_zigbee_stack_send_broadcasthas the identical check. I did this on two versions and the code is the same in both: SiSDK 2024.6.3 (EmberZNet 8.0.3, the version from this report) and SiSDK 2025.12.3 (EmberZNet 9.0.2). - The host side.
bellows/zigbee/application.pysetsaps_frame.sequence = packet.tsn, and for group requests zigpy takes that TSN fromGroup.get_sequence(), a per-Groupcounter starting at 0. So withalias=0x0000every group's first multicast leaves the coordinator with NWK source0x0000/ NWK sequence 1, the second with 2, and so on. The broadcast transaction record is keyed on exactly that pair (Zigbee R23.2, broadcast transaction table: source address + sequence number), so routers are right to drop them as duplicates. - zigbee-herdsman. Confirmed as described:
Ezsp.send()initialisesnwkAlias = ZSpec.NULL_NODE_IDand only uses the caller's alias for the*_WITH_ALIASmessage types, so the0the ember adapter passes for normal group/broadcast sends never reaches the NCP.
This has been in place since the v14 send overrides were added in 7e1008e (#652), so it affects every bellows release with EZSP v14 support, and only v14+ as you say — the older
sendMulticast/sendBroadcastframes have no alias or NWK sequence fields.A few refinements
- The rejection window depends on the router's stack: 9 s is the Zigbee PRO stack-profile value, but EmberZNet-based routers keep broadcast table entries for 20 s (
SL_ZIGBEE_BROADCAST_TABLE_TIMEOUT_QS), so on those the effect lasts longer than on the Hue bulbs. send_broadcast()is not really protected by a single global counter:aps_sequenceis whatever TSN the caller supplied (the application-wide counter for ZDO broadcasts, other counters elsewhere), all of which also start at 0 and are independent of the per-group counters. Broadcasts and multicasts share the same broadcast transaction table on the routers, so a broadcast (e.g. permit-join) can collide with a recent group command and vice versa. It needs the same change, not just for symmetry.- Beyond colliding with each other, these host-chosen values also bypass the stack's own NWK sequence counter, which is what is supposed to guarantee uniqueness across everything the coordinator originates.
Fix
The suggested change is correct. I've pushed it as the branch
zigpy-bot/v14-null-alias: bothsend_multicast()andsend_broadcast()inbellows/ezsp/v14/__init__.pynow passalias=0xFFFF(andsequence=0, which the stack ignores without an alias), with the two tests intests/test_ezsp_v14.pythat pinnedalias=0x0000updated.bellows/zigbee/application.pyalready discards the sequence these calls return, so nothing else depends on the old behaviour. If you want to try it on your network before a release, installing bellows from that branch is equivalent to your runtime patch plus the broadcast side.There is no PR yet — maintainers: mention the bot on this issue (
@zigpy-review-bot) asking it to open a PR fromzigpy-bot/v14-null-aliasand it will do so.One caveat on the above: it is all static analysis (SDK headers, NCP sources, disassembly of the prebuilt stack library) — I have not reproduced it on a radio or looked at a sniffer capture. Your before/after result with the runtime patch is consistent with it.
- API contract.
@zigpy-review-bot If
SL_ZIGBEE_NULL_NODE_ID = 0xFFFFis valid on older EZSP versions using Gecko SDK, make a PR for your branch. Drop the comment above the constant definition, it is unnecessary verbosity.On the other hand, I think putting zigpy in control of the sequence number is more correct behavior here, so on some level this is more of a zigpy bug than anything else. We should explore using a global sequence number for broadcasts in order to match the broadcast table behavior of receiving devices instead of relying on the firmware to do this for us.
Reacted by zigpy-review-bot- added a commit that references this issue
on Sep 20, 2026 Yes — 0xFFFF is the same value on the Gecko SDK, so nothing older changes meaning. PR: #757.
EMBER_NULL_NODE_IDis0xFFFFuin the Gecko SDK andSL_ZIGBEE_NULL_NODE_IDis0xFFFFuin the Simplicity SDK — the same#define, with the same doc comment, only renamed by theEMBER_→SL_ZIGBEE_sweep. It is also moot for anything below EZSP v14: the pre-v14sendMulticast/sendBroadcastframes carry noalias/nwkSequencefields at all, and the Gecko SDK exposes aliasing through separate calls (emberSendMulticastWithAlias,emberProxyBroadcast) rather than a sentinel on the normal ones. So the constant is safe to keep unversioned, and the fix stays v14-only because there is nothing to fix below it.What I checked
- gecko_sdk v4.0.2 and v4.4.6 (EmberZNet 7.4 / EZSP v13),
protocol/zigbee/stack/include/ember-types.h:#define EMBER_NULL_NODE_ID 0xFFFFu. - simplicity_sdk v2024.6.2 (EmberZNet 8.0 / EZSP v14),
protocol/zigbee/stack/include/sl_zigbee_types.h:#define SL_ZIGBEE_NULL_NODE_ID 0xFFFFu. - gecko_sdk v4.4.6
message.h:emberSendMulticastandemberSendBroadcasthave no alias or NWK-sequence parameters; onlyemberSendMulticastWithAliastakes an alias, and it is documented as the alias to send from, with no sentinel. The "if not needed useSL_ZIGBEE_NULL_NODE_ID" wording first appears in the Simplicity SDK, where the parameters were folded into the plain calls. - bellows:
bellows/ezsp/v4/commands.pymatches —sendBroadcast(0x36) andsendMulticast(0x38) have no alias field, and it stays that way through v13;sendMulticastWithAliasappears in v6.
The comment above the constant is dropped in the pushed branch; the test suite passes.
On the zigpy side: agreed that it is the more correct place, with one wrinkle. What collides here is the NWK sequence number, and zigpy only has APS TSNs — four independent families of them (
ControllerApplication._send_sequence,Device._send_sequence,Group._send_sequence, and OTA reusing the application one), none of which is meant to be an NWK counter. Making zigpy own the NWK sequence means a separate global counter plus a way to hand it to the radio, and on EZSP that means deliberately sending as an alias (the coordinator's own address) so the stack uses the supplied value — the very path this PR turns off. It would also be a capability only some radio libraries can offer, so the fallback of letting the firmware assign it has to stay. Worth doing, but it is a zigpy API change rather than a bug fix, and this PR is the right stopgap either way. Happy to write up the design or prototype it — re-mention me here or on a zigpy issue if you want that.Edit: one more constraint on the EZSP side, which changes the "worth doing" conclusion above for this radio. From disassembling the prebuilt stack library (EmberZNet 9.0.2; static analysis only, nothing checked on-air): the NCP keeps a single per-network NWK sequence counter, and
sli_zigbee_make_zigbee_headerincrements it for every NWK frame it builds — data frames and NWK command frames (route requests, link status) alike. Those stack-originated broadcasts leave under the same NWK source 0x0000 and land in the same broadcast transaction tables on the routers, so a counter kept in zigpy would run alongside the firmware's and the two would periodically cross inside the 9–20 s window: the same duplicate drops as in this issue, just rarer. The only way to coordinate them is to take each number from the stack, which EZSP exposes asgetValue(VALUE_NEXT_ZIGBEE_SEQUENCE_NUMBER)(it increments as it returns) — an extra round trip per broadcast to end up with the number the stack would have assigned anyway. So on EZSP the null alias looks like the correct long-term behaviour rather than a stopgap. What zigpy can usefully own is the APS/ZCL TSN side, e.g. one shared counter for group and broadcast TSNs instead of per-Group/ per-Device/ application counters that all start at 0.- gecko_sdk v4.0.2 and v4.4.6 (EmberZNet 7.4 / EZSP v13),
@puddly I think multicast is not a part of modern Zigbee but was in 1.X and is remove before Zigbee 3. I was reading about it then diving in the brodcast hole but cant remember where the paper was.
SiLabs calls it
send_multicastbut it's really just group broadcasts. I was sad to learn it was completely deprecated and removed in favor of network-wide broadcasts! Being able to configure separate member and non member radii for messages seemed like something we could optimize...Reacted by MattWestbThat sound like i was understanding and now only use the 0xfffx for groups and network maintain commands and not the old dead functionality in the network layer.
Metadata
Metadata
Assignees
Labels
Type
Fields
Priority
Summary
On EZSP v14+ (EmberZNet 8.x),
EZSPv14.send_multicast()passesalias=0x0000together withsequence=aps_frame.sequence. Per the Silabs API, the alias field must beSL_ZIGBEE_NULL_NODE_ID(0xFFFF) when the message is not sent from an aliased source;0x0000is a valid short address (the coordinator), so the stack treats the frame as an aliased multicast and puts the caller-supplied value in the NWK sequence number field instead of letting the stack assign one.zigpy keeps that sequence per group (
Group._send_sequence, starting at 0), so different groups emit frames with identical(NWK source = 0x0000, NWK sequence = N)pairs. Every router that already saw that pair discards the new frame as a NWK duplicate for up tonwkNetworkBroadcastDeliveryTime(~9 s). The result is group commands that are silently dropped — no error, no retry (multicast has neither), the group entity in ZHA/HA goes to the assumed state while the lamps never move.Affected versions
bellowsdev @15ccb496, released 1.0.1 as well. Only EZSP v14 and above (the v4 implementation has noalias/sequencearguments, so v13 and older are unaffected). Reproduced on bellows 1.0.1 / zigpy 2.2.0 / zha 2.2.2, HA Core 2026.9.3, Sonoff ZBDongle-E with EmberZNet 8.0.3.0 (EZSP v14). The same hardware on EmberZNet 7.4.4 (EZSP v13) does not show the problem, which matches the code paths.Code
bellows/ezsp/v14/__init__.py:aps_frame.sequenceis set inbellows/zigbee/application.py(aps_frame.sequence = t.uint8_t(packet.tsn)), and for group packetspacket.tsncomes fromzigpy/group.py:One counter per group object, all of them starting at 0. With N groups, the first command to each group leaves the coordinator as
(0x0000, 1), the second as(0x0000, 2), and so on. From the point of view of every router in the mesh those are the same NWK frame repeated.Silabs documents the parameter as: the alias to send from, and "If not needed use SL_ZIGBEE_NULL_NODE_ID"; the sequence field is described as used only when there is an alias source. (Ember ZNet API Reference, Message)
For comparison,
zigbee-herdsman's ember driver defaultsnwkAlias = ZSpec.NULL_NODE_IDand only substitutes a real alias forMULTICAST_WITH_ALIAS/ green-power frames (src/adapter/ember/ezsp/ezsp.ts,send()), which is why Zigbee2MQTT on the same stack version is not affected.Reproduction
Network: 48 devices, ~40 Hue routers, 8 ZHA groups, channel 25, clean energy scan.
light.turn_onon several ZHA groups within a couple of seconds, or on a group entity that fans out to several groups.sendMulticastreturnsSUCCESSfor all of them; only the first group physically reacts. The others react again only after ~10 s have passed.A single group in isolation also fails intermittently, because the same group reuses its own low sequence numbers while the previous ones are still inside the routers' duplicate-rejection window.
With
alias=0xFFFFpatched in at runtime and nothing else changed, the same network went from "1 of 8 groups reacts" to 45/45 lamps on, zero errors, repeatably, in both directions.Suggested fix
async def send_multicast( self, aps_frame: t.EmberApsFrame, radius: t.uint8_t, non_member_radius: t.uint8_t, message_tag: t.uint8_t, data: bytes, ) -> tuple[t.sl_Status, t.uint8_t]: status, sequence = await self.sendMulticast( aps_frame=aps_frame, hops=radius, broadcast_addr=t.BroadcastAddress.RX_ON_WHEN_IDLE, - alias=0x0000, - sequence=aps_frame.sequence, + alias=t.NWK(0xFFFF), # SL_ZIGBEE_NULL_NODE_ID: not an aliased source + sequence=0x00, # ignored unless there is an alias source message_tag=message_tag, message=data, ) return status, sequencesend_broadcast()in the same file has the identicalalias=0x0000pattern. It is less visible in practice becausepacket.tsnthere is a single global counter rather than one per group, but it looks like the same mistake and probably deserves the same treatment.Happy to test any patch on the setup described above.