Skip to content

EZSP v14+: send_multicast passes alias=0x0000, so group commands are dropped as NWK duplicates #756

Description

@gabrielcmlopes

Summary

On EZSP v14+ (EmberZNet 8.x), EZSPv14.send_multicast() passes alias=0x0000 together with sequence=aps_frame.sequence. Per the Silabs API, the alias field must be SL_ZIGBEE_NULL_NODE_ID (0xFFFF) when the message is not sent from an aliased source; 0x0000 is a valid short address (the coordinator), so the stack treats the frame as an aliased multicast and puts the caller-supplied value in the NWK sequence number field instead of letting the stack assign one.

zigpy keeps that sequence per group (Group._send_sequence, starting at 0), so different groups emit frames with identical (NWK source = 0x0000, NWK sequence = N) pairs. Every router that already saw that pair discards the new frame as a NWK duplicate for up to nwkNetworkBroadcastDeliveryTime (~9 s). The result is group commands that are silently dropped — no error, no retry (multicast has neither), the group entity in ZHA/HA goes to the assumed state while the lamps never move.

Affected versions

bellows dev @ 15ccb496, released 1.0.1 as well. Only EZSP v14 and above (the v4 implementation has no alias/sequence arguments, so v13 and older are unaffected). Reproduced on bellows 1.0.1 / zigpy 2.2.0 / zha 2.2.2, HA Core 2026.9.3, Sonoff ZBDongle-E with EmberZNet 8.0.3.0 (EZSP v14). The same hardware on EmberZNet 7.4.4 (EZSP v13) does not show the problem, which matches the code paths.

Code

bellows/ezsp/v14/__init__.py:

async def send_multicast(self, aps_frame, radius, non_member_radius, message_tag, data):
    status, sequence = await self.sendMulticast(
        aps_frame=aps_frame,
        hops=radius,
        broadcast_addr=t.BroadcastAddress.RX_ON_WHEN_IDLE,
        alias=0x0000,                  # <-- should be NULL_NODE_ID (0xFFFF)
        sequence=aps_frame.sequence,   # <-- only meaningful for an aliased source
        message_tag=message_tag,
        message=data,
    )
    return status, sequence

aps_frame.sequence is set in bellows/zigbee/application.py (aps_frame.sequence = t.uint8_t(packet.tsn)), and for group packets packet.tsn comes from zigpy/group.py:

self._send_sequence = 0
...
def get_sequence(self) -> int:
    self._send_sequence = (self._send_sequence + 1) % 256
    return self._send_sequence

One counter per group object, all of them starting at 0. With N groups, the first command to each group leaves the coordinator as (0x0000, 1), the second as (0x0000, 2), and so on. From the point of view of every router in the mesh those are the same NWK frame repeated.

Silabs documents the parameter as: the alias to send from, and "If not needed use SL_ZIGBEE_NULL_NODE_ID"; the sequence field is described as used only when there is an alias source. (Ember ZNet API Reference, Message)

For comparison, zigbee-herdsman's ember driver defaults nwkAlias = ZSpec.NULL_NODE_ID and only substitutes a real alias for MULTICAST_WITH_ALIAS / green-power frames (src/adapter/ember/ezsp/ezsp.ts, send()), which is why Zigbee2MQTT on the same stack version is not affected.

Reproduction

Network: 48 devices, ~40 Hue routers, 8 ZHA groups, channel 25, clean energy scan.

  1. Turn all lights off, so the multicast actually has an observable effect.
  2. Call light.turn_on on several ZHA groups within a couple of seconds, or on a group entity that fans out to several groups.
  3. sendMulticast returns SUCCESS for all of them; only the first group physically reacts. The others react again only after ~10 s have passed.

A single group in isolation also fails intermittently, because the same group reuses its own low sequence numbers while the previous ones are still inside the routers' duplicate-rejection window.

With alias=0xFFFF patched in at runtime and nothing else changed, the same network went from "1 of 8 groups reacts" to 45/45 lamps on, zero errors, repeatably, in both directions.

Suggested fix

     async def send_multicast(
         self,
         aps_frame: t.EmberApsFrame,
         radius: t.uint8_t,
         non_member_radius: t.uint8_t,
         message_tag: t.uint8_t,
         data: bytes,
     ) -> tuple[t.sl_Status, t.uint8_t]:
         status, sequence = await self.sendMulticast(
             aps_frame=aps_frame,
             hops=radius,
             broadcast_addr=t.BroadcastAddress.RX_ON_WHEN_IDLE,
-            alias=0x0000,
-            sequence=aps_frame.sequence,
+            alias=t.NWK(0xFFFF),  # SL_ZIGBEE_NULL_NODE_ID: not an aliased source
+            sequence=0x00,        # ignored unless there is an alias source
             message_tag=message_tag,
             message=data,
         )
 
         return status, sequence

send_broadcast() in the same file has the identical alias=0x0000 pattern. It is less visible in practice because packet.tsn there is a single global counter rather than one per group, but it looks like the same mistake and probably deserves the same treatment.

Happy to test any patch on the setup described above.

Activity

  1. zigpy-review-bot commented on Sep 20, 2026

    @zigpy-review-bot
    Collaborator

    Thanks for the detailed report — I checked this against the firmware itself and the diagnosis holds.

    What I verified

    • API contract. stack/include/message.h in the Simplicity SDK documents alias as "If not needed use SL_ZIGBEE_NULL_NODE_ID" and nwkSequence as "Only used if there is an alias source", for both sl_zigbee_send_multicast and sl_zigbee_send_broadcast.
    • The NCP does not translate the value. The generated EZSP handlers (app/em260/command-handlers-messaging-generated.c) fetch alias and nwkSequence off the wire and the shared handleSendCommand in the prebuilt libncp-pro-library.a passes them to the stack unchanged.
    • The stack only treats 0xFFFF as "no alias". The stack is closed source, so I disassembled application-support.c.obj from the prebuilt libzigbee-pro-stack.a. sli_zigbee_stack_send_multicast compares alias against 0xFFFF and nothing else; for any other value, including 0x0000, it stores the alias and the caller's sequence into two statics (sendFromSource, nwkSequenceNumber). sli_zigbee_send_aps_message then checks sendFromSource != 0xFFFF and, if so, overwrites bytes 4–5 (NWK source) and byte 7 (NWK sequence number) of the already-built NWK header before handing the frame to the network layer. sli_zigbee_stack_send_broadcast has the identical check. I did this on two versions and the code is the same in both: SiSDK 2024.6.3 (EmberZNet 8.0.3, the version from this report) and SiSDK 2025.12.3 (EmberZNet 9.0.2).
    • The host side. bellows/zigbee/application.py sets aps_frame.sequence = packet.tsn, and for group requests zigpy takes that TSN from Group.get_sequence(), a per-Group counter starting at 0. So with alias=0x0000 every group's first multicast leaves the coordinator with NWK source 0x0000 / NWK sequence 1, the second with 2, and so on. The broadcast transaction record is keyed on exactly that pair (Zigbee R23.2, broadcast transaction table: source address + sequence number), so routers are right to drop them as duplicates.
    • zigbee-herdsman. Confirmed as described: Ezsp.send() initialises nwkAlias = ZSpec.NULL_NODE_ID and only uses the caller's alias for the *_WITH_ALIAS message types, so the 0 the ember adapter passes for normal group/broadcast sends never reaches the NCP.

    This has been in place since the v14 send overrides were added in 7e1008e (#652), so it affects every bellows release with EZSP v14 support, and only v14+ as you say — the older sendMulticast/sendBroadcast frames have no alias or NWK sequence fields.

    A few refinements

    • The rejection window depends on the router's stack: 9 s is the Zigbee PRO stack-profile value, but EmberZNet-based routers keep broadcast table entries for 20 s (SL_ZIGBEE_BROADCAST_TABLE_TIMEOUT_QS), so on those the effect lasts longer than on the Hue bulbs.
    • send_broadcast() is not really protected by a single global counter: aps_sequence is whatever TSN the caller supplied (the application-wide counter for ZDO broadcasts, other counters elsewhere), all of which also start at 0 and are independent of the per-group counters. Broadcasts and multicasts share the same broadcast transaction table on the routers, so a broadcast (e.g. permit-join) can collide with a recent group command and vice versa. It needs the same change, not just for symmetry.
    • Beyond colliding with each other, these host-chosen values also bypass the stack's own NWK sequence counter, which is what is supposed to guarantee uniqueness across everything the coordinator originates.

    Fix

    The suggested change is correct. I've pushed it as the branch zigpy-bot/v14-null-alias: both send_multicast() and send_broadcast() in bellows/ezsp/v14/__init__.py now pass alias=0xFFFF (and sequence=0, which the stack ignores without an alias), with the two tests in tests/test_ezsp_v14.py that pinned alias=0x0000 updated. bellows/zigbee/application.py already discards the sequence these calls return, so nothing else depends on the old behaviour. If you want to try it on your network before a release, installing bellows from that branch is equivalent to your runtime patch plus the broadcast side.

    There is no PR yet — maintainers: mention the bot on this issue (@zigpy-review-bot) asking it to open a PR from zigpy-bot/v14-null-alias and it will do so.

    One caveat on the above: it is all static analysis (SDK headers, NCP sources, disassembly of the prebuilt stack library) — I have not reproduced it on a radio or looked at a sniffer capture. Your before/after result with the runtime patch is consistent with it.

  2. puddly commented on Sep 20, 2026

    @puddly
    Contributor

    @zigpy-review-bot If SL_ZIGBEE_NULL_NODE_ID = 0xFFFF is valid on older EZSP versions using Gecko SDK, make a PR for your branch. Drop the comment above the constant definition, it is unnecessary verbosity.

    On the other hand, I think putting zigpy in control of the sequence number is more correct behavior here, so on some level this is more of a zigpy bug than anything else. We should explore using a global sequence number for broadcasts in order to match the broadcast table behavior of receiving devices instead of relying on the firmware to do this for us.

  3. zigpy-review-bot commented on Sep 20, 2026

    @zigpy-review-bot
    Collaborator

    Yes — 0xFFFF is the same value on the Gecko SDK, so nothing older changes meaning. PR: #757.

    EMBER_NULL_NODE_ID is 0xFFFFu in the Gecko SDK and SL_ZIGBEE_NULL_NODE_ID is 0xFFFFu in the Simplicity SDK — the same #define, with the same doc comment, only renamed by the EMBER_ → SL_ZIGBEE_ sweep. It is also moot for anything below EZSP v14: the pre-v14 sendMulticast/sendBroadcast frames carry no alias/nwkSequence fields at all, and the Gecko SDK exposes aliasing through separate calls (emberSendMulticastWithAlias, emberProxyBroadcast) rather than a sentinel on the normal ones. So the constant is safe to keep unversioned, and the fix stays v14-only because there is nothing to fix below it.

    What I checked
    • gecko_sdk v4.0.2 and v4.4.6 (EmberZNet 7.4 / EZSP v13), protocol/zigbee/stack/include/ember-types.h: #define EMBER_NULL_NODE_ID 0xFFFFu.
    • simplicity_sdk v2024.6.2 (EmberZNet 8.0 / EZSP v14), protocol/zigbee/stack/include/sl_zigbee_types.h: #define SL_ZIGBEE_NULL_NODE_ID 0xFFFFu.
    • gecko_sdk v4.4.6 message.h: emberSendMulticast and emberSendBroadcast have no alias or NWK-sequence parameters; only emberSendMulticastWithAlias takes an alias, and it is documented as the alias to send from, with no sentinel. The "if not needed use SL_ZIGBEE_NULL_NODE_ID" wording first appears in the Simplicity SDK, where the parameters were folded into the plain calls.
    • bellows: bellows/ezsp/v4/commands.py matches — sendBroadcast (0x36) and sendMulticast (0x38) have no alias field, and it stays that way through v13; sendMulticastWithAlias appears in v6.

    The comment above the constant is dropped in the pushed branch; the test suite passes.

    On the zigpy side: agreed that it is the more correct place, with one wrinkle. What collides here is the NWK sequence number, and zigpy only has APS TSNs — four independent families of them (ControllerApplication._send_sequence, Device._send_sequence, Group._send_sequence, and OTA reusing the application one), none of which is meant to be an NWK counter. Making zigpy own the NWK sequence means a separate global counter plus a way to hand it to the radio, and on EZSP that means deliberately sending as an alias (the coordinator's own address) so the stack uses the supplied value — the very path this PR turns off. It would also be a capability only some radio libraries can offer, so the fallback of letting the firmware assign it has to stay. Worth doing, but it is a zigpy API change rather than a bug fix, and this PR is the right stopgap either way. Happy to write up the design or prototype it — re-mention me here or on a zigpy issue if you want that.

    Edit: one more constraint on the EZSP side, which changes the "worth doing" conclusion above for this radio. From disassembling the prebuilt stack library (EmberZNet 9.0.2; static analysis only, nothing checked on-air): the NCP keeps a single per-network NWK sequence counter, and sli_zigbee_make_zigbee_header increments it for every NWK frame it builds — data frames and NWK command frames (route requests, link status) alike. Those stack-originated broadcasts leave under the same NWK source 0x0000 and land in the same broadcast transaction tables on the routers, so a counter kept in zigpy would run alongside the firmware's and the two would periodically cross inside the 9–20 s window: the same duplicate drops as in this issue, just rarer. The only way to coordinate them is to take each number from the stack, which EZSP exposes as getValue(VALUE_NEXT_ZIGBEE_SEQUENCE_NUMBER) (it increments as it returns) — an extra round trip per broadcast to end up with the number the stack would have assigned anyway. So on EZSP the null alias looks like the correct long-term behaviour rather than a stopgap. What zigpy can usefully own is the APS/ZCL TSN side, e.g. one shared counter for group and broadcast TSNs instead of per-Group / per-Device / application counters that all start at 0.

  4. MattWestb commented on Sep 20, 2026

    @MattWestb
    Contributor

    @puddly I think multicast is not a part of modern Zigbee but was in 1.X and is remove before Zigbee 3. I was reading about it then diving in the brodcast hole but cant remember where the paper was.

  5. puddly commented on Sep 20, 2026

    @puddly
    Contributor

    SiLabs calls it send_multicast but it's really just group broadcasts. I was sad to learn it was completely deprecated and removed in favor of network-wide broadcasts! Being able to configure separate member and non member radii for messages seemed like something we could optimize...

  6. MattWestb commented on Sep 20, 2026

    @MattWestb
    Contributor

    That sound like i was understanding and now only use the 0xfffx for groups and network maintain commands and not the old dead functionality in the network layer.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Fields

    Priority

    None yet

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions