Skip to content

mDNS rediscovery fails for previously commissioned devices #113

Description

@qwandor

With the issues in the other bugs fixed I am able to successfully commission a number of Matter devices and read attributes from them. However, when I restart my service and attempt to connect to the devices I previously commissioned again, it almost always fails with an error like "operational error: connect to node 2 failed: driver error: device discovery failed: operational node F52AC107C954E38E-0000000000000002 not found via mDNS". (It does occasionally work, but it seems to work less than 5% of the time.)

If I open avahi-discover I can see all these addresses advertised under _matter._tcp just fine, so it seems to be something to do with how the matter-controller crate does mDNS discovery in this case. As another data point, the same devices also reconnect fine if I use the matc crate instead (though that crate has other limitations, so I'd prefer to use matter-controller).

The same issue happens for all the devices I've tried, and it only seems to happen on re-connection; the initial commissioning works fine most of the time. I'm running on Ubuntu 24.04.4, in case that makes a difference.

Activity

  1. hemanpa commented on Aug 18, 2026

    @hemanpa
    Contributor

    Thanks — this one was worth chasing, and it turned up a real bug in our mDNS adapter. I want to be straight with you though: I could not reproduce your failure here, so I can't yet claim this is your fault. What I can say is that what I found is consistent with everything you describe.

    The bug

    mdns-sd keys its listeners by service type — service_queriers: HashMap<ty_domain, Sender> — so there is exactly one listener per type. Our adapter handed out a QueryHandle per query() call and stored a Receiver per handle, as though they were independent browses. They aren't:

    • a second browse("_matter._tcp.local.") replaces the first sender, orphaning the first Receiver — it never receives again;
    • stop_browse then removes the shared querier and wipes that type's cache.

    So any second operational resolve on the same Discovery could silently kill the controller's own browse. Afterwards the actor still held a Some(handle) and therefore never re-browsed, and every later resolve expired with not found via mDNS — permanently, for the life of the process, while avahi-browse kept showing the records. That is a very close match to "almost always fails after a restart, occasionally works, avahi sees everything fine".

    Fixed in matter-controller 0.7.1 (with matter-transport 0.3.2, matter-commissioning 0.5.2): one browse per service type, reference-counted and fanned out to every handle, with stop_browse only on the last release. Records already surfaced are replayed to a handle that attaches later — without that, a late handle would see nothing, because mdns-sd only emits ServiceResolved for changed records.

    Validated on a live LAN with ~14 operational records: two concurrent handles both receive the full set, releasing one leaves the other's browse alive, and a real device reconnects across fresh processes.

    If it still happens, this release will tell us why

    The honest problem with your original report is that our discovery path had no logging at all, and four places where a record could be dropped silently — so the failure was undiagnosable from the outside. That's fixed too.

    If you can reproduce on 0.7.1, please run with:

    RUST_LOG=matter_transport=debug,matter_controller=debug
    

    and paste the output around a failed connect. You'll get a line per browse, per handle attach, and per record surfaced, plus an explicit line for each dropped record with the reason — including the record's ty_domain when it isn't recognised. That last one matters: Matter advertises operational records under _I<compressed-fabric>._sub._matter._tcp subtypes, and if mdns-sd ever surfaces one under the subtype rather than the base type, we'd drop it silently today. I couldn't trigger that here, but the log will say so immediately if it's what you're hitting.

    The error message itself is also more useful now — it reports what discovery actually saw:

    … not found via mDNS (saw 3 operational mDNS record(s), none matching:
      f52ac107c954e38e-0000000000000003, …)
    

    versus (saw 0 operational mDNS records …). That alone distinguishes "nothing reached us" from "records reached us but yours wasn't among them", which is the first fork in diagnosing this.

    Three things that would help if it persists: whether avahi-daemon is running alongside your service, whether the affected devices are IPv4, IPv6 or dual-stack, and whether the same binary fails against all your devices or only some.

  2. qwandor commented on Aug 18, 2026

    @qwandor
    ContributorAuthor

    Thanks for the quick response! Unfortunately the issue still seems to happen with the new version.

    The errors are now:

    operational error: connect to node 2 failed: driver error: device discovery failed: operational node F52AC107C954E38E-0000000000000002 not found via mDNS (saw 6 operational mDNS record(s), none matching: 2c755bdfdbfdf8a2-0000000000840c83, 2c755bdfdbfdf8a2-000000006ec96bf5, 2c755bdfdbfdf8a2-00000000dfcb37cc, 2c755bdfdbfdf8a2-00000000f141dde6, 2c755bdfdbfdf8a2-00000000f25ef2e4, … 1 more)
    operational error: connect to node 3 failed: driver error: device discovery failed: operational node F52AC107C954E38E-0000000000000003 not found via mDNS (saw 6 operational mDNS record(s), none matching: 2c755bdfdbfdf8a2-0000000000840c83, 2c755bdfdbfdf8a2-000000006ec96bf5, 2c755bdfdbfdf8a2-00000000dfcb37cc, 2c755bdfdbfdf8a2-00000000f141dde6, 2c755bdfdbfdf8a2-00000000f25ef2e4, … 1 more)
    operational error: connect to node 4 failed: driver error: device discovery failed: operational node F52AC107C954E38E-0000000000000004 not found via mDNS (saw 6 operational mDNS record(s), none matching: 2c755bdfdbfdf8a2-0000000000840c83, 2c755bdfdbfdf8a2-000000006ec96bf5, 2c755bdfdbfdf8a2-00000000dfcb37cc, 2c755bdfdbfdf8a2-00000000f141dde6, 2c755bdfdbfdf8a2-00000000f25ef2e4, … 1 more)
    

    All three devices I tested with connect over Thread, which I assume means they are IPv6 only.

    I've attached the debug logs: log.txt

  3. qwandor commented on Aug 18, 2026

    @qwandor
    ContributorAuthor

    For additional context, I tried running the query example from the mdns-sd crate, and that seems to see the nodes right away:

    $ cargo run --example query _matter._tcp
        Finished `dev` profile [unoptimized + debuginfo] target(s) in 0.02s
         Running `target/debug/examples/query _matter._tcp`
    At 322.513µs: SearchStarted("_matter._tcp.local. on 2 interfaces [wlp170s0 (3), lo (1)]")
    At 46.647774ms: SearchStarted("_matter._tcp.local. on 2 interfaces [wlp170s0 (3), lo (1)]")
    At 317.055798ms: ServiceFound("_matter._tcp.local.", "2C755BDFDBFDF8A2-00000000AAB08AA7._matter._tcp.local.")
    At 317.094102ms: ServiceFound("_matter._tcp.local.", "2C755BDFDBFDF8A2-00000000056BB72A._matter._tcp.local.")
    At 317.105053ms: ServiceFound("_matter._tcp.local.", "2C755BDFDBFDF8A2-000000000BD7BD08._matter._tcp.local.")
    At 317.111374ms: ServiceFound("_matter._tcp.local.", "2C755BDFDBFDF8A2-0000000077A6EFB7._matter._tcp.local.")
    At 317.115939ms: ServiceFound("_matter._tcp.local.", "2C755BDFDBFDF8A2-000000009B2EEF43._matter._tcp.local.")
    At 317.123104ms: ServiceFound("_matter._tcp.local.", "2C755BDFDBFDF8A2-000000004D852B21._matter._tcp.local.")
    At 317.129643ms: ServiceFound("_matter._tcp.local.", "2C755BDFDBFDF8A2-00000000642B4839._matter._tcp.local.")
    At 317.135777ms: ServiceFound("_matter._tcp.local.", "2C755BDFDBFDF8A2-FFFF000000000000._matter._tcp.local.")
    At 317.142114ms: ServiceFound("_matter._tcp.local.", "2C755BDFDBFDF8A2-0000000061817431._matter._tcp.local.")
    At 317.148434ms: ServiceFound("_matter._tcp.local.", "2C755BDFDBFDF8A2-000000004192B362._matter._tcp.local.")
    At 317.154801ms: ServiceFound("_matter._tcp.local.", "2C755BDFDBFDF8A2-00000000F8CD79FE._matter._tcp.local.")
    At 317.161051ms: ServiceFound("_matter._tcp.local.", "9B5791331E047398-000000000000012D._matter._tcp.local.")
    At 317.167035ms: ServiceFound("_matter._tcp.local.", "9B5791331E047398-000000000000012E._matter._tcp.local.")
    At 317.174405ms: ServiceFound("_matter._tcp.local.", "F52AC107C954E38E-0000000000000002._matter._tcp.local.")
    At 317.181164ms: ServiceFound("_matter._tcp.local.", "F52AC107C954E38E-0000000000000003._matter._tcp.local.")
    At 317.187235ms: ServiceFound("_matter._tcp.local.", "F52AC107C954E38E-0000000000000004._matter._tcp.local.")
    At 317.194232ms: ServiceFound("_matter._tcp.local.", "2C755BDFDBFDF8A2-000000007D343903._matter._tcp.local.")
    At 522.524688ms: ServiceFound("_matter._tcp.local.", "2C755BDFDBFDF8A2-00000000DFCB37CC._matter._tcp.local.")
    At 523.018225ms: Resolved a new service: 2C755BDFDBFDF8A2-00000000DFCB37CC._matter._tcp.local.
     host: 3C8D20E19E8E.local.
     port: 5540
     Address: fd96:7000:b73c:b0d9:ad84:5968:6560:c59e
     Address: 192.168.86.78
     Address: fe80::2425:600b:ae14:e74f%wlp170s0
     Address: fd96:7000:b73c:b0d9:6f7f:9d00:64c8:e17
     Address: fd96:7000:b73c:b0d9:3235:9a93:ab8d:172c
    At 523.104915ms: ServiceFound("_matter._tcp.local.", "2C755BDFDBFDF8A2-0000000000840C83._matter._tcp.local.")
    At 523.402917ms: Resolved a new service: 2C755BDFDBFDF8A2-0000000000840C83._matter._tcp.local.
     host: 14C14EDFB074.local.
     port: 5540
     Address: fe80::4de2:ea09:ee05:7dda%wlp170s0
     Address: 192.168.86.80
     Address: fd96:7000:b73c:b0d9:f8b:17ee:9dc6:b361
     Address: fd96:7000:b73c:b0d9:1015:5f20:59ee:7c33
     Address: fd96:7000:b73c:b0d9:48d8:3273:e105:6916
    At 1.046694972s: SearchStarted("_matter._tcp.local. on 2 interfaces [wlp170s0 (3), lo (1)]")
    At 3.047705459s: SearchStarted("_matter._tcp.local. on 2 interfaces [wlp170s0 (3), lo (1)]")
    At 3.387661735s: ServiceFound("_matter._tcp.local.", "2C755BDFDBFDF8A2-000000006EC96BF5._matter._tcp.local.")
    At 3.387926987s: Resolved a new service: 2C755BDFDBFDF8A2-000000006EC96BF5._matter._tcp.local.
     host: ECDA3B081E84.local.
     port: 5540
     Address: fd96:7000:b73c:b0d9:eeda:3bff:fe08:1e84
     Address: fe80::eeda:3bff:fe08:1e84%wlp170s0
     Address: 192.168.86.28
     Property: SII=5000
     Property: SAI=300
     Property: T=1
    At 7.049050317s: SearchStarted("_matter._tcp.local. on 2 interfaces [wlp170s0 (3), lo (1)]")
    At 7.285577326s: ServiceFound("_matter._tcp.local.", "2C755BDFDBFDF8A2-00000000F141DDE6._matter._tcp.local.")
    At 7.285753229s: Resolved a new service: 2C755BDFDBFDF8A2-00000000F141DDE6._matter._tcp.local.
     host: F4F5D8C0DCE6.local.
     port: 5540
     Address: fd96:7000:b73c:b0d9:d066:7e60:3058:c6ed
    
  4. hemanpa commented on Aug 18, 2026

    @hemanpa
    Contributor

    That's exactly the data I needed — thank you. Your logs identify the fault, and it isn't the aliasing bug I fixed in 0.7.1. That one was real, but it wasn't yours. Sorry for the detour.

    What your logs show

    Your mdns-sd query run is the decisive part:

    At 317.174405ms: ServiceFound(…, "F52AC107C954E38E-0000000000000002._matter._tcp.local.")
    At 317.181164ms: ServiceFound(…, "F52AC107C954E38E-0000000000000003._matter._tcp.local.")
    At 317.187235ms: ServiceFound(…, "F52AC107C954E38E-0000000000000004._matter._tcp.local.")
    

    All three of your nodes are found in the first 317 ms — but none of them is ever resolved. Across the whole run, 18 instances are found and only a handful reach Resolved a new service, roughly one per query cycle (523 ms, 3.4 s, 7.3 s — each just after a SearchStarted, on mdns-sd's exponential backoff).

    ServiceFound means the PTR record arrived. ServiceResolved means SRV + address resolution completed. Our adapter only surfaces ServiceResolved, so an instance stuck at "found" is invisible to us — and with 18 instances on your network resolving at roughly one per cycle, yours simply never make it inside the 30-second budget. That is precisely your ~5% success rate: it works when yours happen to come up early.

    Your debug log confirms it from our side, and rules out our filters: 6 records surfaced, 0 drops — nothing was discarded for an unrecognised service type, a malformed name, or missing addresses. The six that did surface are all 2C755BDFDBFDF8A2… / 9B5791331E047398…, i.e. other fabrics. Yours never arrive at all.

    For completeness: this is not fixed by a newer mdns-sd — I read the 0.20.3 and 0.21.0 changelogs and neither touches this.

    The lever, and a request

    Matter defines a subtype exactly for this: operational nodes also advertise under _I<compressed-fabric-id>._sub._matter._tcp, so a controller can ask for its own fabric rather than every operational node on the LAN. We browse the base type today, which is why we inherit the whole neighbourhood's resolution queue.

    I verified the subtype works with mdns-sd against real devices here — it narrowed a 16-instance browse to the single node on my fabric, resolved immediately:

    --- full browse: _matter._tcp.local. ---
      ServiceFound: 16   ServiceResolved: 16
    --- subtype browse: _IC701306F5C36CEBB._sub._matter._tcp.local. ---
      ServiceFound: 1    ServiceResolved: 1
    

    But my network resolves all 16 fine, so I cannot reproduce your stall and I don't want to ship a fix I can't verify against the failure. You can settle it in one command:

    cargo run --example query _IF52AC107C954E38E._sub._matter._tcp

    If your three nodes resolve promptly there while the plain _matter._tcp browse still leaves them stuck at ServiceFound, that confirms it and I'll switch operational discovery to the fabric subtype (falling back to the base type, since not every responder publishes subtypes).

    If they don't resolve even under the subtype, then the stall is in resolving those specific instances rather than a queue-depth effect, and I'll take it upstream to mdns-sd with your trace — in which case the fix on our side is likely to surface ServiceFound and drive resolution ourselves rather than waiting.

    Either way your report has already paid for itself twice over: it found a genuine aliasing bug, and it found that our discovery path had no diagnostics at all.

  5. qwandor commented on Aug 18, 2026

    @qwandor
    ContributorAuthor
    $ cargo run --example query _IF52AC107C954E38E._sub._matter._tcp
        Finished `dev` profile [unoptimized + debuginfo] target(s) in 0.07s
         Running `target/debug/examples/query _IF52AC107C954E38E._sub._matter._tcp`
    At 603.94µs: SearchStarted("_IF52AC107C954E38E._sub._matter._tcp.local. on 2 interfaces [wlp170s0 (2), lo (1)]")
    At 33.948397ms: SearchStarted("_IF52AC107C954E38E._sub._matter._tcp.local. on 2 interfaces [wlp170s0 (2), lo (1)]")
    At 266.493871ms: ServiceFound("_IF52AC107C954E38E._sub._matter._tcp.local.", "F52AC107C954E38E-0000000000000002._matter._tcp.local.")
    At 266.553504ms: ServiceFound("_IF52AC107C954E38E._sub._matter._tcp.local.", "F52AC107C954E38E-0000000000000003._matter._tcp.local.")
    At 266.557897ms: ServiceFound("_IF52AC107C954E38E._sub._matter._tcp.local.", "F52AC107C954E38E-0000000000000004._matter._tcp.local.")
    At 266.740761ms: Resolved a new service: F52AC107C954E38E-0000000000000004._matter._tcp.local.
     host: CED8B66A8876B184.local.
     port: 5540
     Address: fd36:24e1:62ea:1:2c2:82f:81d1:a9d9
     Property: SAI=2000
     Property: SAT=4000
     Property: SII=800
    At 266.799295ms: Resolved a new service: F52AC107C954E38E-0000000000000003._matter._tcp.local.
     host: C2AAA66B72F7A3EB.local.
     port: 5540
     Address: fd36:24e1:62ea:1:381e:f2ae:5150:61c8
     Property: SAI=2500
     Property: SAT=1000
     Property: SII=15800
    At 266.825998ms: Resolved a new service: F52AC107C954E38E-0000000000000002._matter._tcp.local.
     host: D29F10CD4A981C8B.local.
     port: 5540
     Address: fd36:24e1:62ea:1:8c7a:7b39:1883:c1b7
     Property: SAI=2000
     Property: SAT=4000
     Property: SII=800
    At 1.034450269s: SearchStarted("_IF52AC107C954E38E._sub._matter._tcp.local. on 2 interfaces [wlp170s0 (2), lo (1)]")
    At 3.037044957s: SearchStarted("_IF52AC107C954E38E._sub._matter._tcp.local. on 2 interfaces [wlp170s0 (2), lo (1)]")
    At 7.038412644s: SearchStarted("_IF52AC107C954E38E._sub._matter._tcp.local. on 2 interfaces [wlp170s0 (2), lo (1)]")
    
  6. qwandor commented on Aug 18, 2026

    @qwandor
    ContributorAuthor

    It looks like the subtype works? But it still seems odd that the plain _matter._tcp isn't working. Do you think that's an mdns-sd bug?

  7. qwandor commented on Aug 18, 2026

    @qwandor
    ContributorAuthor

    I also wonder whether running a separate mDNS implementation from mdns-sd is the best way to go here, rather than using Avahi or whatever OS implementation is already provided. It looks like there are two Rust crates for doing that, https://crates.io/crates/mdns-sd-discovery and https://crates.io/crates/zeroconf.

  8. qwandor commented on Aug 18, 2026

    @qwandor
    ContributorAuthor

    Also, if you have a potential fix which you'd like me to test, feel free to push it to a branch here and I can try building against that before you merge it. It won't be until next week though, as I'll be away for the rest of the week.

  9. hemanpa commented on Aug 19, 2026

    @hemanpa
    Contributor

    That's conclusive — thank you. All three nodes resolved at ~266 ms under the subtype, where the base type never resolves them at all.

    Branch ready to test: fix/113-operational-subtype (commit c095afcc).

    matter-controller = { git = "https://github.com/phunapps/matter-rust", branch = "fix/113-operational-subtype" }

    No API change; operational discovery just asks for your fabric instead of the whole neighbourhood. Both browses are opened and whichever delivers first settles the resolve, so a responder that publishes no subtype still works exactly as before — I didn't want to trade your bug for a silent "finds nothing" on someone else's setup.

    One detail that nearly made the fix a no-op, in case it's useful to you: mdns-sd reports the browsed string as a resolved record's ty_domain, so subtype records arrive tagged _I….._sub._matter._tcp.local., and our exact-match type check was dropping every one of them. That's handled and pinned by a test now.

    Verified here against a real Thread device: subtype query on the wire, reconnect and a 204-attribute read both fine. But my network resolves the base type fine, so your run is the one that actually tests the fix. No rush — next week is fine.

    Is it an mdns-sd bug?

    I think so, yes. Your trace shows 18 instances found in the first 317 ms and then resolved at roughly one per query cycle, on exponential backoff — 523 ms, 3.4 s, 7.3 s. Discovery is fast; completing SRV/address resolution for each instance is what crawls. With 18 instances that's minutes, which is why yours never made a 30-second budget and why it looked like a ~5% coin flip.

    I read the 0.20.3 and 0.21.0 changelogs and neither touches this. I'd like to report it upstream with your trace (credited, obviously) — the subtype makes it moot for Matter, but anyone browsing a busy service type will hit the same wall.

    On using Avahi / the OS implementation instead

    A fair question, and you've found a real gap. matter-transport already defines a public Discovery trait, so an Avahi- or Bonjour-backed implementation is entirely possible — but MatterController doesn't currently let you supply one; it constructs MdnsSdDiscovery internally and the injection seam is pub(crate). So today you'd have the trait and no way to use it, which isn't much good to you.

    Opening that seam is additive and cheap, and I'm inclined to do it regardless of this issue — it would let you run zeroconf or a direct Avahi client without waiting on us, and it makes the mDNS stack a choice rather than something we impose. Say the word if that's useful and I'll raise it as a separate issue.

    On the default itself: zeroconf/mdns-sd-discovery bind to the platform daemon, which is more battle-tested and avoids running a second responder next to avahi-daemon — but pulls a C dependency and behaves differently per platform, which is awkward for a library that has to work identically on Linux, macOS and embedded targets. My instinct is to keep the pure-Rust default and make it swappable rather than switch wholesale, but I'd genuinely weigh a strong argument the other way.

    (Small note if you re-run with logs: our own examples don't install a tracing subscriber, so RUST_LOG only does anything inside an application that sets one up — yours clearly does.)

  10. hemanpa commented on Aug 19, 2026

    @hemanpa
    Contributor

    Both done.

    1. Reported upstream: keepsimple1/mdns-sd#493, with your trace and credited to you. I was explicit that I can't reproduce it on my own network — 16 instances resolve 16/16 in under 500 ms here — and asked whether one-instance-per-query-cycle is expected for a browse with many instances, or whether follow-up SRV/address queries should be more aggressive when several are pending. I mentioned you'd offered to gather more data if they want it.

    2. Discovery injection is on main (67d638fc), so you can bring your own mDNS stack:

    MatterController::builder(store)
        .discovery(my_discovery)   // any impl of matter_transport::Discovery
        .build()
        .await?

    Additive — the default is unchanged, and builder(store).build() still gives you mdns-sd with no turbofish. It's a generic method rather than a boxed trait object on purpose: a delegating shim would forward only the methods it explicitly writes, so a defaulted trait method it missed would silently fall back to the default. Concretely, that would mean an Avahi backend's query_operational_fabric override being quietly bypassed and losing the subtype narrowing, with no compile error. Monomorphising avoids that class entirely.

    One scope limit, called out rather than buried: the injected Discovery is used for the controller's own resolution, but not by the OTA provider, serve_provider_once, or the ICD check-in listener. Those run deliberately off the actor on their own sockets and each need an exclusively-owned instance, so they keep constructing their own MdnsSdDiscovery. Closing that properly means taking a discovery factory (.discovery_with(|| …)) instead of a value. I didn't want to guess at the API on your behalf — if you end up using those paths with a non-default backend, say so here and I'll switch it.

    On the default: I'm keeping pure-Rust mdns-sd as the out-of-the-box choice rather than switching to zeroconf, mostly because a C dependency with per-platform behaviour is awkward for a library that has to work the same on Linux, macOS and embedded targets. But that's now your call rather than mine, which was the actual problem with the previous state — the trait was public and there was no way to use it.

    The subtype fix is still on fix/113-operational-subtype waiting on your run whenever you're back. No hurry.

  11. qwandor commented on Aug 19, 2026

    @qwandor
    ContributorAuthor

    I've tested your branch and the subtype fix doesn't seem to have helped unfortunately. I still get errors:

    operational error: connect to node 2 failed: driver error: device discovery failed: operational node F52AC107C954E38E-0000000000000002 not found via mDNS (saw 6 operational mDNS record(s), none matching: 2c755bdfdbfdf8a2-0000000000840c83, 2c755bdfdbfdf8a2-000000006ec96bf5, 2c755bdfdbfdf8a2-00000000dfcb37cc, 2c755bdfdbfdf8a2-00000000f141dde6, 2c755bdfdbfdf8a2-00000000f25ef2e4, … 1 more)
    operational error: connect to node 3 failed: driver error: device discovery failed: operational node F52AC107C954E38E-0000000000000003 not found via mDNS (saw 6 operational mDNS record(s), none matching: 2c755bdfdbfdf8a2-0000000000840c83, 2c755bdfdbfdf8a2-000000006ec96bf5, 2c755bdfdbfdf8a2-00000000dfcb37cc, 2c755bdfdbfdf8a2-00000000f141dde6, 2c755bdfdbfdf8a2-00000000f25ef2e4, … 1 more)
    operational error: connect to node 4 failed: driver error: device discovery failed: operational node F52AC107C954E38E-0000000000000004 not found via mDNS (saw 6 operational mDNS record(s), none matching: 2c755bdfdbfdf8a2-0000000000840c83, 2c755bdfdbfdf8a2-000000006ec96bf5, 2c755bdfdbfdf8a2-00000000dfcb37cc, 2c755bdfdbfdf8a2-00000000f141dde6, 2c755bdfdbfdf8a2-00000000f25ef2e4, … 1 more)
    

    Logs: log.txt

    I'll test further next week.

  12. hemanpa commented on Aug 19, 2026

    @hemanpa
    Contributor

    Your test found the flaw, and it was mine — thank you for running it.

    The fallback was the problem. The branch opened the subtype browse and the base _matter._tcp browse together, so a responder that publishes no subtype would still work. Your log shows what that actually did:

    mDNS browse started service_type=_IF52AC107C954E38E._sub._matter._tcp.local.
    mDNS browse started service_type=_matter._tcp.local.
    mDNS record surfaced browse="_matter._tcp.local." subtype_browse=false  … ×6
    

    Every record came from the base browse; the subtype browse produced nothing at all in 30 s, with zero drops — so its records never arrived rather than being filtered by us.

    The reason, given the underlying mdns-sd behaviour: the subtype helps only by limiting how many instances get discovered. Opening the base browse re-discovers all 18, which puts your three straight back into the same one-per-cycle resolution queue. A subtype browse can't hand you a resolved record for an instance mdns-sd hasn't resolved yet. So my "safe" fallback reintroduced precisely the bug the subtype was meant to sidestep.

    Worth saying plainly: my rig couldn't have caught this. Both browses concurrently work fine here — subtype resolves 1, base resolves 16 — because everything on my network resolves in under half a second. A network where the base type is pathological is the only place the difference shows, which is exactly why v1 got past me.

    Corrected — same branch, commit 6bd88693

    The subtype is now the only browse for the first 2 seconds. The base browse opens only if nothing has matched by then, so a subtype-less responder still resolves with ~28 s of budget left. 2 s is about 8× the healthy subtype response time (~266 ms on yours, ~250 ms on mine) and absorbs a retransmit, so the fallback shouldn't fire for you at all.

    Verified on the wire here: 13 subtype queries, 0 base-type queries — the fallback correctly stayed shut — and the device reconnected.

    Both regression tests were confirmed to fail against the old design before being kept, so this specific mistake can't come back silently.

    When you get a chance, the same RUST_LOG will tell us which path won: records now carry browse= / subtype_browse=, and there's a new debug line if the fallback opens ("no subtype match after 2s"). If you see that line, the subtype browse isn't answering for you at all and this is a different problem again — in which case I'd want to know, because your manual subtype run resolving all three in 266 ms says it should.

    Upstream report is keepsimple1/mdns-sd#493 if you want to follow along; no response yet.

  13. hemanpa commented on Aug 20, 2026

    @hemanpa
    Contributor

    Upstream moved fast: the mdns-sd maintainer confirmed it looks like a real bug and has opened keepsimple1/mdns-sd#494.

    His diagnosis matches your trace precisely: resolution of a found instance is driven by a per-instance Command::Resolve retry chain that fires every 500 ms and gives up after 3 tries (~1.5 s). After that the only route to resolution is an unsolicited or browse-elicited cache update — which is exactly why your resolutions appeared one at a time, immediately after each SearchStarted. The fix drives follow-up SRV/address queries off the browse retransmission cycle instead.

    When you're back, that PR is worth testing before our workaround — if it fixes it at the source, it fixes it for the plain _matter._tcp browse and for anyone else hitting this, not just Matter. To try it:

    [patch.crates-io]
    mdns-sd = { git = "https://github.com/keepsimple1/mdns-sd", branch = "fix/issue-493" }

    One wrinkle: the branch is 0.21.0 and we currently require 0.20, so the patch is ignored until you also bump that — easiest is to test it against our fix/113-operational-subtype branch with matter-transport's mdns-sd requirement changed to "0.21". I verified 0.21 works fine with our adapter, so that bump is safe.

    I ran the patch here and reported back on the PR, but only as no-regression evidence — 16/16 instances still resolve, the subtype still resolves, and a real reconnect plus a 204-attribute read are unaffected. My network can't confirm the fix because it never had the problem. Yours is the only real test, for the upstream patch as much as for our branch.

    So there are now two independent things you could try, and it'd be genuinely useful to know which is doing the work:

    1. upstream patch alone, on plain main of ours — does the base _matter._tcp browse now resolve your nodes?
    2. our branch alone (fix/113-operational-subtype, now at 6bd88693 with the subtype-exclusive window) — does the subtype path resolve them?

    If (1) works, the right outcome is probably that we wait for an mdns-sd release and keep the subtype narrowing as a belt-and-braces improvement rather than a workaround. No rush on either.

  14. qwandor commented on Aug 25, 2026

    @qwandor
    ContributorAuthor

    fix/113-operational-subtype now works, it connects to the devices quickly. The upstream patch to mdns-sd doesn't seem to help. I'll post logs for that on the upstream issue.

  15. hemanpa commented on Aug 25, 2026

    @hemanpa
    Contributor

    Thank you — that confirmation is the whole ballgame. This bug was invisible on our own LAN, where every instance resolves fast enough to hide it, so your network was the only real test it ever had.

    Merged and released:

    Crate Version
    matter-transport 0.4.0
    matter-commissioning 0.6.0
    matter-controller 0.10.0
    matter-controller = "0.10"

    Operational resolves now browse _I<compressed-fabric-id>._sub._matter._tcp (Matter Core Spec §4.3.1), narrowing discovery to our own fabric, with the base _matter._tcp browse kept only as a delayed fallback. The delay is the part your first test earned: v1 ran both browses concurrently, the base browse re-discovered all 18 instances, and resolution starved exactly as before. It now opens only once the subtype window closes with nothing parked — or up front if the subtype browse cannot be opened at all.

    Version note: matter-transport 0.4.0 is additive (it gains operational_fabric_subtype), but a 0.x minor is a compatibility break, so matter-commissioning and matter-controller moved with it. If you name matter_transport types in your own code — implementing Discovery, say — you will need to bump that dependency in step.

    Noted that the upstream mdns-sd patch doesn't help, and thanks for taking the logs to keepsimple1/mdns-sd#494 directly. Worth saying plainly: the subtype browse is the right fix regardless of what upstream lands. Browsing our own fabric's subtype is what the spec defines it for, and it means we do not depend on any particular resolver's fairness under load. If upstream does improve the retry chain, that is a second layer of robustness under this, not a replacement.

    Closing as fixed — please do reopen if it resurfaces.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions