Repository navigation
mDNS rediscovery fails for previously commissioned devices #113
Description
Activity
Thanks — this one was worth chasing, and it turned up a real bug in our mDNS adapter. I want to be straight with you though: I could not reproduce your failure here, so I can't yet claim this is your fault. What I can say is that what I found is consistent with everything you describe.
The bug
mdns-sdkeys its listeners by service type —service_queriers: HashMap<ty_domain, Sender>— so there is exactly one listener per type. Our adapter handed out aQueryHandleperquery()call and stored aReceiverper handle, as though they were independent browses. They aren't:- a second
browse("_matter._tcp.local.")replaces the first sender, orphaning the firstReceiver— it never receives again; stop_browsethen removes the shared querier and wipes that type's cache.
So any second operational resolve on the same
Discoverycould silently kill the controller's own browse. Afterwards the actor still held aSome(handle)and therefore never re-browsed, and every later resolve expired withnot found via mDNS— permanently, for the life of the process, whileavahi-browsekept showing the records. That is a very close match to "almost always fails after a restart, occasionally works, avahi sees everything fine".Fixed in matter-controller 0.7.1 (with matter-transport 0.3.2, matter-commissioning 0.5.2): one browse per service type, reference-counted and fanned out to every handle, with
stop_browseonly on the last release. Records already surfaced are replayed to a handle that attaches later — without that, a late handle would see nothing, because mdns-sd only emitsServiceResolvedfor changed records.Validated on a live LAN with ~14 operational records: two concurrent handles both receive the full set, releasing one leaves the other's browse alive, and a real device reconnects across fresh processes.
If it still happens, this release will tell us why
The honest problem with your original report is that our discovery path had no logging at all, and four places where a record could be dropped silently — so the failure was undiagnosable from the outside. That's fixed too.
If you can reproduce on 0.7.1, please run with:
RUST_LOG=matter_transport=debug,matter_controller=debugand paste the output around a failed connect. You'll get a line per browse, per handle attach, and per record surfaced, plus an explicit line for each dropped record with the reason — including the record's
ty_domainwhen it isn't recognised. That last one matters: Matter advertises operational records under_I<compressed-fabric>._sub._matter._tcpsubtypes, and if mdns-sd ever surfaces one under the subtype rather than the base type, we'd drop it silently today. I couldn't trigger that here, but the log will say so immediately if it's what you're hitting.The error message itself is also more useful now — it reports what discovery actually saw:
… not found via mDNS (saw 3 operational mDNS record(s), none matching: f52ac107c954e38e-0000000000000003, …)versus
(saw 0 operational mDNS records …). That alone distinguishes "nothing reached us" from "records reached us but yours wasn't among them", which is the first fork in diagnosing this.Three things that would help if it persists: whether
avahi-daemonis running alongside your service, whether the affected devices are IPv4, IPv6 or dual-stack, and whether the same binary fails against all your devices or only some.- a second
Thanks for the quick response! Unfortunately the issue still seems to happen with the new version.
The errors are now:
operational error: connect to node 2 failed: driver error: device discovery failed: operational node F52AC107C954E38E-0000000000000002 not found via mDNS (saw 6 operational mDNS record(s), none matching: 2c755bdfdbfdf8a2-0000000000840c83, 2c755bdfdbfdf8a2-000000006ec96bf5, 2c755bdfdbfdf8a2-00000000dfcb37cc, 2c755bdfdbfdf8a2-00000000f141dde6, 2c755bdfdbfdf8a2-00000000f25ef2e4, … 1 more) operational error: connect to node 3 failed: driver error: device discovery failed: operational node F52AC107C954E38E-0000000000000003 not found via mDNS (saw 6 operational mDNS record(s), none matching: 2c755bdfdbfdf8a2-0000000000840c83, 2c755bdfdbfdf8a2-000000006ec96bf5, 2c755bdfdbfdf8a2-00000000dfcb37cc, 2c755bdfdbfdf8a2-00000000f141dde6, 2c755bdfdbfdf8a2-00000000f25ef2e4, … 1 more) operational error: connect to node 4 failed: driver error: device discovery failed: operational node F52AC107C954E38E-0000000000000004 not found via mDNS (saw 6 operational mDNS record(s), none matching: 2c755bdfdbfdf8a2-0000000000840c83, 2c755bdfdbfdf8a2-000000006ec96bf5, 2c755bdfdbfdf8a2-00000000dfcb37cc, 2c755bdfdbfdf8a2-00000000f141dde6, 2c755bdfdbfdf8a2-00000000f25ef2e4, … 1 more)All three devices I tested with connect over Thread, which I assume means they are IPv6 only.
I've attached the debug logs: log.txt
For additional context, I tried running the
queryexample from themdns-sdcrate, and that seems to see the nodes right away:$ cargo run --example query _matter._tcp Finished `dev` profile [unoptimized + debuginfo] target(s) in 0.02s Running `target/debug/examples/query _matter._tcp` At 322.513µs: SearchStarted("_matter._tcp.local. on 2 interfaces [wlp170s0 (3), lo (1)]") At 46.647774ms: SearchStarted("_matter._tcp.local. on 2 interfaces [wlp170s0 (3), lo (1)]") At 317.055798ms: ServiceFound("_matter._tcp.local.", "2C755BDFDBFDF8A2-00000000AAB08AA7._matter._tcp.local.") At 317.094102ms: ServiceFound("_matter._tcp.local.", "2C755BDFDBFDF8A2-00000000056BB72A._matter._tcp.local.") At 317.105053ms: ServiceFound("_matter._tcp.local.", "2C755BDFDBFDF8A2-000000000BD7BD08._matter._tcp.local.") At 317.111374ms: ServiceFound("_matter._tcp.local.", "2C755BDFDBFDF8A2-0000000077A6EFB7._matter._tcp.local.") At 317.115939ms: ServiceFound("_matter._tcp.local.", "2C755BDFDBFDF8A2-000000009B2EEF43._matter._tcp.local.") At 317.123104ms: ServiceFound("_matter._tcp.local.", "2C755BDFDBFDF8A2-000000004D852B21._matter._tcp.local.") At 317.129643ms: ServiceFound("_matter._tcp.local.", "2C755BDFDBFDF8A2-00000000642B4839._matter._tcp.local.") At 317.135777ms: ServiceFound("_matter._tcp.local.", "2C755BDFDBFDF8A2-FFFF000000000000._matter._tcp.local.") At 317.142114ms: ServiceFound("_matter._tcp.local.", "2C755BDFDBFDF8A2-0000000061817431._matter._tcp.local.") At 317.148434ms: ServiceFound("_matter._tcp.local.", "2C755BDFDBFDF8A2-000000004192B362._matter._tcp.local.") At 317.154801ms: ServiceFound("_matter._tcp.local.", "2C755BDFDBFDF8A2-00000000F8CD79FE._matter._tcp.local.") At 317.161051ms: ServiceFound("_matter._tcp.local.", "9B5791331E047398-000000000000012D._matter._tcp.local.") At 317.167035ms: ServiceFound("_matter._tcp.local.", "9B5791331E047398-000000000000012E._matter._tcp.local.") At 317.174405ms: ServiceFound("_matter._tcp.local.", "F52AC107C954E38E-0000000000000002._matter._tcp.local.") At 317.181164ms: ServiceFound("_matter._tcp.local.", "F52AC107C954E38E-0000000000000003._matter._tcp.local.") At 317.187235ms: ServiceFound("_matter._tcp.local.", "F52AC107C954E38E-0000000000000004._matter._tcp.local.") At 317.194232ms: ServiceFound("_matter._tcp.local.", "2C755BDFDBFDF8A2-000000007D343903._matter._tcp.local.") At 522.524688ms: ServiceFound("_matter._tcp.local.", "2C755BDFDBFDF8A2-00000000DFCB37CC._matter._tcp.local.") At 523.018225ms: Resolved a new service: 2C755BDFDBFDF8A2-00000000DFCB37CC._matter._tcp.local. host: 3C8D20E19E8E.local. port: 5540 Address: fd96:7000:b73c:b0d9:ad84:5968:6560:c59e Address: 192.168.86.78 Address: fe80::2425:600b:ae14:e74f%wlp170s0 Address: fd96:7000:b73c:b0d9:6f7f:9d00:64c8:e17 Address: fd96:7000:b73c:b0d9:3235:9a93:ab8d:172c At 523.104915ms: ServiceFound("_matter._tcp.local.", "2C755BDFDBFDF8A2-0000000000840C83._matter._tcp.local.") At 523.402917ms: Resolved a new service: 2C755BDFDBFDF8A2-0000000000840C83._matter._tcp.local. host: 14C14EDFB074.local. port: 5540 Address: fe80::4de2:ea09:ee05:7dda%wlp170s0 Address: 192.168.86.80 Address: fd96:7000:b73c:b0d9:f8b:17ee:9dc6:b361 Address: fd96:7000:b73c:b0d9:1015:5f20:59ee:7c33 Address: fd96:7000:b73c:b0d9:48d8:3273:e105:6916 At 1.046694972s: SearchStarted("_matter._tcp.local. on 2 interfaces [wlp170s0 (3), lo (1)]") At 3.047705459s: SearchStarted("_matter._tcp.local. on 2 interfaces [wlp170s0 (3), lo (1)]") At 3.387661735s: ServiceFound("_matter._tcp.local.", "2C755BDFDBFDF8A2-000000006EC96BF5._matter._tcp.local.") At 3.387926987s: Resolved a new service: 2C755BDFDBFDF8A2-000000006EC96BF5._matter._tcp.local. host: ECDA3B081E84.local. port: 5540 Address: fd96:7000:b73c:b0d9:eeda:3bff:fe08:1e84 Address: fe80::eeda:3bff:fe08:1e84%wlp170s0 Address: 192.168.86.28 Property: SII=5000 Property: SAI=300 Property: T=1 At 7.049050317s: SearchStarted("_matter._tcp.local. on 2 interfaces [wlp170s0 (3), lo (1)]") At 7.285577326s: ServiceFound("_matter._tcp.local.", "2C755BDFDBFDF8A2-00000000F141DDE6._matter._tcp.local.") At 7.285753229s: Resolved a new service: 2C755BDFDBFDF8A2-00000000F141DDE6._matter._tcp.local. host: F4F5D8C0DCE6.local. port: 5540 Address: fd96:7000:b73c:b0d9:d066:7e60:3058:c6edThat's exactly the data I needed — thank you. Your logs identify the fault, and it isn't the aliasing bug I fixed in 0.7.1. That one was real, but it wasn't yours. Sorry for the detour.
What your logs show
Your
mdns-sdqueryrun is the decisive part:At 317.174405ms: ServiceFound(…, "F52AC107C954E38E-0000000000000002._matter._tcp.local.") At 317.181164ms: ServiceFound(…, "F52AC107C954E38E-0000000000000003._matter._tcp.local.") At 317.187235ms: ServiceFound(…, "F52AC107C954E38E-0000000000000004._matter._tcp.local.")All three of your nodes are found in the first 317 ms — but none of them is ever resolved. Across the whole run, 18 instances are found and only a handful reach
Resolved a new service, roughly one per query cycle (523 ms, 3.4 s, 7.3 s — each just after aSearchStarted, on mdns-sd's exponential backoff).ServiceFoundmeans the PTR record arrived.ServiceResolvedmeans SRV + address resolution completed. Our adapter only surfacesServiceResolved, so an instance stuck at "found" is invisible to us — and with 18 instances on your network resolving at roughly one per cycle, yours simply never make it inside the 30-second budget. That is precisely your ~5% success rate: it works when yours happen to come up early.Your debug log confirms it from our side, and rules out our filters: 6 records surfaced, 0 drops — nothing was discarded for an unrecognised service type, a malformed name, or missing addresses. The six that did surface are all
2C755BDFDBFDF8A2…/9B5791331E047398…, i.e. other fabrics. Yours never arrive at all.For completeness: this is not fixed by a newer mdns-sd — I read the 0.20.3 and 0.21.0 changelogs and neither touches this.
The lever, and a request
Matter defines a subtype exactly for this: operational nodes also advertise under
_I<compressed-fabric-id>._sub._matter._tcp, so a controller can ask for its own fabric rather than every operational node on the LAN. We browse the base type today, which is why we inherit the whole neighbourhood's resolution queue.I verified the subtype works with mdns-sd against real devices here — it narrowed a 16-instance browse to the single node on my fabric, resolved immediately:
--- full browse: _matter._tcp.local. --- ServiceFound: 16 ServiceResolved: 16 --- subtype browse: _IC701306F5C36CEBB._sub._matter._tcp.local. --- ServiceFound: 1 ServiceResolved: 1But my network resolves all 16 fine, so I cannot reproduce your stall and I don't want to ship a fix I can't verify against the failure. You can settle it in one command:
cargo run --example query _IF52AC107C954E38E._sub._matter._tcp
If your three nodes resolve promptly there while the plain
_matter._tcpbrowse still leaves them stuck atServiceFound, that confirms it and I'll switch operational discovery to the fabric subtype (falling back to the base type, since not every responder publishes subtypes).If they don't resolve even under the subtype, then the stall is in resolving those specific instances rather than a queue-depth effect, and I'll take it upstream to mdns-sd with your trace — in which case the fix on our side is likely to surface
ServiceFoundand drive resolution ourselves rather than waiting.Either way your report has already paid for itself twice over: it found a genuine aliasing bug, and it found that our discovery path had no diagnostics at all.
$ cargo run --example query _IF52AC107C954E38E._sub._matter._tcp Finished `dev` profile [unoptimized + debuginfo] target(s) in 0.07s Running `target/debug/examples/query _IF52AC107C954E38E._sub._matter._tcp` At 603.94µs: SearchStarted("_IF52AC107C954E38E._sub._matter._tcp.local. on 2 interfaces [wlp170s0 (2), lo (1)]") At 33.948397ms: SearchStarted("_IF52AC107C954E38E._sub._matter._tcp.local. on 2 interfaces [wlp170s0 (2), lo (1)]") At 266.493871ms: ServiceFound("_IF52AC107C954E38E._sub._matter._tcp.local.", "F52AC107C954E38E-0000000000000002._matter._tcp.local.") At 266.553504ms: ServiceFound("_IF52AC107C954E38E._sub._matter._tcp.local.", "F52AC107C954E38E-0000000000000003._matter._tcp.local.") At 266.557897ms: ServiceFound("_IF52AC107C954E38E._sub._matter._tcp.local.", "F52AC107C954E38E-0000000000000004._matter._tcp.local.") At 266.740761ms: Resolved a new service: F52AC107C954E38E-0000000000000004._matter._tcp.local. host: CED8B66A8876B184.local. port: 5540 Address: fd36:24e1:62ea:1:2c2:82f:81d1:a9d9 Property: SAI=2000 Property: SAT=4000 Property: SII=800 At 266.799295ms: Resolved a new service: F52AC107C954E38E-0000000000000003._matter._tcp.local. host: C2AAA66B72F7A3EB.local. port: 5540 Address: fd36:24e1:62ea:1:381e:f2ae:5150:61c8 Property: SAI=2500 Property: SAT=1000 Property: SII=15800 At 266.825998ms: Resolved a new service: F52AC107C954E38E-0000000000000002._matter._tcp.local. host: D29F10CD4A981C8B.local. port: 5540 Address: fd36:24e1:62ea:1:8c7a:7b39:1883:c1b7 Property: SAI=2000 Property: SAT=4000 Property: SII=800 At 1.034450269s: SearchStarted("_IF52AC107C954E38E._sub._matter._tcp.local. on 2 interfaces [wlp170s0 (2), lo (1)]") At 3.037044957s: SearchStarted("_IF52AC107C954E38E._sub._matter._tcp.local. on 2 interfaces [wlp170s0 (2), lo (1)]") At 7.038412644s: SearchStarted("_IF52AC107C954E38E._sub._matter._tcp.local. on 2 interfaces [wlp170s0 (2), lo (1)]")It looks like the subtype works? But it still seems odd that the plain
_matter._tcpisn't working. Do you think that's anmdns-sdbug?I also wonder whether running a separate mDNS implementation from
mdns-sdis the best way to go here, rather than using Avahi or whatever OS implementation is already provided. It looks like there are two Rust crates for doing that, https://crates.io/crates/mdns-sd-discovery and https://crates.io/crates/zeroconf.Also, if you have a potential fix which you'd like me to test, feel free to push it to a branch here and I can try building against that before you merge it. It won't be until next week though, as I'll be away for the rest of the week.
That's conclusive — thank you. All three nodes resolved at ~266 ms under the subtype, where the base type never resolves them at all.
Branch ready to test:
fix/113-operational-subtype(commitc095afcc).matter-controller = { git = "https://github.com/phunapps/matter-rust", branch = "fix/113-operational-subtype" }
No API change; operational discovery just asks for your fabric instead of the whole neighbourhood. Both browses are opened and whichever delivers first settles the resolve, so a responder that publishes no subtype still works exactly as before — I didn't want to trade your bug for a silent "finds nothing" on someone else's setup.
One detail that nearly made the fix a no-op, in case it's useful to you: mdns-sd reports the browsed string as a resolved record's
ty_domain, so subtype records arrive tagged_I….._sub._matter._tcp.local., and our exact-match type check was dropping every one of them. That's handled and pinned by a test now.Verified here against a real Thread device: subtype query on the wire, reconnect and a 204-attribute read both fine. But my network resolves the base type fine, so your run is the one that actually tests the fix. No rush — next week is fine.
Is it an mdns-sd bug?
I think so, yes. Your trace shows 18 instances found in the first 317 ms and then resolved at roughly one per query cycle, on exponential backoff — 523 ms, 3.4 s, 7.3 s. Discovery is fast; completing SRV/address resolution for each instance is what crawls. With 18 instances that's minutes, which is why yours never made a 30-second budget and why it looked like a ~5% coin flip.
I read the 0.20.3 and 0.21.0 changelogs and neither touches this. I'd like to report it upstream with your trace (credited, obviously) — the subtype makes it moot for Matter, but anyone browsing a busy service type will hit the same wall.
On using Avahi / the OS implementation instead
A fair question, and you've found a real gap.
matter-transportalready defines a publicDiscoverytrait, so an Avahi- or Bonjour-backed implementation is entirely possible — butMatterControllerdoesn't currently let you supply one; it constructsMdnsSdDiscoveryinternally and the injection seam ispub(crate). So today you'd have the trait and no way to use it, which isn't much good to you.Opening that seam is additive and cheap, and I'm inclined to do it regardless of this issue — it would let you run
zeroconfor a direct Avahi client without waiting on us, and it makes the mDNS stack a choice rather than something we impose. Say the word if that's useful and I'll raise it as a separate issue.On the default itself:
zeroconf/mdns-sd-discoverybind to the platform daemon, which is more battle-tested and avoids running a second responder next toavahi-daemon— but pulls a C dependency and behaves differently per platform, which is awkward for a library that has to work identically on Linux, macOS and embedded targets. My instinct is to keep the pure-Rust default and make it swappable rather than switch wholesale, but I'd genuinely weigh a strong argument the other way.(Small note if you re-run with logs: our own examples don't install a
tracingsubscriber, soRUST_LOGonly does anything inside an application that sets one up — yours clearly does.)- added a commit that references this issue
on Aug 19, 2026 Both done.
1. Reported upstream: keepsimple1/mdns-sd#493, with your trace and credited to you. I was explicit that I can't reproduce it on my own network — 16 instances resolve 16/16 in under 500 ms here — and asked whether one-instance-per-query-cycle is expected for a browse with many instances, or whether follow-up SRV/address queries should be more aggressive when several are pending. I mentioned you'd offered to gather more data if they want it.
2.
Discoveryinjection is onmain(67d638fc), so you can bring your own mDNS stack:MatterController::builder(store) .discovery(my_discovery) // any impl of matter_transport::Discovery .build() .await?
Additive — the default is unchanged, and
builder(store).build()still gives youmdns-sdwith no turbofish. It's a generic method rather than a boxed trait object on purpose: a delegating shim would forward only the methods it explicitly writes, so a defaulted trait method it missed would silently fall back to the default. Concretely, that would mean an Avahi backend'squery_operational_fabricoverride being quietly bypassed and losing the subtype narrowing, with no compile error. Monomorphising avoids that class entirely.One scope limit, called out rather than buried: the injected
Discoveryis used for the controller's own resolution, but not by the OTA provider,serve_provider_once, or the ICD check-in listener. Those run deliberately off the actor on their own sockets and each need an exclusively-owned instance, so they keep constructing their ownMdnsSdDiscovery. Closing that properly means taking a discovery factory (.discovery_with(|| …)) instead of a value. I didn't want to guess at the API on your behalf — if you end up using those paths with a non-default backend, say so here and I'll switch it.On the default: I'm keeping pure-Rust
mdns-sdas the out-of-the-box choice rather than switching tozeroconf, mostly because a C dependency with per-platform behaviour is awkward for a library that has to work the same on Linux, macOS and embedded targets. But that's now your call rather than mine, which was the actual problem with the previous state — the trait was public and there was no way to use it.The subtype fix is still on
fix/113-operational-subtypewaiting on your run whenever you're back. No hurry.I've tested your branch and the subtype fix doesn't seem to have helped unfortunately. I still get errors:
operational error: connect to node 2 failed: driver error: device discovery failed: operational node F52AC107C954E38E-0000000000000002 not found via mDNS (saw 6 operational mDNS record(s), none matching: 2c755bdfdbfdf8a2-0000000000840c83, 2c755bdfdbfdf8a2-000000006ec96bf5, 2c755bdfdbfdf8a2-00000000dfcb37cc, 2c755bdfdbfdf8a2-00000000f141dde6, 2c755bdfdbfdf8a2-00000000f25ef2e4, … 1 more) operational error: connect to node 3 failed: driver error: device discovery failed: operational node F52AC107C954E38E-0000000000000003 not found via mDNS (saw 6 operational mDNS record(s), none matching: 2c755bdfdbfdf8a2-0000000000840c83, 2c755bdfdbfdf8a2-000000006ec96bf5, 2c755bdfdbfdf8a2-00000000dfcb37cc, 2c755bdfdbfdf8a2-00000000f141dde6, 2c755bdfdbfdf8a2-00000000f25ef2e4, … 1 more) operational error: connect to node 4 failed: driver error: device discovery failed: operational node F52AC107C954E38E-0000000000000004 not found via mDNS (saw 6 operational mDNS record(s), none matching: 2c755bdfdbfdf8a2-0000000000840c83, 2c755bdfdbfdf8a2-000000006ec96bf5, 2c755bdfdbfdf8a2-00000000dfcb37cc, 2c755bdfdbfdf8a2-00000000f141dde6, 2c755bdfdbfdf8a2-00000000f25ef2e4, … 1 more)Logs: log.txt
I'll test further next week.
- added a commit that references this issue
on Aug 19, 2026 Your test found the flaw, and it was mine — thank you for running it.
The fallback was the problem. The branch opened the subtype browse and the base
_matter._tcpbrowse together, so a responder that publishes no subtype would still work. Your log shows what that actually did:mDNS browse started service_type=_IF52AC107C954E38E._sub._matter._tcp.local. mDNS browse started service_type=_matter._tcp.local. mDNS record surfaced browse="_matter._tcp.local." subtype_browse=false … ×6Every record came from the base browse; the subtype browse produced nothing at all in 30 s, with zero drops — so its records never arrived rather than being filtered by us.
The reason, given the underlying mdns-sd behaviour: the subtype helps only by limiting how many instances get discovered. Opening the base browse re-discovers all 18, which puts your three straight back into the same one-per-cycle resolution queue. A subtype browse can't hand you a resolved record for an instance mdns-sd hasn't resolved yet. So my "safe" fallback reintroduced precisely the bug the subtype was meant to sidestep.
Worth saying plainly: my rig couldn't have caught this. Both browses concurrently work fine here — subtype resolves 1, base resolves 16 — because everything on my network resolves in under half a second. A network where the base type is pathological is the only place the difference shows, which is exactly why v1 got past me.
Corrected — same branch, commit
6bd88693The subtype is now the only browse for the first 2 seconds. The base browse opens only if nothing has matched by then, so a subtype-less responder still resolves with ~28 s of budget left. 2 s is about 8× the healthy subtype response time (~266 ms on yours, ~250 ms on mine) and absorbs a retransmit, so the fallback shouldn't fire for you at all.
Verified on the wire here: 13 subtype queries, 0 base-type queries — the fallback correctly stayed shut — and the device reconnected.
Both regression tests were confirmed to fail against the old design before being kept, so this specific mistake can't come back silently.
When you get a chance, the same
RUST_LOGwill tell us which path won: records now carrybrowse=/subtype_browse=, and there's a new debug line if the fallback opens ("no subtype match after 2s"). If you see that line, the subtype browse isn't answering for you at all and this is a different problem again — in which case I'd want to know, because your manual subtype run resolving all three in 266 ms says it should.Upstream report is keepsimple1/mdns-sd#493 if you want to follow along; no response yet.
Upstream moved fast: the mdns-sd maintainer confirmed it looks like a real bug and has opened keepsimple1/mdns-sd#494.
His diagnosis matches your trace precisely: resolution of a found instance is driven by a per-instance
Command::Resolveretry chain that fires every 500 ms and gives up after 3 tries (~1.5 s). After that the only route to resolution is an unsolicited or browse-elicited cache update — which is exactly why your resolutions appeared one at a time, immediately after eachSearchStarted. The fix drives follow-up SRV/address queries off the browse retransmission cycle instead.When you're back, that PR is worth testing before our workaround — if it fixes it at the source, it fixes it for the plain
_matter._tcpbrowse and for anyone else hitting this, not just Matter. To try it:[patch.crates-io] mdns-sd = { git = "https://github.com/keepsimple1/mdns-sd", branch = "fix/issue-493" }
One wrinkle: the branch is
0.21.0and we currently require0.20, so the patch is ignored until you also bump that — easiest is to test it against ourfix/113-operational-subtypebranch withmatter-transport'smdns-sdrequirement changed to"0.21". I verified 0.21 works fine with our adapter, so that bump is safe.I ran the patch here and reported back on the PR, but only as no-regression evidence — 16/16 instances still resolve, the subtype still resolves, and a real reconnect plus a 204-attribute read are unaffected. My network can't confirm the fix because it never had the problem. Yours is the only real test, for the upstream patch as much as for our branch.
So there are now two independent things you could try, and it'd be genuinely useful to know which is doing the work:
- upstream patch alone, on plain
mainof ours — does the base_matter._tcpbrowse now resolve your nodes? - our branch alone (
fix/113-operational-subtype, now at6bd88693with the subtype-exclusive window) — does the subtype path resolve them?
If (1) works, the right outcome is probably that we wait for an mdns-sd release and keep the subtype narrowing as a belt-and-braces improvement rather than a workaround. No rush on either.
- upstream patch alone, on plain
- added a commit that references this issue
on Aug 25, 2026 fix/113-operational-subtypenow works, it connects to the devices quickly. The upstream patch tomdns-sddoesn't seem to help. I'll post logs for that on the upstream issue.- added a commit that references this issue
on Aug 25, 2026 Thank you — that confirmation is the whole ballgame. This bug was invisible on our own LAN, where every instance resolves fast enough to hide it, so your network was the only real test it ever had.
Merged and released:
Crate Version matter-transport0.4.0 matter-commissioning0.6.0 matter-controller0.10.0 matter-controller = "0.10"
Operational resolves now browse
_I<compressed-fabric-id>._sub._matter._tcp(Matter Core Spec §4.3.1), narrowing discovery to our own fabric, with the base_matter._tcpbrowse kept only as a delayed fallback. The delay is the part your first test earned: v1 ran both browses concurrently, the base browse re-discovered all 18 instances, and resolution starved exactly as before. It now opens only once the subtype window closes with nothing parked — or up front if the subtype browse cannot be opened at all.Version note:
matter-transport0.4.0 is additive (it gainsoperational_fabric_subtype), but a0.xminor is a compatibility break, somatter-commissioningandmatter-controllermoved with it. If you namematter_transporttypes in your own code — implementingDiscovery, say — you will need to bump that dependency in step.Noted that the upstream
mdns-sdpatch doesn't help, and thanks for taking the logs to keepsimple1/mdns-sd#494 directly. Worth saying plainly: the subtype browse is the right fix regardless of what upstream lands. Browsing our own fabric's subtype is what the spec defines it for, and it means we do not depend on any particular resolver's fairness under load. If upstream does improve the retry chain, that is a second layer of robustness under this, not a replacement.Closing as fixed — please do reopen if it resurfaces.
- added a commit that references this issue
on Sep 8, 2026
With the issues in the other bugs fixed I am able to successfully commission a number of Matter devices and read attributes from them. However, when I restart my service and attempt to connect to the devices I previously commissioned again, it almost always fails with an error like "operational error: connect to node 2 failed: driver error: device discovery failed: operational node F52AC107C954E38E-0000000000000002 not found via mDNS". (It does occasionally work, but it seems to work less than 5% of the time.)
If I open
avahi-discoverI can see all these addresses advertised under_matter._tcpjust fine, so it seems to be something to do with how thematter-controllercrate does mDNS discovery in this case. As another data point, the same devices also reconnect fine if I use thematccrate instead (though that crate has other limitations, so I'd prefer to usematter-controller).The same issue happens for all the devices I've tried, and it only seems to happen on re-connection; the initial commissioning works fine most of the time. I'm running on Ubuntu 24.04.4, in case that makes a difference.