Runnable simulation of one-trace-per-message with batch fan-in via links.
The pipeline uses nothing but System.Diagnostics.ActivitySource, which maps 1:1 to
OTel spans — the OTLP exporter is a second listener bolted on in Otlp.cs, and the
model is unchanged either way.
dotnet run # 5-message walkthrough, every mechanism, readable output
dotnet run -- scale # 100k messages off a queue, drained into batches of 100
dotnet run -- scale is the shape people actually have: messages pile up on a queue, one
service drains them into batches, each batch is forwarded onward as a unit.
--messages N default 100000
--batch-size B default 100 -> 1000 batches
--order ORD-... which order to look up
--batch B000.. which batch to open
It exists to answer two questions without scanning anything:
| Question | Mechanism | Cost |
|---|---|---|
what order ids went out in B00042? |
batch.id attribute on every message root span |
index lookup |
which batch did ORD-004711 land in? |
declared business key order.id, same span |
index lookup |
Both directions come off the same two attributes. The link does not answer either of them — it points message -> batch, one hop per message, so batch -> its 100k members would mean walking 100k links. Putting those links on the batch span instead breaks at 128 (SDK limit) and ~5 MB on one span. The attribute is the answer; the link is for opening the batch from a message you are already looking at.
The batch trace is 3 spans (batch -> drain_queue, publish_batch) whether the batch holds
5 messages or 100k. Batch cost is O(batches), message cost is O(messages), never the product.
Measured, 100k messages / 1000 batches, exporting to a local OTLP endpoint:
502,400 spans 101,000 traces ~239 MB OTLP 1.8s to produce
100,000 order ids indexed, 0 in more than one batch, 200 rejected and still indexed
Spans go to http://localhost:4318 (OTLP http/protobuf) by default — a collector, or any
backend that accepts OTLP. The in-memory Collector stays alongside it; that is what answers
the index questions in the console output.
dotnet run # -> http://localhost:4318
OTEL_SDK_DISABLED=true dotnet run # console only, no export
OTEL_EXPORTER_OTLP_ENDPOINT=https://otlp.example.com \
OTEL_EXPORTER_OTLP_HEADERS="Authorization=Bearer $TOKEN" dotnet run
| Env var | Default |
|---|---|
OTEL_EXPORTER_OTLP_ENDPOINT |
http://localhost:4318 (/v1/traces appended) |
OTEL_EXPORTER_OTLP_HEADERS |
none — set it if your endpoint needs auth |
OTEL_SERVICE_NAME |
mtrace-simulation |
OTEL_SDK_DISABLED |
unset |
The same two questions in a backend — here a ClickHouse OTel schema; adjust the table name to yours:
-- what order ids went out in B00042?
SELECT SpanAttributes['order.id'] AS order_id, TraceId, StatusCode
FROM telemetry.traces
WHERE ServiceName = 'mtrace-simulation'
AND SpanAttributes['batch.id'] = 'B00042'
AND SpanName = 'message_received'
ORDER BY order_id;
-- which batch did ORD-004711 land in?
SELECT SpanAttributes['batch.id'], TraceId
FROM telemetry.traces
WHERE ServiceName = 'mtrace-simulation'
AND SpanAttributes['order.id'] = 'ORD-004711'
AND SpanName = 'message_received';
| Concern | Mechanism |
|---|---|
| Batch runtime (how long, how many) | its own short trace, 1 span |
| A-to-Z for one message | one trace per message, ~5 spans |
| message -> batch | ActivityLink, link.type=batch |
| batch -> its 100k messages | batch.id attribute + index |
| "did order 123 arrive?" | business-key index -> trace id |
| retry / replay | new trace, link.type=retry, same business key |
| next hop, same process or not | W3C traceparent on the wire |
| lossy boundary (file drop) | out-of-band context registry (preferred) |
| lossy boundary, no registry | trace id derived from business key (fallback) |
StartActivity(..., default(ActivityContext), ...)does not create a root — it falls back toActivity.Current. Null outActivity.Currentfirst, or every message silently joins the batch trace. SeeTrace.StartRoot.- Links are creation-time on .NET 8 and earlier;
Activity.AddLinkis .NET 9+. - Derived trace ids collapse retries into one trace. Per-connector setting, never global.
- Business keys must be declared so only those get indexed. Indexing every attribute is what makes people think high cardinality is impossible.
- Build the
TracerProviderbefore the first span. AnActivityListeneradded later never sees spans that have already started — the batch trace would just vanish. - Setting
OtlpExporterOptions.Endpointin code suppresses the SDK's/v1/tracesappend.Otlp.TracesPathdoes it instead, so the env var behaves as expected. - Short-lived process:
ForceFlushbefore exit, or the batch processor is disposed with the last spans still queued. - The batch processor drops the overflow silently. A drain loop fills the export queue
faster than the exporter empties it; the surplus is discarded with no error, no log and a
successful-looking flush. First 100k run here: 502,400 spans produced, 208,986 stored,
exit code 0. The defaults (queue 2048, export batch 512, 5s schedule) are sized for a
service emitting spans as requests arrive, not for a drain loop.
Otlp.TuneForVolumewidens them and the scenario flushes every 100 batches, which is the backpressure a real connector needs anyway. Always reconcile produced against stored — the failure mode is a plausible subset, not an error.
Apache-2.0. See LICENSE.