Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

Repository files navigation

Message tracing — trace/span ID structure for high-volume integration

Runnable simulation of one-trace-per-message with batch fan-in via links. The pipeline uses nothing but System.Diagnostics.ActivitySource, which maps 1:1 to OTel spans — the OTLP exporter is a second listener bolted on in Otlp.cs, and the model is unchanged either way.

dotnet run              # 5-message walkthrough, every mechanism, readable output
dotnet run -- scale     # 100k messages off a queue, drained into batches of 100

The volume scenario

dotnet run -- scale is the shape people actually have: messages pile up on a queue, one service drains them into batches, each batch is forwarded onward as a unit.

--messages N       default 100000
--batch-size B     default 100          -> 1000 batches
--order ORD-...    which order to look up
--batch B000..     which batch to open

It exists to answer two questions without scanning anything:

Question Mechanism Cost
what order ids went out in B00042? batch.id attribute on every message root span index lookup
which batch did ORD-004711 land in? declared business key order.id, same span index lookup

Both directions come off the same two attributes. The link does not answer either of them — it points message -> batch, one hop per message, so batch -> its 100k members would mean walking 100k links. Putting those links on the batch span instead breaks at 128 (SDK limit) and ~5 MB on one span. The attribute is the answer; the link is for opening the batch from a message you are already looking at.

The batch trace is 3 spans (batch -> drain_queue, publish_batch) whether the batch holds 5 messages or 100k. Batch cost is O(batches), message cost is O(messages), never the product.

Measured, 100k messages / 1000 batches, exporting to a local OTLP endpoint:

502,400 spans   101,000 traces   ~239 MB OTLP   1.8s to produce
100,000 order ids indexed, 0 in more than one batch, 200 rejected and still indexed

Exporting over OTLP

Spans go to http://localhost:4318 (OTLP http/protobuf) by default — a collector, or any backend that accepts OTLP. The in-memory Collector stays alongside it; that is what answers the index questions in the console output.

dotnet run                                              # -> http://localhost:4318
OTEL_SDK_DISABLED=true dotnet run                       # console only, no export
OTEL_EXPORTER_OTLP_ENDPOINT=https://otlp.example.com \
OTEL_EXPORTER_OTLP_HEADERS="Authorization=Bearer $TOKEN" dotnet run
Env var Default
OTEL_EXPORTER_OTLP_ENDPOINT http://localhost:4318 (/v1/traces appended)
OTEL_EXPORTER_OTLP_HEADERS none — set it if your endpoint needs auth
OTEL_SERVICE_NAME mtrace-simulation
OTEL_SDK_DISABLED unset

The same two questions in a backend — here a ClickHouse OTel schema; adjust the table name to yours:

-- what order ids went out in B00042?
SELECT SpanAttributes['order.id'] AS order_id, TraceId, StatusCode
FROM telemetry.traces
WHERE ServiceName = 'mtrace-simulation'
  AND SpanAttributes['batch.id'] = 'B00042'
  AND SpanName = 'message_received'
ORDER BY order_id;

-- which batch did ORD-004711 land in?
SELECT SpanAttributes['batch.id'], TraceId
FROM telemetry.traces
WHERE ServiceName = 'mtrace-simulation'
  AND SpanAttributes['order.id'] = 'ORD-004711'
  AND SpanName = 'message_received';

The model

Concern Mechanism
Batch runtime (how long, how many) its own short trace, 1 span
A-to-Z for one message one trace per message, ~5 spans
message -> batch ActivityLink, link.type=batch
batch -> its 100k messages batch.id attribute + index
"did order 123 arrive?" business-key index -> trace id
retry / replay new trace, link.type=retry, same business key
next hop, same process or not W3C traceparent on the wire
lossy boundary (file drop) out-of-band context registry (preferred)
lossy boundary, no registry trace id derived from business key (fallback)

Things that will bite you

  • StartActivity(..., default(ActivityContext), ...) does not create a root — it falls back to Activity.Current. Null out Activity.Current first, or every message silently joins the batch trace. See Trace.StartRoot.
  • Links are creation-time on .NET 8 and earlier; Activity.AddLink is .NET 9+.
  • Derived trace ids collapse retries into one trace. Per-connector setting, never global.
  • Business keys must be declared so only those get indexed. Indexing every attribute is what makes people think high cardinality is impossible.
  • Build the TracerProvider before the first span. An ActivityListener added later never sees spans that have already started — the batch trace would just vanish.
  • Setting OtlpExporterOptions.Endpoint in code suppresses the SDK's /v1/traces append. Otlp.TracesPath does it instead, so the env var behaves as expected.
  • Short-lived process: ForceFlush before exit, or the batch processor is disposed with the last spans still queued.
  • The batch processor drops the overflow silently. A drain loop fills the export queue faster than the exporter empties it; the surplus is discarded with no error, no log and a successful-looking flush. First 100k run here: 502,400 spans produced, 208,986 stored, exit code 0. The defaults (queue 2048, export batch 512, 5s schedule) are sized for a service emitting spans as requests arrive, not for a drain loop. Otlp.TuneForVolume widens them and the scenario flushes every 100 batches, which is the backpressure a real connector needs anyway. Always reconcile produced against stored — the failure mode is a plausible subset, not an error.

License

Apache-2.0. See LICENSE.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages