Repository navigation
fix(tracing): keep metadata within the span attribute limit - #1957
Conversation
Each metadata key is written as its own OTel attribute. Once a span holds more than SpanLimits.max_span_attributes (default 128) attributes, the OTel SDK evicts the oldest ones, which are the Langfuse core attributes such as input, model and version. Before writing observation attributes, count the span's existing attribute keys plus the new keys and drop the excess new metadata keys from the tail. Non-metadata attributes and overwrites of existing keys are never dropped. Log one warning per write that dropped keys instead of raising. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
8c564d9 to
f097372
Compare
|
@claude review |
There was a problem hiding this comment.
Claude Code Review
No review was started: this request came from a bot account. Manual reviews can only be requested by someone with write access to this repository. Ask a maintainer to comment @claude review, or have your automation post the comment from a user account with write access.
Tip: disable this comment in your organization's Code Review settings.
| new_metadata_keys = [key for key in new_keys if _is_observation_metadata_key(key)] | ||
| free_slots = max( | ||
| 0, max_attributes - len(existing) - (len(new_keys) - len(new_metadata_keys)) | ||
| ) | ||
| dropped = new_metadata_keys[free_slots:] | ||
| if not dropped: | ||
| return attributes |
There was a problem hiding this comment.
Later updates lose earlier fields
When metadata fills a span at creation, a later update(output=..., usage_details=...) still exceeds the limit. The helper only drops new metadata keys, so an update without metadata returns unchanged. OpenTelemetry then evicts earlier core or trace attributes. The OpenAI integration follows this start-then-update pattern.
Reserve space for fields added later, or reclaim metadata space before adding them. Add a test that fills the span before updating its output and usage.
Knowledge Base Used: Observations, spans, and traces
Prompt To Fix With AI
This is a comment left during a code review.
Path: langfuse/_client/span.py
Line: 95-101
Comment:
**Later updates lose earlier fields**
When metadata fills a span at creation, a later `update(output=..., usage_details=...)` still exceeds the limit. The helper only drops new metadata keys, so an update without metadata returns unchanged. OpenTelemetry then evicts earlier core or trace attributes. The OpenAI integration follows this start-then-update pattern.
Reserve space for fields added later, or reclaim metadata space before adding them. Add a test that fills the span before updating its output and usage.
**Knowledge Base Used:** [Observations, spans, and traces](https://app.greptile.com/personal-org-4986/-/custom-context/knowledge-base/langfuse/langfuse-python/-/docs/observations-spans-and-traces.md)
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.| def test_metadata_over_limit_at_start_keeps_core_attributes( | ||
| langfuse_memory_client, get_span, caplog | ||
| ): | ||
| metadata = {f"key_{i}": f"value_{i}" for i in range(130)} |
There was a problem hiding this comment.
These tests assume a limit of 128, but langfuse_memory_client creates its provider without clearing or setting OTEL_SPAN_ATTRIBUTE_COUNT_LIMIT. If the environment sets a valid custom limit such as 256, the overflow and warning assertions fail even when the code works.
Set the limit explicitly before creating the provider for these tests, as small_limit_client already does.
Prompt To Fix With AI
This is a comment left during a code review.
Path: tests/unit/test_metadata_attribute_limit.py
Line: 38-41
Comment:
**Custom limits break tests**
These tests assume a limit of 128, but `langfuse_memory_client` creates its provider without clearing or setting `OTEL_SPAN_ATTRIBUTE_COUNT_LIMIT`. If the environment sets a valid custom limit such as 256, the overflow and warning assertions fail even when the code works.
Set the limit explicitly before creating the provider for these tests, as `small_limit_client` already does.
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
There was a problem hiding this comment.
Beyond the inline findings, I also checked whether the third call site's unfiltered-looking attributes dict (span.py:745-747) could pass None values into _drop_metadata_over_attribute_limit — it can't, since create_span_attributes/create_generation_attributes already strip None before returning. Also confirmed str.removeprefix is safe given this package requires Python >=3.11.
Extended reasoning...
The diff adds a helper that counts a span's existing OTel attributes plus pending new ones and, if over max_span_attributes, drops excess new observation-metadata keys (by key prefix) rather than letting the OTel SDK evict the oldest (core) attributes; it's wired into the three set_attributes call sites in span.py. No auth/crypto/injection surface; this is local attribute bookkeeping backed by a new unit test file. The already-identified candidate issues (reliance on private _limits.max_span_attributes, non-metadata keys still able to trigger eviction, propagate_attributes and cross-thread races not covered) are real residual limitations but are explicitly acknowledged as out of scope or accepted tradeoffs by the PR description, and none of them are regressions versus current main.
|
@claude review 🤖 Written by Claude (an AI agent) on behalf of Niklas |
Capping metadata at exactly the span attribute limit left no room for attributes written in later updates, so e.g. update(output=...) made the OTel SDK evict the oldest attribute (langfuse.internal.is_app_root, then input, model, ...). Metadata now also leaves one slot for each reserved observation attribute that is not on the span or in the current write yet, matching langfuse-js. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
@claude review 🤖 Written by Claude (an AI agent) on behalf of Niklas |
The tests assert the OTel default of 128 attributes, so a custom OTEL_SPAN_ATTRIBUTE_COUNT_LIMIT or OTEL_ATTRIBUTE_COUNT_LIMIT in the environment made them fail. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
@claude review 🤖 Written by Claude (an AI agent) on behalf of Niklas |
There was a problem hiding this comment.
Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.
Beyond the inline finding, I also checked whether the per-write warning logging added in _drop_metadata_over_attribute_limit (span.py:133-144) could spam logs without throttling on repeated overflowing writes — it logs once per call that actually drops keys, which matches the PR's stated design and isn't a bug.
Extended reasoning...
The diff adds attribute-budget enforcement for observation metadata in langfuse/_client/span.py (a new helper plus a reserved-keys tuple applied at three set_attributes call sites) with accompanying unit tests; no auth/crypto/injection surface is touched, this is purely SDK-local span-attribute bookkeeping. A prior run already flagged core-attribute eviction and orphaned media uploads, and the author pushed two follow-up commits (5b626a2, f9a20c2) that added the reserved-slot logic and pinned the test limit, addressing the second point; this run's new confirmed finding (reserved slots wasted for non-generation observation types) is a real but narrow design gap, not large enough to warrant restating here since it is posted inline.
| _RESERVED_OBSERVATION_ATTRIBUTE_KEYS = ( | ||
| LangfuseOtelSpanAttributes.OBSERVATION_TYPE, | ||
| LangfuseOtelSpanAttributes.OBSERVATION_LEVEL, | ||
| LangfuseOtelSpanAttributes.OBSERVATION_STATUS_MESSAGE, | ||
| LangfuseOtelSpanAttributes.VERSION, | ||
| LangfuseOtelSpanAttributes.ENVIRONMENT, | ||
| LangfuseOtelSpanAttributes.OBSERVATION_INPUT, | ||
| LangfuseOtelSpanAttributes.OBSERVATION_OUTPUT, | ||
| LangfuseOtelSpanAttributes.OBSERVATION_MODEL, | ||
| LangfuseOtelSpanAttributes.OBSERVATION_USAGE_DETAILS, | ||
| LangfuseOtelSpanAttributes.OBSERVATION_COST_DETAILS, | ||
| LangfuseOtelSpanAttributes.OBSERVATION_COMPLETION_START_TIME, | ||
| LangfuseOtelSpanAttributes.OBSERVATION_MODEL_PARAMETERS, | ||
| LangfuseOtelSpanAttributes.OBSERVATION_PROMPT_NAME, | ||
| LangfuseOtelSpanAttributes.OBSERVATION_PROMPT_VERSION, |
There was a problem hiding this comment.
🟡 (optional) Every span/event observation (non-generation) permanently loses metadata capacity for 7 reserved keys it will never write. _RESERVED_OBSERVATION_ATTRIBUTE_KEYS (span.py:72-86) always includes OBSERVATION_MODEL, MODEL_PARAMETERS, USAGE_DETAILS, COST_DETAILS, COMPLETION_START_TIME, PROMPT_NAME, PROMPT_VERSION, but update() for span/event types only ever calls create_span_attributes (span.py:758), which never emits those keys. _drop_metadata_over_attribute_limit is a module function with no access to self._observation_type, so it reserves all 14 slots unconditionally, dropping up to 7 more metadata keys than needed on every plain span or event. Fix: make the reserved-key set type-aware so span/event observations keep their full metadata budget.
Why this was flagged
Trigger: any non-generation observation (default as_type="span", or "event") with metadata large enough to approach the span attribute limit, e.g. the small_limit_client tests with limit=40. _drop_metadata_over_attribute_limit (span.py:116-121) counts all 14 _RESERVED_OBSERVATION_ATTRIBUTE_KEYS as reserved even though create_span_attributes (span.py:758) never writes OBSERVATION_MODEL/MODEL_PARAMETERS/USAGE_DETAILS/COST_DETAILS/COMPLETION_START_TIME/PROMPT_NAME/PROMPT_VERSION for this observation type. This wastes 7 of the 14 reserved slots that could otherwise hold metadata, so more metadata keys are dropped than the stated design (reserve only for fields that may still be written) intends. On base there was no reservation at all (and no dropping), so this is a new, avoidable over-drop introduced by the fix, beyond what the PR's own documented "up to 14 fewer metadata keys" cost already accounts for since it conflates generation-only keys with span-only observations.
Verification: nit. reserved_count (langfuse/_client/span.py:116-120) counts all 14 entries of _RESERVED_OBSERVATION_ATTRIBUTE_KEYS (span.py:72-86) unconditionally. For span/event, update() takes the else branch at span.py:756-773 calling create_span_attributes, which never writes the 7 generation-only reserved keys. These still count toward reserved_count, shrinking free_slots (lines 126-128), so up to 7 metadata keys are dropped that would have fit.
…ibute overflow The Python OTel SDK evicts the oldest attribute once a span is full, while OTel JS drops the newest. After capping metadata, the limit helper now also drops excess new non-metadata keys from the tail when the span has no physical capacity left, so Langfuse writes never evict existing attributes. Overwrites of existing keys are always kept, and at most one warning is logged per write. Langfuse write paths on observation spans now go through the guard: the observation constructor, update(), media-processed updates, set_trace_as_public, the experiment item attribute writes, as_root marking, propagated attributes in the span processor and propagate_attributes, and the app-root marker. The warning now names SpanLimits.max_span_attributes alongside OTEL_SPAN_ATTRIBUTE_COUNT_LIMIT, matching langfuse-js. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…n limit The task span metadata put experiment_name, experiment_run_name and the dataset ids after item and experiment metadata, so the span attribute limit dropped them first. Insert the run keys first so they survive truncation, and again last so they still win over user keys of the same name. Mirrors langfuse-js. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Span-like observations and events never write model, usage, cost, completion start time, model parameters or prompt attributes, so reserving slots for them wasted 7 metadata keys. Reserve the 7 shared keys (type, level, status message, version, environment, input, output) for span-like types and events, and all 14 for generation-like types and for writes whose observation type is unknown (as_root, propagated attributes, app-root marker). Mirrors langfuse-js. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Experiment metadata and experiment item metadata were propagated as one attribute per flattened key (langfuse.experiment.item.metadata.<path>), so large item metadata filled the task span and its children and left no room for later attributes such as the output. Send each as a single serialized JSON attribute (langfuse.experiment.metadata and langfuse.experiment.item.metadata), like langfuse-js. The server parses the blob at the bare prefix and flattens it to the same paths. Experiment attributes are set by the SDK and skip the 200-character validation of user-propagated values, matching langfuse-js. Remove the now unused _flatten_and_serialize_metadata_values. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…butes The experiment task span got its observation metadata at start, before the experiment attributes were copied onto it. The metadata cap could then fill the span up to the reserved slots, and the copied attributes used those slots, so the later output was dropped. Start the task span without metadata and set it inside the propagation block, before the task runs, so the cap counts the copied attributes and trims the observation metadata instead. The experiment run keys still go first. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Problem
Observation metadata is written as one OTel attribute per key (
langfuse.observation.metadata.<key>). The OTel SDK caps attributes per span (SpanLimits.max_span_attributes, default 128, configurable viaOTEL_SPAN_ATTRIBUTE_COUNT_LIMIT), andBoundedAttributesevicts the oldest attribute on overflow. Langfuse core attributes are written first, so large metadata evicts them.Repro on
main: a generation started withinput,modeland 130 metadata keys ended with 128 attributes anddropped_attributes=9.input,model, version and similar attributes were gone; onlylangfuse.observation.typeandoutputsurvived. Propagated trace attributes (langfuse.trace.metadata.*, user id, session id, tags) use up the same budget.Behavior changes
What users will notice compared to
main:experiment_name,experiment_run_name,dataset_id,dataset_item_id). Run keys still win over user keys with the same name.langfuse.experiment.metadata,langfuse.experiment.item.metadata) instead of one attribute per flattened key, as in langfuse-js. The server stores the same paths. Custom span processors, exporters or masking functions that read the per-key attributes will see the new format.Fix
_drop_attributes_over_span_limitinlangfuse/_client/span.py(via_set_span_attributes_within_limit) runs before every Langfuse attribute write on a span: the observation constructor (span start, which coversstart_observation,start_as_current_observation,@observe, LangChain and OpenAI),update()(which also coversupdate_current_span/update_current_generation),_set_processed_span_attributes,set_trace_as_public, the experiment item attribute writes and parent backfill inrun_experiment,as_rootmarking, propagated attributes (span processoron_startandpropagate_attributeson the current span), and the app-root marker.It counts the span's existing attribute keys (
ReadableSpan.attributes) plus the new non-Nonekeys that aren't already on the span. If that exceedsspan._limits.max_span_attributes, it works in two steps, both keeping insertion order:This matches OTel JS, which drops the newest attribute once a span is full, whereas Python's
BoundedAttributesevicts the oldest. Langfuse writes no longer evict an existing attribute. Overwriting a key that's already on the span costs nothing and is always kept. Spans without SDK limits (non-recording or non-SDK spans) are left alone. It keeps no side state: the OTel span already holds the key set.Experiment run keys
run_experimentused to putexperiment_name,experiment_run_name,dataset_idanddataset_item_idafter item and experiment metadata in the task span metadata, so they were the first keys to be dropped. They now go first in insertion order (so they survive truncation) and again last (so they still win over user keys of the same name), like langfuse-js.Experiment metadata wire format
Experiment metadata and experiment item metadata used to be propagated as one attribute per flattened key (
langfuse.experiment.item.metadata.<path>), so an item with many metadata keys filled the task span and every child span. They are now sent as one serialized JSON attribute each (langfuse.experiment.metadata,langfuse.experiment.item.metadata), like langfuse-js (serializeValue(...)). The server (extractMetadataFromPrefix) parses both the blob at the bare prefix and per-key attributes and flattens them to the same stored paths. The blob is also the format the server parsed first: experiment ingestion (langfuse/langfuse#10004, Nov 2025) only didJSON.parseon the bare attribute, and per-key parsing was added later (langfuse/langfuse#13368, Apr 2026). So the blob works on older self-hosted servers too. Like JS, these SDK-set experiment attributes skip the 200-character validation for user-propagated values._flatten_and_serialize_metadata_valuesis now unused and removed.Experiment task span write order
The task span used to get its observation metadata at start, before the experiment attributes were copied onto it, so the copied attributes could take the slot reserved for the output. It now starts without metadata and gets
update(metadata=...)inside the propagation block, before the task runs. The metadata cap then counts the copied attributes and trims the duplicate observation metadata, so the output still lands. The item-run span already writes its metadata in the finalupdate()after the backfill, so it needed no change.Behavior notes
langfuse_logger.warning, in the same format as langfuse-js:Dropped <n> metadata key(s) from observation '<name>' to stay within the span attribute limit of <limit> (SpanLimits.max_span_attributes / OTEL_SPAN_ATTRIBUTE_COUNT_LIMIT). Dropped keys include: <up to 5 keys>. When the hard guard also drops non-metadata keys, it readsDropped <n> metadata key(s) and <m> other attribute(s) from observation ....langfuse.trace.*), environment, release, and so on, because they all live on the same span.langfuse.trace.metadata.*,langfuse.experiment.metadata,langfuse.experiment.item.metadata) is never trimmed by the metadata cap, because the server reads it from the trace root / experiment item root span.Reserved slots
Capping metadata so the span holds exactly
max_span_attributesisn't enough. The next new non-metadata attribute would overflow the span: without the hard guard,BoundedAttributessilently evicts the oldest one, and with it the new attribute (e.g. the output) would be dropped. Repro (limit 128): a generation started withinput,modeland 130 metadata keys ends up with 124 metadata keys and 128 attributes. A followingupdate(output=...)then evictslangfuse.internal.is_app_root(dropped_attributes=1), and more updates (usage, cost, ...) would evictinput,modeland so on.The metadata budget therefore also holds back one slot for each observation attribute that a later update may still write and that isn't on the span or in the current write yet. The set depends on the observation type, like langfuse-js #1004:
_RESERVED_SPAN_ATTRIBUTE_KEYS)._RESERVED_OBSERVATION_ATTRIBUTE_KEYS).Reserved keys already on the span cost nothing.
Cost: when metadata is very large, up to 7 (spans) or 14 (generations) fewer metadata keys are kept than strictly fit. With the fix, the repro above ends with 113 metadata keys, 118 attributes and
dropped_attributes=0, with input, model, output andis_app_rootall present. A plain span started with only input and 130 metadata keys keeps 7 more metadata keys than a generation started the same way.Context
This is one of three small PRs that replace the abandoned single-JSON-blob approach in #1944 (LFE-17153).
Verification
tests/unit/test_metadata_attribute_limit.py: 16 tests, all pass, including the reserved set by observation type (a span keeps 7 more metadata keys than a generation; a generation still keeps room for model, usage, cost and output). New in this round: the hard guard (span full, then new non-metadata attributes:dropped_attributes == 0, earlier attributes kept, newest dropped, one warning), overwrites kept and the tail dropped at the helper level, the helper never raises, the exact warning format, and the experiment run keys surviving truncation on the task span (the user'sexperiment_nameis overridden). The experiment test fails without the reordering.test_experiment_with_large_metadata_keeps_output_and_experiment_attributes: an experiment item with ~150 metadata keys including a nested dict. On the task span: output present,langfuse.experiment.item.metadatais one JSON string that round-trips to the item metadata, experiment metadata / id / name / item id / root observation id present, run keys in the observation metadata,dropped_attributes == 0. A child span created in the task has both blobs and isn't flooded. Fails without the write-order change.tests/unit/test_propagate_attributes.pyandtests/unit/test_experiment.pynow assert the single JSON attributes (168 passed).uv run --frozen ruff check .: passesuv run --frozen mypy langfuse --no-error-summary: passesuv run --frozen ruff format --check .: passes for all changed files. It flagstests/unit/test_media.py, which is already unformatted on the base branch and was left alone to keep this PR scoped.uv run --frozen pytest -n auto --dist worksteal tests/unitwith dummyLANGFUSE_PUBLIC_KEY/LANGFUSE_SECRET_KEY/LANGFUSE_BASE_URL: 798 passed, 2 skipped. Without those variables, the 18 tests intests/unit/test_prompt.pyerror with "Langfuse client is not initialized", on the base branch too.🤖 Written by Claude (an AI agent) on behalf of Niklas
🤖 Generated with Claude Code
The PR needs a fix for later output and usage writes on spans already filled by metadata before merging.
Summary
Adds a shared filter that drops excess new observation metadata before writing span attributes. Adds seven tests for limits, updates, overwrites, warnings, and propagated attributes.
Diagram
%%{init: {'theme': 'neutral'}}%% flowchart TD A[Create or update observation] --> B[Process media and mask values] B --> C[Count existing and new attributes] C --> D[Drop excess new metadata] D --> E[Write span attributes] E --> F{Does the write still exceed the limit?} F -->|No| G[Keep earlier fields] F -->|Yes: later non-metadata fields| H[OpenTelemetry evicts earlier fields]Reviews (1) · Last reviewed commit: "fix(tracing): keep metadata within the s..." · Reviewed by Greptile