Skip to content

feat(batch-evaluation)!: evaluate observations via the v2 observations API - #1936

Merged
hassiebp merged 59 commits into
prepare-v5-releasefrom
lfe-17058-python-sdk-v5-rework-batch-evaluation-for-langfuse-platform
Oct 7, 2026
Merged

hassiebp merged 59 commits into
prepare-v5-releasefrom
lfe-17058-python-sdk-v5-rework-batch-evaluation-for-langfuse-platform

Conversation

@hassiebp

@hassiebp hassiebp commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator

What does this PR do?

This PR now carries three changes, because #1935 and #1937 were squash-merged into this branch.

Batch evaluation on the v2 observations API (this PR's original change)

  • run_batched_evaluation reads only through api.observations.get_many (v2), with cursor pagination. The v3 GET /traces and v1 GET /observations paths are removed (they return 404 on Langfuse v4). It builds on community PR fix(batch_evaluation): route batch_evaluation through v2 observations API #1898.
  • It evaluates observations only. The scope parameter and the root-observation / trace mode are removed. Every observation matching filter is scored on the observation.
  • fields (v2 field groups, default "core,basic,io,metadata") replaces fetch_trace_fields. Mappers receive ObservationV2, where input/output are raw strings.
  • BatchEvaluationResumeToken carries the v2 cursor. Resuming reuses the token's filter and rejects a different one. Duplicate rows that the events table briefly returns for one observation are evaluated once.
  • _additional_trace_tags is dropped, since it relied on trace-create ingestion. Scores still go through the existing create_score path.

Remove trace-level IO setters (from #1935)

  • Langfuse.set_current_trace_io(), span.set_trace_io() and the langfuse.trace.input/output attributes and media handling are removed. Set IO on the root observation instead.

Test suites on v4, CI on events_only (from #1937)

  • New v4 test helpers read through v2 observations and scores_v3. All e2e and live-provider suites are ported, and CI drops the dual-write override.

Base: #1934's branch. Merge after #1934, then #1929 can merge at any point.

Type of change

  • Bug fix
  • New feature
  • Breaking change
  • Refactor
  • Documentation update
  • Tooling, CI, or repo maintenance

Verification

uv run --frozen ruff check .                       # All checks passed!
uv run --frozen mypy langfuse --no-error-summary   # clean
uv run --frozen pytest -n auto tests/unit          # 751 passed, 2 skipped
# against a local Langfuse server (events_only):
uv run --frozen pytest -n 4 tests/e2e/test_batch_evaluation.py  # 7 passed

Checklist

  • I self-reviewed the diff using code_review.md.
  • I added or updated tests for behavior changes.
  • I updated docs, examples, or .env.template if needed.
  • I did not hand-edit generated files; if generated files changed, I used the upstream regeneration path.
  • I did not commit secrets or credentials.
Open in Web Open in Cursor 

RetriggerConfidence Score: 2/5

The PR is not yet safe to merge because resumed runs can miss observations and root-scope runs can score a trace more than once.

Summary

The PR moves batch evaluation reads to the v2 observations API, adds root-observation scope and field selection, and replaces page-based fetching with cursor pagination.

  • Resume handling needs to preserve the original filter and handle tied timestamps when a cursor is unavailable.
  • Root-observation filtering does not by itself guarantee a single trace-level evaluation.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[Scope and filter] --> B[v2 observations API]
  B --> C[Cursor-paginated pages]
  C --> D[Mapper and evaluators]
  D --> E{Scope}
  E -->|observations| F[Observation scores]
  E -->|root_observations| G[Trace scores]
  C --> H[Resume token]
Loading

Reviews (1) · Last reviewed commit: "test(batch-evaluation): rewrite e2e suit..."

cursoragent and others added 11 commits October 6, 2026 14:15
Python 3.10 reaches end of life in October 2026. Raise requires-python
to >=3.11, target py311 in ruff, drop 3.10 from the CI unit matrix, and
remove the pre-3.11 fallbacks for asyncio.create_task(context=...) and
typing.NotRequired.

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
The OpenAI integration now always uses the v1 client shapes. Remove the
pre-1.0 method definitions and every _is_openai_v1() branch, and raise
the dev dependency floor to openai>=1.0.0.

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
The default OTLP exporter now sends x-langfuse-ingestion-version: 4 so
servers in dual-write or preview mode process exported spans directly.
It is a no-op on events_only deployments, and additional_headers can
still override it.

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
blocked_instrumentation_scopes was deprecated in favor of
should_export_span. Remove it from Langfuse, the resource manager and
the span processor, and drop the leftover is_active entry from the
create_prompt docstring. LANGFUSE_HOST and host keep working, but now
log a deprecation warning when they decide the base URL.

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
propagate_attributes(metadata=...) coerced non-string values with str(),
so a boolean became "True" and a list became its Python repr. Serialize
them as JSON instead, like observation metadata, and drop None values
instead of sending "None".

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
…taset_run

These helpers call the dataset-run endpoints that Langfuse v4 no longer
serves, so they always fail with a 404. Read experiment runs through
langfuse.api.experiments.list() and list_items() instead.

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
Generate the experiment id client-side instead of creating dataset run
items via POST /dataset-run-items. Dataset-backed runs derive a stable id
from project, dataset and run name (matching the server's derivation), so
re-runs with the same run name keep grouping into one experiment. Local
runs keep a random id.

- set langfuse.experiment.item.version from the pinned dataset version
- always persist run evaluations as scores on the experiment, incl. local data
- return experiment_id / experiment_url (v4 results page) on results;
  dataset_run_id / dataset_run_url remain as deprecated aliases

BREAKING CHANGE: run_experiment no longer creates Postgres dataset runs;
dataset_run_id / dataset_run_url are deprecated aliases that are now also
set for local-data experiments, and the URL points to the experiments page.

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
Replace dataset run, dataset run item and legacy trace reads with the
experiments, observations and v3 scores APIs, and cover stable experiment
ids, run-level scores for local data and the recorded item version.

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
…s API

Batch evaluation now reads exclusively from GET /api/public/v2/observations
with cursor pagination, so it works on Langfuse platform v4.

BREAKING CHANGES:
- scope="traces" is replaced by scope="root_observations", which evaluates
  the logical root observation of each trace and attaches scores to the
  trace. scope="observations" is unchanged in intent.
- Mappers receive an ObservationV2. input/output are raw strings and only
  the requested field groups are populated.
- fetch_trace_fields is replaced by fields (v2 field groups), defaulting to
  "core,basic,io,metadata" in both the client and the runner.
- filter must be a JSON array in the v2 observations filter format.
- Resume tokens carry the pagination cursor and are also returned when
  max_items is reached while more items exist.
- The private _additional_trace_tags option is removed.

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown

@claude review

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

No review was started: this request came from a bot account. Manual reviews can only be requested by someone with write access to this repository. Ask a maintainer to comment @claude review, or have your automation post the comment from a user account with write access.

Tip: disable this comment in your organization's Code Review settings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 638034e310

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread langfuse/batch_evaluation.py Outdated
Comment thread langfuse/batch_evaluation.py Outdated
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-07T20:17:31.870335Z 4ea561d New commits
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

Comment thread langfuse/batch_evaluation.py Outdated
Comment thread langfuse/batch_evaluation.py
Comment thread langfuse/batch_evaluation.py Outdated

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed this PR and didn't find any bugs. Because it's a large, breaking rework of the batch-evaluation read path (new v2 observations API, cursor-based pagination, and a redesigned resume token), a human look would still be worthwhile.

What was reviewed:

  • Cursor-pagination/resume loop and max_items clamping in batch_evaluation.py
  • Filter construction (v2 JSON-array validation, scope/resume conditions) and scope-based score attachment (observations vs root_observations)
  • Ruled out: observations with no trace_id now fail per-item rather than being silently scored; cursor-less resume fallback uses strict "<" on startTime so same-timestamp boundary items could be skipped; max_items<=0 reports has_more_items=True despite fetching nothing
Extended reasoning...

The change reworks langfuse/batch_evaluation.py and langfuse/_client/client.py to move batch evaluation off the legacy v1 traces/observations APIs onto the v2 observations API, replacing timestamp-based pagination with cursor-based pagination and a redesigned resume token; tests were updated accordingly (unit suite grew substantially, e2e suite was heavily trimmed). It doesn't touch auth/crypto, but it changes the request/filter contract sent to the API and the pagination/resume correctness guarantees of a data-processing feature. It's a large (1000+ line), self-declared breaking change with real design tradeoffs (cursor vs. timestamp resume, root-observation semantics replacing traces), which is why a human look is still warranted even though no bugs were found.

cursoragent and others added 12 commits October 6, 2026 15:02
Use compact separators, keep non-ASCII characters, and send None as
"null", so the Python and JS SDKs emit byte-identical propagated
metadata values.

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
…ataset_run-are' into lfe-17047-python-sdk-v5-update-the-experiment-runner

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
…nto lfe-17058-python-sdk-v5-rework-batch-evaluation-for-langfuse-platform

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
…ata-align-with' into lfe-16769-bug-get_dataset_run-get_dataset_runs-delete_dataset_run-are

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
inspect.iscoroutinefunction ignores asyncio's _is_coroutine marker,
which asgiref's markcoroutinefunction sets on Python < 3.12. Accept
both so such functions keep the async wrapper and their observation
ends after the coroutine runs.

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
…o lfe-17074-python-sdk-v5-send-x-langfuse-ingestion-version-4-by-default

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
…scopes-and' into lfe-17061-python-sdk-v5-json-serialize-propagated-metadata-align-with

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
…ataset_run-are' into lfe-17047-python-sdk-v5-update-the-experiment-runner

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
…nto lfe-17058-python-sdk-v5-rework-batch-evaluation-for-langfuse-platform

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
…on-4-by-default' into lfe-17078-python-sdk-v5-remove-blocked_instrumentation_scopes-and

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
…ata-align-with' into lfe-16769-bug-get_dataset_run-get_dataset_runs-delete_dataset_run-are

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
…0' into lfe-17080-python-sdk-v5-drop-openai-sdk-10-support

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
cursoragent and others added 16 commits October 7, 2026 08:57
…ata-align-with' into lfe-16769-bug-get_dataset_run-get_dataset_runs-delete_dataset_run-are

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
…ataset_run-are' into lfe-17047-python-sdk-v5-update-the-experiment-runner

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
…nto lfe-17058-python-sdk-v5-rework-batch-evaluation-for-langfuse-platform

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
Integers keep their exact digits as JSON numbers at any depth, and
propagated metadata values containing NaN or Infinity are dropped with a
warning instead of being stored, matching the JS SDK.

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
…ata-align-with' into lfe-16769-bug-get_dataset_run-get_dataset_runs-delete_dataset_run-are

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
…ataset_run-are' into lfe-17047-python-sdk-v5-update-the-experiment-runner

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
…nto lfe-17058-python-sdk-v5-rework-batch-evaluation-for-langfuse-platform

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
…-blocked_instrumentation_scopes-and

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
…scopes-and' into lfe-17061-python-sdk-v5-json-serialize-propagated-metadata-align-with

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
…ata-align-with' into lfe-16769-bug-get_dataset_run-get_dataset_runs-delete_dataset_run-are

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
…ataset_run-are' into lfe-17047-python-sdk-v5-update-the-experiment-runner

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
…nto lfe-17058-python-sdk-v5-rework-batch-evaluation-for-langfuse-platform

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
…-the-experiment-runner

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
Comment thread langfuse/_client/client.py Outdated
scope: Which observations to evaluate. Must be one of:
- "observations": Every observation matching the filter (spans,
generations, events, ...). Scores are attached to the observation.
- "root_observations": Only the logical root observation of each trace.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

remove this option, let's keep it simple to only observations for now

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in 8299f52: batch evaluation now runs on observations only.

  • The scope parameter is removed from run_batched_evaluation and the runner. Removing it entirely rather than keeping a one-value option means a scope can come back later as an additive change.
  • Root-observation mode is gone: the isRootObservation filter, the one-root-per-trace dedupe and trace-level score attachment. Every matching observation is scored on the observation (plus the trace via the private _add_observation_scores_to_trace, as before).
  • BatchEvaluationResumeToken no longer has a scope field. The filter check on resume stays.
  • Docstrings and examples are updated. The root-scope unit and e2e tests are removed, and the remaining e2e tests now cover all observations. Unit: 750 passed; batch-eval e2e against an events_only server: 7 passed.

Remove the deprecated Langfuse.set_current_trace_io() and
span.set_trace_io(), the langfuse.trace.input/output attributes, and
their media handling. On Langfuse v4 a trace's input and output come
from its root observation: set them with update() on the root span or
update_current_span().

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 8063f8467f

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread langfuse/batch_evaluation.py Outdated
Comment thread langfuse/batch_evaluation.py Outdated
Comment thread langfuse/batch_evaluation.py Outdated
* feat(tracing)!: remove trace-level input/output setters

Remove the deprecated Langfuse.set_current_trace_io() and
span.set_trace_io(), the langfuse.trace.input/output attributes, and
their media handling. On Langfuse v4 a trace's input and output come
from its root observation: set them with update() on the root span or
update_current_span().

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>

* test: read e2e assertions through the v4 observations API

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>

* ci: run e2e against an events_only server

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>

* test: pin that create_score logs invalid input without enqueueing

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>

* test: remove the v3 wait_for_trace helper

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>

* test: keep scalar text IO as strings and read full response_format

The v2 observations API returns IO as raw strings, so only JSON objects
and arrays are parsed back; text like "2" stays a string. The
response_format assertion requests the untruncated metadata value.

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>

* test: dedupe observation rows and assert streamed text as a string

The events table can briefly return several rows for one span, so the
observation helper keeps the latest row per id. A streamed completion of
"2" is text and is now asserted as a string.

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
@hassiebp
hassiebp requested a review from a team as a code owner October 7, 2026 20:12

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 4ea561df36

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread langfuse/batch_evaluation.py Outdated
Comment thread langfuse/batch_evaluation.py
Comment thread langfuse/batch_evaluation.py Outdated
cursoragent and others added 3 commits October 7, 2026 20:17
Remove the scope parameter and root-observation mode, so batch
evaluation always runs on the observations matching the filter and
scores each observation. Resume tokens no longer carry a scope.

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
…nto lfe-17058-python-sdk-v5-rework-batch-evaluation-for-langfuse-platform

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
The events table can briefly return several rows for one observation,
within a page or across a page boundary. Keep the most recently updated
row per id so each observation is evaluated and scored once. Also build
exception messages before raising.

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
Base automatically changed from lfe-17047-python-sdk-v5-update-the-experiment-runner to prepare-v5-release October 7, 2026 20:25
cursoragent and others added 2 commits October 7, 2026 20:25
…-batch-evaluation-for-langfuse-platform

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
…-batch-evaluation-for-langfuse-platform

Co-authored-by: Hassieb Pakzad <hassiebp@users.noreply.github.com>
@hassiebp
hassiebp merged commit b4c284b into prepare-v5-release Oct 7, 2026
12 checks passed
@hassiebp
hassiebp deleted the lfe-17058-python-sdk-v5-rework-batch-evaluation-for-langfuse-platform branch October 7, 2026 20:30
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants