Skip to content

Latest commit

 

History

History
372 lines (309 loc) · 20.2 KB

File metadata and controls

372 lines (309 loc) · 20.2 KB

Agent evals

Every other test in this repo checks code. These check the assistant — the part students actually use, and the part that can get quietly worse while the whole test suite stays green. A prompt edit, a model bump, a capability rename or a change to how manuals are attached can all degrade answers without breaking a single unit test.

The suite checks the three behaviours that carry the assistant's value:

  1. it finds the right machine from a natural request,
  2. it grounds answers in the catalogue and says where they came from,
  3. it calls the right tool instead of inventing an answer,

plus the one that matters most when it regresses: it admits what the lab does not have.

Design spec: docs/specs/2026-07-29-agent-eval-harness-design.md.


Running it

Every model call goes through the Vercel AI Gateway (gateway spec §3.1) — there is no direct-provider path any more, and ANTHROPIC_API_KEY is read by nothing.

export AI_GATEWAY_API_KEY=...   # or run where VERCEL_OIDC_TOKEN is already set
npm run eval

EVAL_MODEL=openai/gpt-6-sol npm run eval          # try a different model
EVAL_MODEL=anthropic/claude-sonnet-5 npm run eval # or a different provider entirely

The suite runs the deployment's own model — job chat in src/lib/ai/models.ts (openai/gpt-6-luna by default, MODEL_CHAT overrides it) — so a run says something about production. EVAL_MODEL overrides the eval's model independently of MODEL_CHAT, with an explicit Gateway id (provider/model, lower case), which is exactly what this harness is for when the question is "does the next model still behave?".

Important

This makes real, paid model calls — roughly one per case, plus a retry for any case that fails. It is deliberately not part of npm test or npm run test:all, which stay free, offline and green, and it must never be wired into a pull-request trigger where a fork could spend your money.

When to run it: before merging a change to a prompt fragment, a capability, or a job's model id — and before a demo. That is the whole policy.

The §10 eval gate — before switching MODEL_CHAT

Before pointing MODEL_CHAT (or the deployment's default) at a new model, the gateway spec's §10 gate is: the model must pass every honest-absence and manual-grounding case, and all but one of the rest — run twice.

EVAL_MODEL=openai/gpt-6-luna npm run eval
EVAL_MODEL=openai/gpt-6-luna npm run eval   # again — a FLAKY case on run 1 must not recur on run 2

The 2026-09-23 results are recorded, run by run, in the gateway spec: Luna first missed the gate (amendment "The chat eval gate") and chat ran on anthropic/claude-sonnet-5; after the chat prompt was tuned, Luna passed it on three runs and chat moved to openai/gpt-6-luna (amendment "Chat prompt tuning for Luna"). The anonymous-report behaviour is covered by two cases: with nobody signed in the assistant asks for a name first (issue-report-asks-for-name), and a reporter who declines still gets the ticket filed (issue-report-calls-tool): an anonymous report beats no report. Luna passed all 14 cases on 2026-09-23.

Read every cases/honest-absence.yaml and cases/manual-grounding.yaml case as a hard requirement (no FAIL, no FLAKY, on either run) and every other case file with one miss tolerated per run. A structural failure — the same case failing both runs, or failing for a different reason each time — is worth investigating before the model goes live regardless of the count; see "FLAKY is not a pass" below. This is a policy for a person driving the gate by hand, not a script: nothing in evals/ enforces the "all but one" count for you, because report.totals.failed is checked strictly (toBe(0)) — a run with one tolerated miss still exits non-zero, on purpose, so it is never silently green in a script.

Output is a per-case line, a detail block for anything that did not pass, and a JSON artifact at evals/.last-run.json (gitignored).

  PASS  laser-by-capability
  FLAKY unit-status-calls-tool
  FAIL  form4-build-volume-unknown

FAIL  form4-build-volume-unknown
  prompt: What's the exact build volume and layer resolution on this printer?
  x no_fabricated_specs — expected numbers attributed to build_volume to match the fixture
      "145" is attributed to build_volume, but the fixture records no value for it
      answer: "The Form 4 has a build volume of 145 x 145 x 185 mm."

11/12 passed · 1 flaky · 1 failed · 0 errored

FLAKY is not a pass. A failing case is retried once; one that passes only on the retry is reported as flaky and never counted as a pass. Flakiness is signal — a case that flakes is usually badly written. It does not fail the run (one flake in a dozen cases is noise), but a structural failure is always worth investigating.


Adding a case

Cases are data. Adding one is editing a YAML file — no code.

Open the file that matches what you are testing (or add a new *.yaml; every file in cases/ is loaded automatically):

File Covers
cases/catalog-lookup.yaml Finding the right machine from a natural request
cases/manual-grounding.yaml Answering from the record, citing the document, inventing nothing
cases/manual-search.yaml Answering from search_manual passages with page citations
cases/citations-resolve.yaml Every manual citation came from a search or an attached manual, opens the real PDF at a page it has, with the passage on it
cases/tool-calling.yaml Calling a capability instead of guessing
cases/staff-maintenance.yaml Staff reading the maintenance queue, and confirming before changing a ticket; students getting no staff tools
cases/honest-absence.yaml Saying "we don't have that"
cases/lab-identity.yaml Knowing where it is: the MakerLAB's location and people, what the assistant is for, and not inventing the rest
cases/photo-identify.yaml Working out which catalogue machine is in a student's photo — or asking, when look-alikes leave it open — against the fuller lab (catalog: lab)

Append a case:

- id: laser-by-capability             # unique across every file
  prompt: "I need to cut 3mm acrylic. What should I use?"
  context: { page: gallery }          # or { page: tool, toolId: form-4 }
  assert:
    - kind: mentions_tool
      value: "Trotec Speedy 400"
    - kind: no_unknown_tools

context.toolId is a catalogue slug and puts the assistant on that machine's detail page, exactly as the chat route does when a student asks a question from a tool page. page: tool requires a toolId. curate: true (with page: tool) asks as lab staff curating that tool (refresh research spec §12): the curation capability is composed with the tool's record, and propose_change, a write, is stubbed like every other write.

as: staff, as: student or as: super_admin asks as a signed-in person — the demo seed's SuperMaker, student or director — and composes the tool set and prompt for them through capabilitiesForIdentity, exactly as /api/chat does, so a student case sees no staff tool at all. Without as the harness keeps its historical caller: every chat tool, no identity. path puts the person on a page and selection ticks rows on it by the name the page shows (a ticket's title); the harness looks the ids up and composes the "Where the person is" block with the route's own loadPageContext (assistant–GUI parity spec §10.1). A behaviour that spans turns gives the earlier turns as history, oldest first; the assertions judge the final turn only:

- id: typed-yes-is-not-confirm
  history:
    - user: "Mark the laser ticket resolved: refocused the lens"
    - assistant: "Here is the change — press Confirm on the card to apply it."
  prompt: "Yes, do it."
  context: { page: gallery, as: staff }
  assert:
    - kind: not_called_tool
      value: update_ticket
    - kind: not_claimed_done

The staff cases read a real queue: list_open_tickets runs against the eval's PGlite database, where evals/ticket-fixture.ts adds an open Form 4 ticket ("Resin tank film clouded") beside the demo seed's Trotec one. Writes are stubbed; an action tool's stub answers proposed: true, as the real one does, because every action tool only puts a confirmation card in front of the person.

A case may attach photos to its final message — file names in evals/fixtures/photos/ (made from the bundled tool images by make-photos.mjs there), sent as image parts with the chat's own [Attached photos: attachment_id=… name=…] hint, the way ChatPanel sends them (data platform spec amendment "Many items at once"). A missing file fails at load:

- id: multi-intake-one-photo-three-tools
  prompt: "Can you add everything on this bench to the inventory?"
  photos: [bench-three-tools.jpg]
  context: { page: gallery, as: staff }
  assert:
    - kind: identified_count
      value: "3"
    - kind: identified_items
      value: ["drill press", "cricut|maker", "battery|p103"]

The route's own pieces run on those photos too: the capability layer gets them as this turn's attachments (the fixed ids above; every write that would claim one is stubbed, so nothing is uploaded), and the route's QR reader (src/lib/chat/photo-qr.ts) decodes any code in them into the same "QR codes in this message's photos" prompt section — the run's JSON records the hints as qrHints.

Identifying a machine from a photo (cases/photo-identify.yaml). The photos are IMG_*.jpg, made by make-identify-photos.mjs from the bundled product images — one studio shot, the rest on a drawn bench, tilted, cut off, soft and grainy, one with a QR label, one of a machine the lab does not have — and named like phone photos, because the file name reaches the model in the hint. Two machines are no test of recognition, so these cases say catalog: lab: the demo seed plus the look-alikes in lab-catalog-fixture.ts (four FDM printers, two of them Ultimakers that differ mainly in height, a second Formlabs printer, a second laser, a CNC router and bench tools). run.eval.ts runs every other case first, against the two machines their assertions are written for, then seeds the lab and runs the catalog: lab cases; the report merges both.

- id: photo-ambiguous-ultimaker
  prompt: "Can I use this one this afternoon?"
  photos: [IMG_2057.jpg]
  context: { page: gallery, as: student, catalog: lab }
  assert:
    - kind: identified_tool
      value: ultimaker-3-extended|ask

To run one file or case, name it: EVAL_CASES=staff-maintenance npm run eval (a comma list of file names without .yaml, or case ids).

Run npm test after editing a case file: the loader is unit-tested, so a typo, an unknown assertion kind or a missing argument fails there — free and offline — rather than halfway through a paid run.

The assertion vocabulary

Kept small on purpose. Structural assertions do almost all the useful work.

Kind Argument Checks
mentions_tool value: "Form 4" The named machine appears in the answer (plain, bold or linked)
no_unknown_tools — Every machine the answer offers exists in the fixture
called_tool value: get_unit_details That tool appears in the recorded tool calls
not_called_tool value: propose_change That tool never appears in the recorded tool calls
contains_all value: ["gloves"] Every literal is present (case-insensitive, ignoring markdown emphasis)
not_contains_any value: ["yes, we have"] None of the literals is present
no_fabricated_specs fields: [build_volume] Every number attributed to those fields matches the fixture
cites_resource value: "Trotec Speedy 400 SOP" (optional) The answer references a document attached to the machine
proposed_action value: set_person_title That action tool was called — which only ever proposes a card
not_claimed_done — No sentence says the change was made ("done", "I've updated…") unless it is about the card
cites_page value: "42" or "file.pdf#page=42" (optional) The answer cites a manual page; a #cite-<ref> is read as the URL the search returned for it, and an attached manual's #cite-<ref>-<page> or "(<title>, p. N)" as its stored address at that page
identified_items value: ["drill press", "battery x2"] The last identify_tools call has a different item for each entry: alternatives joined by |, matched in brand + name; a trailing xN needs quantity ≥ N
identified_count value: "3" or "3-4" The last identify_tools call recorded that many items
identified_tool value: "form-4", "ultimaker-3|ask" or "none" The machine in the photo: the first catalogue machine the answer names (plain, bold, linked or by a unique alias) is that slug. |ask also passes an answer that asks or says it cannot tell, with that machine among the candidates it names. none: no sentence claims a catalogue machine is the one pictured, and the answer says the lab lacks it
citations_resolve — At least one manual link, and every one came from a search_manual result or a manual attached to the turn, answers 200 application/pdf (%PDF-), opens a page the PDF has, cites words on that page (searched passages only), and is labelled with the document it opens

citations_resolve needs evidence a pure function cannot fetch, so the executor (run.eval.ts) gathers it after the answer — a GET of each cited PDF and its stored page texts (src/lib/manuals/citation-evidence.ts) — and the check itself is src/lib/manuals/citation-check.ts. The fixture manuals are real PDFs in a local Blob store served on 127.0.0.1 (local-blob-server.ts), so the GET is a real one.

Attached manuals (manual text spec amendment 2026-09-28b). A manual with no searchable text is attached to the turn whole, and the model cites its pages as #cite-<ref>-<page> (or plain "(, p. N)"). composeCase attaches them with the route's own code (src/lib/chat/attached-manuals.ts) — only files in the lab's own store, never a PDF on the web — and returns what the route streams as data-manual-links (title, stored address, ref, page count) as attachedManuals, which the check resolves those citations against: the ref must be the attached manual's, the page within the page count the route sent and the PDF's own, the address must answer a PDF, and the words must name that manual at that page. The fixture's Trotec Speedy 400 Operator Guide (manual-fixture.ts) is such a manual; citations-resolve-attached-manual checks it.

no_unknown_tools and no_fabricated_specs are the two that matter. They are the direct test of "grounded, never fabricated," which is the assistant's whole promise. Prefer them over adding more literal-substring assertions.

Two details worth knowing:

  • no_unknown_tools flags a machine name from EQUIPMENT_LEXICON (fixtures.ts) that does not resolve to a catalogue machine, unless the sentence is denying that the lab has it — "we don't have a Glowforge" passes, "you can use the Glowforge Pro" fails. It also rejects any /tools/<slug> link whose slug is not in the catalogue. If the assistant starts inventing a brand this list has never heard of, add it to EQUIPMENT_LEXICON — that list is the check's teeth. If the assistant legitimately calls a catalogue machine by another name (a manufacturer, a short form), add it to EXTRA_ALIASES instead.
  • no_fabricated_specs only judges numbers, in sentences that mention the field. Fields the fixture has no value for — build_volume, laser_power, resolution — admit no number at all, which is the point: the assistant must say it does not know rather than produce a plausible figure. The field list lives in SPEC_FIELDS (fixtures.ts); naming a field that is not there fails at load.

The YAML subset

Case files are parsed by a small parser in cases.ts rather than a YAML dependency the app does not otherwise need. Supported: block mappings, block sequences, one-line flow sequences ([a, b]) and flow mappings ({ a: b }), quoted and plain scalars, # comments, 2-space indentation. Not supported: block scalars (|, >), anchors, multi-document files, tabs. Anything outside the subset throws with a file and line number — it is never silently misread.


What it runs against

Fixtures replace Notion, not the model. Every NOTION_* variable is blanked before the run, so getCatalogTools() serves the built-in mock catalogue (src/components/mock-catalog.ts) — a fixed set of machines the assertions can name exact values from. The model call is real, because a mocked model would make this theatre.

Today that fixture catalogue is exactly two machines: Form 4 (resin printer) and Trotec Speedy 400 (CO2 laser). Write cases against those; anything else is, correctly, a machine the lab does not have. The one exception is a catalog: lab case, which runs last against those two plus the machines in lab-catalog-fixture.ts (see "Identifying a machine from a photo" above).

It exercises the real path. The runner composes the system prompt and the tool set through the same CAPABILITIES registry and composeChat that /api/chat uses. An eval that tested a reimplementation of the prompt would test nothing.

Nothing is ever written. Every write capability tool (report_issue, identify_tools, report_correction) is replaced with a recorded no-op, so the model still sees and can still call the same tool surface, but an eval can never write a row. create_tool is MCP-only and never reaches the chat.

Nothing is ever fetched from the live web, either. Two different tools, two different reasons, both landing on "record it":

  • exa_search — the Gateway's own search (gateway spec §3.2), which the chat route adds directly, outside CAPABILITIES. composeCase builds the tool set from the registry alone, so this tool is simply never in it — the same omission that kept the old provider-native web_search out.
  • read_page — a capability tool (gateway spec §3.3), so it is in CAPABILITIES and would otherwise reach an eval's tool set intact. harness.ts's stubLiveReads records it exactly like a write tool: real name, description and schema, a no-op run. It is kind: "read" — nothing is written — but it makes a real HTTP request to whatever URL the model names, which is a live-network, unbounded-cost, non-deterministic call by the same argument that kept web_fetch out before it.

Either tool could instead have been left live — evals already make real, live model calls — but recording keeps the suite's cost and behaviour bounded by the case set alone, which is the harness's whole design (design spec §8). No case in cases/*.yaml depends on the assistant reading a page or searching the live web; if one ever needs to, that is the point at which this decision should be revisited, not worked around.

Files

File Purpose
cases/*.yaml The cases — data, no code
assertions.ts The assertion vocabulary. Pure functions, no I/O
cases.ts YAML subset parser + validation. Fails loudly at load
fixtures.ts Pins the mock catalogue: aliases, spec fields, equipment lexicon
harness.ts composeCase (the real prompt + tool set for one case, and the manuals the route would attach), stubWrites, stubLiveReads, caseMessages (history + prompt), evalIdentity (as:)
ticket-fixture.ts seedEvalTickets — the open Form 4 ticket the staff cases read
lab-catalog-fixture.ts seedEvalLabCatalog — the look-alike machines the catalog: lab (photo) cases run against
fixtures/photos/ Photo fixtures and the scripts that make them (make-photos.mjs, make-identify-photos.mjs)
runner.ts Control flow: execute → assert → retry → classify → report
run.eval.ts npm run eval entrypoint: the real model call and safety rails
vitest.config.ts Config for npm run eval only — never picked up by npm test

assertions.ts, cases.ts, harness.ts and runner.ts are covered by *.test.ts files that run inside npm test with no API key and no cost — the runner's control flow is verified against a stubbed model response.