Every other test in this repo checks code. These check the assistant — the part students actually use, and the part that can get quietly worse while the whole test suite stays green. A prompt edit, a model bump, a capability rename or a change to how manuals are attached can all degrade answers without breaking a single unit test.
The suite checks the three behaviours that carry the assistant's value:
- it finds the right machine from a natural request,
- it grounds answers in the catalogue and says where they came from,
- it calls the right tool instead of inventing an answer,
plus the one that matters most when it regresses: it admits what the lab does not have.
Design spec: docs/specs/2026-07-29-agent-eval-harness-design.md.
Every model call goes through the Vercel AI Gateway (gateway spec §3.1) — there
is no direct-provider path any more, and ANTHROPIC_API_KEY is read by
nothing.
export AI_GATEWAY_API_KEY=... # or run where VERCEL_OIDC_TOKEN is already set
npm run eval
EVAL_MODEL=openai/gpt-6-sol npm run eval # try a different model
EVAL_MODEL=anthropic/claude-sonnet-5 npm run eval # or a different provider entirelyThe suite runs the deployment's own model — job chat in src/lib/ai/models.ts
(openai/gpt-6-luna by default, MODEL_CHAT overrides it) — so a run says
something about production. EVAL_MODEL overrides the eval's model
independently of MODEL_CHAT, with an explicit Gateway id (provider/model,
lower case), which is exactly what this harness is for when the question is
"does the next model still behave?".
Important
This makes real, paid model calls — roughly one per case, plus a retry for
any case that fails. It is deliberately not part of npm test or
npm run test:all, which stay free, offline and green, and it must never be
wired into a pull-request trigger where a fork could spend your money.
When to run it: before merging a change to a prompt fragment, a capability, or a job's model id — and before a demo. That is the whole policy.
Before pointing MODEL_CHAT (or the deployment's default) at a new model, the
gateway spec's §10 gate is: the model must pass every honest-absence and
manual-grounding case, and all but one of the rest — run twice.
EVAL_MODEL=openai/gpt-6-luna npm run eval
EVAL_MODEL=openai/gpt-6-luna npm run eval # again — a FLAKY case on run 1 must not recur on run 2The 2026-09-23 results are recorded, run by run, in the gateway spec: Luna first
missed the gate (amendment "The chat eval gate") and chat ran on
anthropic/claude-sonnet-5; after the chat prompt was tuned, Luna passed it on
three runs and chat moved to openai/gpt-6-luna (amendment "Chat prompt tuning
for Luna"). The anonymous-report behaviour is covered by two cases: with nobody
signed in the assistant asks for a name first (issue-report-asks-for-name), and
a reporter who declines still gets the ticket filed (issue-report-calls-tool):
an anonymous report beats no report. Luna passed all 14 cases on 2026-09-23.
Read every cases/honest-absence.yaml and cases/manual-grounding.yaml case as
a hard requirement (no FAIL, no FLAKY, on either run) and every other case
file with one miss tolerated per run. A structural failure — the same case
failing both runs, or failing for a different reason each time — is worth
investigating before the model goes live regardless of the count; see
"FLAKY is not a pass" below. This is a policy for a person driving the gate by
hand, not a script: nothing in evals/ enforces the "all but one" count for
you, because report.totals.failed is checked strictly (toBe(0)) — a
run with one tolerated miss still exits non-zero, on purpose, so it is never
silently green in a script.
Output is a per-case line, a detail block for anything that did not pass, and a
JSON artifact at evals/.last-run.json (gitignored).
PASS laser-by-capability
FLAKY unit-status-calls-tool
FAIL form4-build-volume-unknown
FAIL form4-build-volume-unknown
prompt: What's the exact build volume and layer resolution on this printer?
x no_fabricated_specs — expected numbers attributed to build_volume to match the fixture
"145" is attributed to build_volume, but the fixture records no value for it
answer: "The Form 4 has a build volume of 145 x 145 x 185 mm."
11/12 passed · 1 flaky · 1 failed · 0 errored
FLAKY is not a pass. A failing case is retried once; one that passes only on the retry is reported as flaky and never counted as a pass. Flakiness is signal — a case that flakes is usually badly written. It does not fail the run (one flake in a dozen cases is noise), but a structural failure is always worth investigating.
Cases are data. Adding one is editing a YAML file — no code.
Open the file that matches what you are testing (or add a new *.yaml; every
file in cases/ is loaded automatically):
| File | Covers |
|---|---|
cases/catalog-lookup.yaml |
Finding the right machine from a natural request |
cases/manual-grounding.yaml |
Answering from the record, citing the document, inventing nothing |
cases/manual-search.yaml |
Answering from search_manual passages with page citations |
cases/citations-resolve.yaml |
Every manual citation came from a search or an attached manual, opens the real PDF at a page it has, with the passage on it |
cases/tool-calling.yaml |
Calling a capability instead of guessing |
cases/staff-maintenance.yaml |
Staff reading the maintenance queue, and confirming before changing a ticket; students getting no staff tools |
cases/honest-absence.yaml |
Saying "we don't have that" |
cases/lab-identity.yaml |
Knowing where it is: the MakerLAB's location and people, what the assistant is for, and not inventing the rest |
cases/photo-identify.yaml |
Working out which catalogue machine is in a student's photo — or asking, when look-alikes leave it open — against the fuller lab (catalog: lab) |
Append a case:
- id: laser-by-capability # unique across every file
prompt: "I need to cut 3mm acrylic. What should I use?"
context: { page: gallery } # or { page: tool, toolId: form-4 }
assert:
- kind: mentions_tool
value: "Trotec Speedy 400"
- kind: no_unknown_toolscontext.toolId is a catalogue slug and puts the assistant on that machine's
detail page, exactly as the chat route does when a student asks a question from
a tool page. page: tool requires a toolId. curate: true (with page: tool) asks as lab staff curating that tool (refresh research spec §12): the
curation capability is composed with the tool's record, and propose_change,
a write, is stubbed like every other write.
as: staff, as: student or as: super_admin asks as a signed-in person —
the demo seed's SuperMaker, student or director — and composes the tool set
and prompt for them through capabilitiesForIdentity, exactly as /api/chat
does, so a student case sees no staff tool at all. Without as the harness
keeps its historical caller: every chat tool, no identity. path puts the
person on a page and selection ticks rows on it by the name the page shows
(a ticket's title); the harness looks the ids up and composes the "Where the
person is" block with the route's own loadPageContext (assistant–GUI parity
spec §10.1). A behaviour that spans turns gives the earlier turns as
history, oldest first; the assertions judge the final turn only:
- id: typed-yes-is-not-confirm
history:
- user: "Mark the laser ticket resolved: refocused the lens"
- assistant: "Here is the change — press Confirm on the card to apply it."
prompt: "Yes, do it."
context: { page: gallery, as: staff }
assert:
- kind: not_called_tool
value: update_ticket
- kind: not_claimed_doneThe staff cases read a real queue: list_open_tickets runs against the eval's
PGlite database, where evals/ticket-fixture.ts adds an open Form 4 ticket
("Resin tank film clouded") beside the demo seed's Trotec one. Writes are
stubbed; an action tool's stub answers proposed: true, as the real one does,
because every action tool only puts a confirmation card in front of the
person.
A case may attach photos to its final message — file names in
evals/fixtures/photos/ (made from the bundled tool images by
make-photos.mjs there), sent as image parts with the chat's own
[Attached photos: attachment_id=… name=…] hint, the way ChatPanel sends
them (data platform spec amendment "Many items at once"). A missing file fails
at load:
- id: multi-intake-one-photo-three-tools
prompt: "Can you add everything on this bench to the inventory?"
photos: [bench-three-tools.jpg]
context: { page: gallery, as: staff }
assert:
- kind: identified_count
value: "3"
- kind: identified_items
value: ["drill press", "cricut|maker", "battery|p103"]The route's own pieces run on those photos too: the capability layer gets them
as this turn's attachments (the fixed ids above; every write that would
claim one is stubbed, so nothing is uploaded), and the route's QR reader
(src/lib/chat/photo-qr.ts) decodes any code in them into the same "QR codes
in this message's photos" prompt section — the run's JSON records the hints
as qrHints.
Identifying a machine from a photo (cases/photo-identify.yaml). The
photos are IMG_*.jpg, made by make-identify-photos.mjs from the bundled
product images — one studio shot, the rest on a drawn bench, tilted, cut off,
soft and grainy, one with a QR label, one of a machine the lab does not have —
and named like phone photos, because the file name reaches the model in the
hint. Two machines are no test of recognition, so these cases say
catalog: lab: the demo seed plus the look-alikes in lab-catalog-fixture.ts
(four FDM printers, two of them Ultimakers that differ mainly in height, a
second Formlabs printer, a second laser, a CNC router and bench tools).
run.eval.ts runs every other case first, against the two machines their
assertions are written for, then seeds the lab and runs the catalog: lab
cases; the report merges both.
- id: photo-ambiguous-ultimaker
prompt: "Can I use this one this afternoon?"
photos: [IMG_2057.jpg]
context: { page: gallery, as: student, catalog: lab }
assert:
- kind: identified_tool
value: ultimaker-3-extended|askTo run one file or case, name it: EVAL_CASES=staff-maintenance npm run eval
(a comma list of file names without .yaml, or case ids).
Run npm test after editing a case file: the loader is unit-tested, so a typo,
an unknown assertion kind or a missing argument fails there — free and offline —
rather than halfway through a paid run.
Kept small on purpose. Structural assertions do almost all the useful work.
| Kind | Argument | Checks |
|---|---|---|
mentions_tool |
value: "Form 4" |
The named machine appears in the answer (plain, bold or linked) |
no_unknown_tools |
— | Every machine the answer offers exists in the fixture |
called_tool |
value: get_unit_details |
That tool appears in the recorded tool calls |
not_called_tool |
value: propose_change |
That tool never appears in the recorded tool calls |
contains_all |
value: ["gloves"] |
Every literal is present (case-insensitive, ignoring markdown emphasis) |
not_contains_any |
value: ["yes, we have"] |
None of the literals is present |
no_fabricated_specs |
fields: [build_volume] |
Every number attributed to those fields matches the fixture |
cites_resource |
value: "Trotec Speedy 400 SOP" (optional) |
The answer references a document attached to the machine |
proposed_action |
value: set_person_title |
That action tool was called — which only ever proposes a card |
not_claimed_done |
— | No sentence says the change was made ("done", "I've updated…") unless it is about the card |
cites_page |
value: "42" or "file.pdf#page=42" (optional) |
The answer cites a manual page; a #cite-<ref> is read as the URL the search returned for it, and an attached manual's #cite-<ref>-<page> or "(<title>, p. N)" as its stored address at that page |
identified_items |
value: ["drill press", "battery x2"] |
The last identify_tools call has a different item for each entry: alternatives joined by |, matched in brand + name; a trailing xN needs quantity ≥ N |
identified_count |
value: "3" or "3-4" |
The last identify_tools call recorded that many items |
identified_tool |
value: "form-4", "ultimaker-3|ask" or "none" |
The machine in the photo: the first catalogue machine the answer names (plain, bold, linked or by a unique alias) is that slug. |ask also passes an answer that asks or says it cannot tell, with that machine among the candidates it names. none: no sentence claims a catalogue machine is the one pictured, and the answer says the lab lacks it |
citations_resolve |
— | At least one manual link, and every one came from a search_manual result or a manual attached to the turn, answers 200 application/pdf (%PDF-), opens a page the PDF has, cites words on that page (searched passages only), and is labelled with the document it opens |
citations_resolve needs evidence a pure function cannot fetch, so the executor
(run.eval.ts) gathers it after the answer — a GET of each cited PDF and its
stored page texts (src/lib/manuals/citation-evidence.ts) — and the check
itself is src/lib/manuals/citation-check.ts. The fixture manuals are real
PDFs in a local Blob store served on 127.0.0.1 (local-blob-server.ts), so
the GET is a real one.
Attached manuals (manual text spec amendment 2026-09-28b). A manual with no
searchable text is attached to the turn whole, and the model cites its pages as
#cite-<ref>-<page> (or plain "(, p. N)"). composeCase attaches
them with the route's own code (src/lib/chat/attached-manuals.ts) — only
files in the lab's own store, never a PDF on the web — and returns what the
route streams as data-manual-links (title, stored address, ref, page count)
as attachedManuals, which the check resolves those citations against: the
ref must be the attached manual's, the page within the page count the route
sent and the PDF's own, the address must answer a PDF, and the words must name
that manual at that page. The fixture's Trotec Speedy 400 Operator Guide
(manual-fixture.ts) is such a manual; citations-resolve-attached-manual
checks it.
no_unknown_tools and no_fabricated_specs are the two that matter. They
are the direct test of "grounded, never fabricated," which is the assistant's
whole promise. Prefer them over adding more literal-substring assertions.
Two details worth knowing:
no_unknown_toolsflags a machine name fromEQUIPMENT_LEXICON(fixtures.ts) that does not resolve to a catalogue machine, unless the sentence is denying that the lab has it — "we don't have a Glowforge" passes, "you can use the Glowforge Pro" fails. It also rejects any/tools/<slug>link whose slug is not in the catalogue. If the assistant starts inventing a brand this list has never heard of, add it toEQUIPMENT_LEXICON— that list is the check's teeth. If the assistant legitimately calls a catalogue machine by another name (a manufacturer, a short form), add it toEXTRA_ALIASESinstead.no_fabricated_specsonly judges numbers, in sentences that mention the field. Fields the fixture has no value for —build_volume,laser_power,resolution— admit no number at all, which is the point: the assistant must say it does not know rather than produce a plausible figure. The field list lives inSPEC_FIELDS(fixtures.ts); naming a field that is not there fails at load.
Case files are parsed by a small parser in cases.ts rather than a YAML
dependency the app does not otherwise need. Supported: block mappings, block
sequences, one-line flow sequences ([a, b]) and flow mappings ({ a: b }),
quoted and plain scalars, # comments, 2-space indentation. Not supported:
block scalars (|, >), anchors, multi-document files, tabs. Anything outside
the subset throws with a file and line number — it is never silently misread.
Fixtures replace Notion, not the model. Every NOTION_* variable is blanked
before the run, so getCatalogTools() serves the built-in mock catalogue
(src/components/mock-catalog.ts) — a fixed set of machines the assertions can
name exact values from. The model call is real, because a mocked model would
make this theatre.
Today that fixture catalogue is exactly two machines: Form 4 (resin printer)
and Trotec Speedy 400 (CO2 laser). Write cases against those; anything else
is, correctly, a machine the lab does not have. The one exception is a
catalog: lab case, which runs last against those two plus the machines in
lab-catalog-fixture.ts (see "Identifying a machine from a photo" above).
It exercises the real path. The runner composes the system prompt and the
tool set through the same CAPABILITIES registry and composeChat that
/api/chat uses. An eval that tested a reimplementation of the prompt would
test nothing.
Nothing is ever written. Every write capability tool (report_issue,
identify_tools, report_correction) is replaced with a recorded no-op, so the
model still sees and can still call the same tool surface, but an eval can never
write a row. create_tool is MCP-only and never reaches the chat.
Nothing is ever fetched from the live web, either. Two different tools, two different reasons, both landing on "record it":
exa_search— the Gateway's own search (gateway spec §3.2), which the chat route adds directly, outsideCAPABILITIES.composeCasebuilds the tool set from the registry alone, so this tool is simply never in it — the same omission that kept the old provider-nativeweb_searchout.read_page— a capability tool (gateway spec §3.3), so it is inCAPABILITIESand would otherwise reach an eval's tool set intact.harness.ts'sstubLiveReadsrecords it exactly like a write tool: real name, description and schema, a no-oprun. It iskind: "read"— nothing is written — but it makes a real HTTP request to whatever URL the model names, which is a live-network, unbounded-cost, non-deterministic call by the same argument that keptweb_fetchout before it.
Either tool could instead have been left live — evals already make real,
live model calls — but recording keeps the suite's cost and behaviour bounded
by the case set alone, which is the harness's whole design (design spec §8).
No case in cases/*.yaml depends on the assistant reading a page or searching
the live web; if one ever needs to, that is the point at which this decision
should be revisited, not worked around.
| File | Purpose |
|---|---|
cases/*.yaml |
The cases — data, no code |
assertions.ts |
The assertion vocabulary. Pure functions, no I/O |
cases.ts |
YAML subset parser + validation. Fails loudly at load |
fixtures.ts |
Pins the mock catalogue: aliases, spec fields, equipment lexicon |
harness.ts |
composeCase (the real prompt + tool set for one case, and the manuals the route would attach), stubWrites, stubLiveReads, caseMessages (history + prompt), evalIdentity (as:) |
ticket-fixture.ts |
seedEvalTickets — the open Form 4 ticket the staff cases read |
lab-catalog-fixture.ts |
seedEvalLabCatalog — the look-alike machines the catalog: lab (photo) cases run against |
fixtures/photos/ |
Photo fixtures and the scripts that make them (make-photos.mjs, make-identify-photos.mjs) |
runner.ts |
Control flow: execute → assert → retry → classify → report |
run.eval.ts |
npm run eval entrypoint: the real model call and safety rails |
vitest.config.ts |
Config for npm run eval only — never picked up by npm test |
assertions.ts, cases.ts, harness.ts and runner.ts are covered by
*.test.ts files that run inside npm test with no API key and no cost —
the runner's control flow is verified against a stubbed model response.