Practical runbook for the v5 test suite. For the full design rationale and the
per-file coverage matrix, see the design doc:
docs/specs/2026-05-29-v5-test-suite-design.md
(repo root). For harness internals (fixtures, mocks, render helper, the exact
import paths), see test/README.md.
The suite has four layers, all fully mocked — no live services. Tests never hit Postgres, Notion, the Vercel AI Gateway, or Upstash; there are no API keys and no network cost. Everything is deterministic and offline.
| Layer | What | Where it lives | Runner |
|---|---|---|---|
| Unit | src/lib, src/i18n pure logic |
colocated *.test.ts next to source |
Vitest |
| Integration | API routes (/api/*) with HTTP + module mocks |
colocated route.test.ts next to the route |
Vitest |
| Component | React UI via React Testing Library | colocated *.test.tsx next to the component |
Vitest |
| E2E | Full app against the PGlite demo catalog | e2e/*.spec.ts |
Playwright |
Shared infra lives in test/ (MSW server + handlers, fixtures, mocks, the RTL
render helper). Tests import from there; don't edit package.json or run
npm install — the harness deps and scripts are already wired.
npm run test:all # everything: lint + typecheck + vitest + playwright (one command)
npm test # vitest run — one-shot, runs unit + integration + component
npm run test:watch # vitest watch mode
npm run test:coverage # vitest run --coverage (v8 reporter: text + html)
npm run test:e2e # playwright test (E2E)
npm run test:e2e:ui # playwright test --ui (interactive)Prerequisite for E2E: run this once before your first npm run test:e2e
— the harness installs @playwright/test but not the browser binary:
npx playwright install chromiumPlaywright boots its own dev server (see E2E notes below), so no separate
npm run dev is required.
- MSW (Mock Service Worker) intercepts all outbound HTTP: Notion
(
api.notion.com/v1/*), the Upstash*/pipelineREST endpoint, and — for the workflow tier only, whichvi.mockcannot reach — the Gateway (https://ai-gateway.vercel.sh, viatest/gateway/msw.ts'sgatewayHandlers). Everywhere else, the model is stubbed one layer up, at the registry (test/ai/models-stub.ts) — see "Stubbing the model" below. The nodeserverlives intest/msw/server.ts; default handlers intest/msw/handlers.ts. Lifecycle (start / reset / stop) is managed invitest.setup.ts. Unhandled outbound requests fail the test by design (onUnhandledRequest: "error"). - The Notion mirror (
src/lib/mirror/*) is tested against a stateful in-memory Notion,test/fakes/notion-fake.ts(createNotionFake), which checks the bearer token and theNotion-Versionheader, validates page properties against the database schema, and can be told to fail (failNext, including 429 withRetry-After). Install it into a test's MSW server withuseNotionFake(server, fake)fromtest/msw/notion-mirror.ts— imported under another name (import { useNotionFake as installNotionFake }) outside a component, because ESLint's hooks rule reads anyuse*call as a hook. Workflow tests use the same fake through MSW, sincevi.mockdoes not reach step code. PGlite's session time zone is the machine's, so comparetimestamptztext in SQL ($1::timestamptz = …), never as strings. - Resend (staff email) is never reached. A test that sends stubs
RESEND_API_KEYandEMAIL_FROMand installsuseResendFake(server)fromtest/msw/resend.ts: it records every request, answers from a script of statuses (fake.script(503)), and treats a repeatedIdempotency-Keyas the same email (fake.delivered()). Workflow tests use it too. vi.mock("next/cache", …)—catalog.tsusescacheTag/cacheLifeandadmin/revalidate/route.tsusesrevalidateTag; these only work inside a Next build. Mock them with thenextCacheMock()factory fromtest/mocks/next-cache.ts.server-onlyis aliased to an empty stub (test/mocks/server-only.ts) invitest.config.ts, sorate-limit.tsand its importers load under Vitest.- The PGlite rule.
getCatalogTools()/getCatalogTool(id)read from an in-process PGlite database — real Postgres, compiled to WebAssembly, migrated and seeded with demo data (two tools:form-4,trotec-speedy-400) — wheneverDATABASE_URLis unset, which is the default in every test. So:- Demo-seed path (default in tests): leave
DATABASE_URLunset → the seeded PGlite database, no MSW needed, no Notion env at all. Put// @vitest-environment nodeat the top of any file that touches PGlite. - An isolated database:
createPgliteDb()fromsrc/lib/db/pglite.tsreturns a fresh instance for a test that seeds its own rows. - Writes (tickets, corrections, project submission, uploads, intake) go
to the same PGlite database — no write touches Notion. Only the one-way
mirror and the one-time import talk to Notion; their tests use the fake in
test/fakes/notion-fake.tsor the MSW handlers intest/msw/.
- Demo-seed path (default in tests): leave
Place tests next to the code they cover. describe / it / expect / vi and
the lifecycle hooks are globals — no imports needed.
Add a unit test — create src/lib/<name>.test.ts:
import { isSupportedLocale } from "@/i18n/config";
it("recognizes a supported locale", () => {
expect(isSupportedLocale("en")).toBe(true);
});Add a component test — create src/components/<Name>.test.tsx and use the
custom render (it wraps the component in NextIntlClientProvider with the
en messages, which i18n-aware components need):
import { render, screen, userEvent } from "../../test/utils/render";
import { ToolCard } from "./ToolCard";
import { availableTool } from "../../test/fixtures/catalog";
it("renders the tool name", () => {
render(<ToolCard tool={availableTool} />);
expect(screen.getByText(availableTool.name)).toBeInTheDocument();
});Add an MSW override — defaults live in handlers.ts; override per test with
server.use(...) (the afterEach reset undoes it):
import { server } from "../../test/msw/server";
import { http, HttpResponse } from "msw";
server.use(
http.post("https://api.notion.com/v1/databases/:id/query", () =>
HttpResponse.json({ object: "error" }, { status: 500 })
)
);Add a fixture — extend test/fixtures/notion.ts (raw NotionPage shapes +
the notionQueryResponse(pages, { hasMore }) pagination helper) or
test/fixtures/catalog.ts (resolved MakerLabTool / MakerLabUnit objects for
component tests). Reuse the existing exports before adding new ones.
Env stubbing for module-load-time reads. site-config.ts reads
NEXT_PUBLIC_* / AUDIENCE at module load, and rate-limit.ts computes its
Upstash branch from UPSTASH_* at load. vi.stubEnv after import won't change
those captured values — stub, then vi.resetModules(), then dynamic import():
it("honors NEXT_PUBLIC_SITE_NAME override", async () => {
vi.stubEnv("NEXT_PUBLIC_SITE_NAME", "Acme Lab");
vi.resetModules();
const { siteConfig } = await import("@/lib/site-config");
expect(siteConfig.name).toBe("Acme Lab");
});vi.unstubAllEnvs() and vi.restoreAllMocks() run automatically after every
test (the setup file). The in-memory rate limiter is a per-process singleton
Map — use distinct keys per test, or resetModules() + re-import for a fresh
window.
Blob mode. The setup file sets BLOB_LOCAL_DISABLE=1, so with no
BLOB_READ_WRITE_TOKEN a test sees "no store" (blob_not_configured), as a
deploy without one does. To test the local .blob-data/ store, stub
BLOB_LOCAL_DISABLE / VERCEL to "" and NODE_ENV to "development", and
point process.cwd() at a temp folder (src/lib/blob-local.test.ts).
BLOB_LOCAL_DIR (test-only) moves the folder and also allows the local store
in a production build — the intake E2E server's switch; never on Vercel.
Stubbing the model (chat route and elsewhere). The chat route's tool
execute functions are inline and its helpers are module-private, so don't
unit-test them directly. Instead mock the registry — vi.mock("@/lib/ai/models", …)
through test/ai/models-stub.ts — call POST(req), and assert on the
response; recordedCalls(model) gives you back the { prompt, tools, providerOptions } a stubbed model received. Never mock @ai-sdk/gateway
directly. The workflow tier, which vi.mock cannot reach,
stubs the Gateway's own HTTP boundary instead with test/gateway/msw.ts. The
full verified snippet and both seams are in
test/README.md.
- Playwright's
webServerbuilds and bootsnpx next start -p 3100withDATABASE_URLunset, so the app serves the seeded PGlite demo database regardless of your dev shell's environment.testDiris./e2e;baseURLishttp://localhost:3100. /api/chatis intercepted inside each spec at the network layer viapage.route()returning a UI-message stream chunk — no real model call.- Except
e2e/intake.spec.ts(gateway spec §10, formerly data platform spec §10 scenario 5), which needsidentify_toolsand the research workflow — including the image stage — to run on the server. Playwright boots a second web server,e2e/stubs/gateway-stub.tson port 3101 (plainnode:http, servingtest/gateway/wire.ts's builders undernode --experimental-strip-types), and the app reaches it throughAI_GATEWAY_BASE_URLwith a fakeAI_GATEWAY_API_KEY—READ_PAGE_TEST_ORIGINalso points at it, so the read step's and the chat'sread_pagefetches are exempted from the SSRF guard that would otherwise refuse a loopback address. The model is stubbed at the Gateway's own wire format, and everything else is the real app, including the Workflow SDK's local world. It is its own Playwright project (intake) that depends onchromium, so it runs after every other spec, and it runs against its own server: the same build started again withnpx next start -p 3103and a local Blob folder (BLOB_LOCAL_DIR=.blob-data-e2e, git-ignored). So the image stage cuts the backdrop out of the stub's one candidate (a product on plain white; the cutout is deterministic and calls no model), the review page preselects the cleaned copy, and the test approves it and checks the gallery card shows the published copy — while port 3100 keeps no Blob store forprojects.spec.ts's "uploads unavailable" assertion. Its demo database is its own. Run it alone withnpx playwright test --project=intake --no-deps(Playwright still starts every web server, including the 3100 build). - And
e2e/mirror.spec.ts(§10 scenario 8), the same shape for Notion: the mirror calls Notion from server actions and workflow steps, so a third web server,e2e/stubs/notion-stub.tson port 3102, answers/v1/*with the in-memory fake the Vitest suites use (test/fakes/notion-fake.ts), and the app reaches it throughNOTION_API_BASE_URL(test-only; production never sets it). Fixture values (fake token, page id and URL) are ine2e/stubs/notion-fixture.ts. It is its own project (mirror) that depends onintake, so it runs last of all: while a mirror is connected every write in the app schedules a push, and no other spec may be writing then. One test,retries: 0— each step is the next one's precondition. Run it alone withnpx playwright test --project=mirror --no-deps. reuseExistingServer: false— Playwright always boots its own fresh PGlite-backed server on the dedicated port 3100. This means E2E never collides with (or accidentally reuses) anext devyou have running on the default port 3000 against a real database, so results are deterministic no matter what you have running locally. Port 3000 is left untouched.
npm run test:coverage produces a v8 coverage report (text to stdout +
html). There is no enforced threshold — by design. Coverage is a
diagnostic for developers, not a gate.
.github/workflows/ci.yml runs on every PR and every push to main: one job
for npm run lint, npm run typecheck and npm run spec:coverage -- --ci,
and npx vitest run (both projects) split four ways with --shard, and the E2E
suite (npm run test:e2e) as its own job, e2e (playwright), which is not a
required check yet. After every production deployment,
.github/workflows/deploy-smoke.yml runs scripts/smoke-production.sh against
the live site (npm run smoke:production by hand; docs/deploy.md §7).
Actions are pinned to full commit SHAs (with the version in a comment) and
checkout runs with persist-credentials: false; .github/dependabot.yml opens
the monthly bump PRs. test/ci-hardening.test.ts fails if a step goes back to a
tag or drops the setting.