What to switch on once the site is live, so that a failure reaches a person instead of waiting to be noticed. Setting the site up:
deploy.md. Day-to-day running:handover.md. Design and reasons: the operational hardening spec.
Everything here is free or nearly so, and none of it is a service to run. Send every alert to a shared lab address, never one person's inbox.
Five steps, in order of value. The first three take about fifteen minutes together.
| # | What | Where | Alert on |
|---|---|---|---|
| 1 | Uptime check on /api/health |
UptimeRobot (free) or Better Stack (free) | HTTP status ≠ 200 |
| 2 | Nightly job heartbeat | Healthchecks.io (free) | A /fail ping, or no ping in 25 hours |
| 3 | AI Gateway budget | Vercel → AI Gateway | Spend reaching the monthly budget |
| 4 | Vercel usage and spend | Vercel → Settings → Billing / Usage | Approaching a plan limit or the spend amount |
| 5 | Log drain (optional) | Vercel → Log Drains, or a Marketplace integration | — |
Create an HTTP(S) monitor for https://<your-site>/api/health:
- Interval: 5 minutes (UptimeRobot's free floor; Better Stack's free tier allows 3).
- Alert on the status code, not the body. The endpoint answers 200 when the database is reachable and 503 when it is not — the 503 is the whole contract. A keyword check on the body would miss it.
- Confirm after 2 failures (or "incident after 2 checks") so a single slow cold start does not page anybody.
- Recipients: the shared lab address.
The endpoint is public, unauthenticated, cached for about 30 seconds and rate-limited to 30 a minute per IP, so a monitor polling every few minutes costs nothing. It does not check the model provider on purpose: a model outage does not make the catalogue wrong.
The nightly job (/api/cron/daily, 07:17 UTC) already fails loudly: any stage that breaks
makes the run non-200 and logs [cron] <stage> failed. But Vercel tells nobody — a red
run sits in Project → Cron Jobs until someone looks. A heartbeat monitor closes that gap, and
also catches a job that never ran at all (a cron that was unregistered looks exactly like one
that is working).
- On Healthchecks.io (free, 20 checks), add a check with
Schedule: cron
17 7 * * *, time zone UTC, Grace time: 1 hour. Better Stack's free heartbeat works the same way (period 1 day, grace 1 hour). - Copy its ping URL into Vercel as
CRON_HEARTBEAT_URL(Production only; preview deployments do not run cron). It must behttps. Treat it as a secret: anyone holding it can mark the job healthy. - Redeploy, then trigger the job once by hand (below) and watch the check turn green.
Each run that passes the secret check pings <url> on success and <url>/fail on any
failure, including "no Blob store linked". A heartbeat that cannot be delivered is logged and
never fails the job. The code is src/lib/cron/heartbeat.ts.
Also on /admin: a super admin on the live database sees a warning on the admin home when
the newest backup is more than 36 hours old, when there is none, when no Blob store is linked,
or when the store could not be read (src/lib/cron/backup-freshness.ts). It judges by the
backup files themselves, so a run that reported success without writing still shows.
To run the job by hand:
GET https://<your-site>/api/cron/daily
Header: x-admin-secret: <ADMIN_REVALIDATE_SECRET>
Inference is the one cost that scales with use and the one that can run away. Vercel
dashboard → AI Gateway → set a monthly budget with an email alert. Keep auto
top-up off, or capped: with it off, spend can never exceed the credit balance, and an
exhausted balance shows as assistant errors (the troubleshooting table in deploy.md says to
check this first). Every model call the app makes — chat, research, embeddings, evals — goes
through the Gateway, so this one limit covers all of them.
- Pro: Team → Settings → Billing → Spend Management. Set a spend amount; Vercel emails as it is approached. Leave "pause projects" off unless an outage is preferable to the bill.
- Hobby: there is no overage bill — limits are hard and the project stops serving when one is hit. Look at Usage monthly; Blob operations and function duration are the ones this app could plausibly approach (the nightly job, manual archiving and research runs).
Vercel keeps runtime logs only briefly, which is enough to read a failed cron run the same
day and not enough for "when did this start?". If that question comes up more than once,
add a log drain (Pro) or a Marketplace logging integration and send it the [cron],
[admin] and DbUnavailableError lines (backup failures log as [cron] backup failed:). Sentry is the documented next step after that
(ops spec §3.5) and is deliberately not the default: one more account to hand over.
One JSON file a night in the private Blob store, backups/YYYY-MM-DD.json: every
Postgres table, discovered from the schema so a new table is included without anyone
remembering. Three kinds of thing are left out on purpose
(src/lib/cron/backup-policy.ts):
- Credentials.
session,verificationandoauth_access_tokenare skipped; the tokens onaccount, the client secret onoauth_applicationand the Notion token onnotion_mirrorsare blanked. A backup must not double as a way to sign in. - Manual search data.
manual_pages(page text) andmanual_chunks(passages and their embeddings) are most of the bytes and are rebuilt from the stored PDFs.manual_documentsis kept. - Files. Photos and manual PDFs live in Blob, not in the file.
The file contains student names and email addresses, and quarterly copies keep them for up to three years — including rows the app itself has since deleted, such as discarded intake items. It stays private and belongs in the university's data inventory with that retention.
The same job prunes, by the date in each file's name (src/lib/cron/backup-retention.ts):
| Age | Kept |
|---|---|
| Under 7 days | Every night |
| 7 days to 1 month | The newest of each ISO week |
| 1 month to 1 year | The newest of each calendar month |
| 1 to 3 years | The newest of each quarter |
| Over 3 years | Nothing |
About 30–35 files at steady state. Two files on the same day keep the newer; a pathname that is not a backup is never deleted.
There is no automated restore. [dev]
-
Download the file from Vercel → Storage → the private Blob store →
backups/. -
Load its
tablesinto an empty, migrated database (npm run db:migratefirst), parents before children.rowCountbeside each table says what to expect. -
Rebuild manual search, from the repo root:
npm run manuals:index -- --force
--forceis required: the restoredmanual_documentsrows already carry the current extractor and chunker versions, so without it the script finds nothing to do. It re-reads every stored PDF (free) and re-embeds the passages through the AI Gateway (a few cents for the whole library), so it needsDATABASE_URL, Blob access and Gateway auth — seedeploy.mdStages 2d–2e. -
Everyone signs in again (sessions are not backed up), and each connected MCP app and Notion mirror asks for its token again.
A reviewed data bundle in data/inventory-cleanup-2026-09-28/ — manual PDF links for tools
that had none, display-name and unit-label fixes, SEO-spam tags removed, starter questions for
tools with none, tracking parameters off resource links. report.md there lists every change.
From the repo root, with the dev server stopped (the local database is single-process):
PGLITE_DATA_DIR=.pglite-data npm run inventory:cleanup -- --dry-run # what would change
PGLITE_DATA_DIR=.pglite-data npm run inventory:cleanup -- --apply
npm run manuals:index # after the archive copies the new PDFsHosted: DATABASE_URL instead of PGLITE_DATA_DIR; --revalidate <site> drops the catalogue
cache (ADMIN_REVALIDATE_SECRET). Writes go through the data layer (updateTool, and
tool-touching transactions for units and resources, as the editor does); every change is
conditional on the value it replaces, so a second run writes nothing. Low-confidence manual
links are skipped unless --include-low.