Skip to content

Devices, password-hash import, external billing, an operator console, and operator permissions - #39

Closed
libworky wants to merge 97 commits into
rekey-dev:mainfrom
libworky:contrib/rekey-2.2-features
Closed

libworky wants to merge 97 commits into
rekey-dev:mainfrom
libworky:contrib/rekey-2.2-features

Conversation

@libworky

@libworky libworky commented Sep 3, 2026 •

Copy link
Copy Markdown
Contributor

This branch carries the work from running Rekey as the licensing and entitlement layer for a desktop product. The earlier revision of this PR covered devices, password-hash import and a bring-your-own billing provider. This revision adds the operator surface those features needed, plus a permission model that the operator surface turned out to require. Each part is its own commit series and stands alone. If only some of it is interesting to you, tell me which and I will split the rest out.

1. Devices, password-hash import, external billing provider

Unchanged from the earlier revision. Devices are a table with a status of ACTIVE, RELEASED or BLOCKED, optionally bound to sessions through a dev claim, with the cap sold as an entitlement (max_devices). Password hashes can be imported so users of an older system sign in without a reset. The external billing provider accepts subscription events from a billing system Rekey does not talk to, signed with a shared secret, with a documented contract for the event feed.

2. An operator console for support work

Before this, an operator looking at an end-user could read a lockout counter and a device list and act on neither. Every remedy needed a developer with an API client.

The end-user page is now a set of routed tabs: overview, subscriptions, devices, credits, security, data. Seven routes back it: unlock, resend verification, send a password reset (with a reason that lands in the audit trail, since an unrequested reset mail looks the same to the recipient as an attacker who reached the panel), list and revoke sessions, and release every device. Operators can also grant and cancel subscriptions from the panel, which used to be super-admin only.

Two role floors moved while doing this. Erasing an end-user and the plain cascade delete are both workspace-owner only now. The cascade delete was gated on write alone, so a member with an application grant could delete a person while the erase form correctly refused them.

Three API bugs came out of this. Device security events were written with an application id but no workspace id, and the read path filtered on workspace id, so the entire device audit trail was recorded and shown nowhere. The fix is on the read side: deriving the workspace during the write put a database read on the audit path and reordered webhook deliveries. Open-ended subscription grants were cancelled immediately rather than at period end, because with no period there is nothing to schedule against.

3. Email control

There was no way to stop an application sending email: not globally, not per event, not for one address. A product whose own backend already sends transactional mail delivered two of everything.

Three gates now sit in dispatch(), in order: a per-application master switch (a column, because the credentials save rewrites the config JSON wholesale and a flag stored there would be dropped), a per-event switch, and a suppression list keyed on lowercased address. Each suppression writes a log row with status: suppressed and the reason, so the delivery view can answer why a mail did not go.

One rule here is load-bearing. A suppressed send returns error, not no_transport. no_transport is the documented contract that hands the raw reset or magic-link token back to the caller so a self-hosted deployment without email can deliver it. Reporting a suppression that way would have turned "we switched email off" into "the API now returns live reset tokens".

Disabling an email the live auth config depends on (password reset while password sign-in is on, verification while verification is required) is refused rather than warned about, and the same check runs when the auth config is changed, from both the REST route and the MCP tool, so the coupling holds in both directions.

The email page is split into settings, templates, delivery and suppressions.

4. Subscription import

The event feed covers everything from the moment a billing system connects. It cannot cover what was sold before, which on migration day is the whole book of business.

An import is a run, not a call. Starting one reads the provider and records, per row, what would happen and why: matched by email, created as an unlinked user, skipped because no plan maps, skipped because the user already has a live subscription, or invalid. Nothing is written until the operator applies the preview behind a typed confirmation. An import never overwrites live entitlement and never matches an erased user.

The external provider gains a read side: one paginated endpoint the customer hosts, with a documented contract (docs/external-billing-pull.md). The request signature is deliberately not the webhook signature. The webhook helper signs the body, and a GET has no body, so reusing it would sign a constant per second. This direction signs the method and path. The fetch gets the same hardening as an outbound webhook delivery: resolved-address checks, pinned connections, refused redirects, capped body reads.

5. Operator permissions

A workspace had one lever for a member: which applications they could touch, and one of three roles on each. Every grant conferred read on everything. A support agent given an application to work in saw its revenue and every payment, and the only alternative was refusing them the application.

Rekey already had scopes on personal access tokens and MCP tokens. This attaches one registry of scopes to the membership itself: seven domains (end-users, billing, auth-config, developer, organizations, activity, overview) at read or write, with write implying read. The three grant roles become presets over the same vocabulary, and the effective set for a request is the intersection of the grant preset and the membership's scopes, so neither can widen the other. Tokens intersect with their holder's scopes at request time.

Every operator route declares what it needs on its route config, and a test reads the live route table and fails the build if any route is unclassified. The decision that used to exist in three diverging copies (two REST helpers and an MCP one) is one function over a plain context. The scope check runs after the existing existence and grant checks, so an application the caller cannot see still answers 404 and only one they can see answers 403 with the missing scope named.

Grants, roles, invitations, lifecycle, impersonation and erasure are floors that no scope can unlock. A team:write scope on a member would let them grant themselves everything and rewrite their own scopes, so it does not exist.

/me and the application read return the caller's resolved scopes, and the panel renders navigation from them instead of inferring an operator's role from a redacted field, which it did. The team page gets a scope editor. MCP tools consult the same registry, with role as the ceiling: an owner or admin passes the gate before the row is read, and promoting a member clears their stored scopes. The security event each tool call records names the scope that admitted it, and the request log carries the same per request, since under scopes "who did it" no longer implies "what they were allowed to do".

Nothing changes for an existing deployment until an admin restricts somebody. The two new membership columns carry constant defaults and the migration rewrites no rows. One deliberate change for existing APP_BILLING holders: provider credentials and refunds were plain write before, so the billing role could do neither. Under the preset it can.

6. A per-workspace switch for the operator MCP server

Two kill switches existed already, one for the whole deployment and one per application for the end-user MCP server. Neither let a workspace owner say that no agent may act as any of their operators. A workspace flag does that now. Off refuses every access token presented for the workspace at auth time and grants no new consent. It does not gate refresh: an agent holding a refresh token keeps rotating it while the workspace is off, and what it mints is refused like any other token, so nothing is revoked and turning the flag back on restores every agent without a re-consent. The flag is read through the membership row auth already fetches, so it costs no extra query per call, and flipping it is a security event.

Token lifetimes as deployment settings

The end-user and operator access tokens were fixed at 15 minutes and both refresh tokens at 30 days. Four settings replace the constants, with the old values as defaults and ceilings of a day, half a day and a year: a desktop application that opens once a month wants a longer refresh window, and an operator on a shift wants an access token that outlasts it. Refresh tokens are sliding, so the window only has to cover the longest gap between uses.

A longer access token needed a way to end a session early, since password change, sign-out everywhere, session revoke and device release only revoked refresh tokens. Both user models carry a "sessions invalid before" instant that every one of those paths stamps; the session middlewares and the SDK's user read refuse an access token issued before it on its next use, with the code SDKs already refresh on, so sessions that were not ended renew silently and only the ended one lands on sign-in. The panel and portal set their session cookies from the expiries the auth responses already carried instead of constants of their own.

Panel navigation

The end-user page had reached three stacked tab strips and two back links before showing anything about the user. Nested pages now get at most two tab strips; a third level becomes a segmented control inside the record's own header, and a breadcrumb replaces the back links. Three panel bugs were fixed along the way: nested tab strips never set aria-current, nested headings rendered at the same size as the page heading, and two comments cited a design document that did not exist.

Panel parity

A route-by-route comparison of the operator API against the panel found four things the API could do that the panel could not show or do. Both event logs rendered the event type and actor only, so what an event actually said (which tool an agent called and under which scope, whether a switch went on or off, which plan a grant named and why) was written and shown nowhere; a whitelisted set of metadata keys now renders as chips. The request log's read schema never declared the admitting scope, so the serializer dropped it before any client saw it; it is declared and shown. Adjusting a subscription's entitlement overrides and ending every live impersonation of an end-user existed only as routes; both have a panel surface now, and the billing view returns a subscription's current overrides so the dialog can show what it is changing.

Security events by the person they are about

The end-user page asks one question of the audit log: what happened to this person. The log stored the answer in two places. A person's own events name them as the actor; anything done to them by an operator or the system names them in metadata.endUserId, with somebody else as the actor. Nothing could filter on "either", so the panel read the application's latest 200 events three times, once per actor type, and matched both fields in memory. That was 600 rows fetched to render twenty on every view, and it missed a quiet user's history on a busy application entirely.

recordSecurityEvent now derives a subject_end_user_id at the one place events are written, so no emit site changed, and GET /tenant/security-events accepts ?endUserId= as an indexed equality. The migration backfills existing rows with the same rule. The explicit subject wins over the actor when both are present; in recorded data they never disagreed. A JSON expression index was the alternative and was rejected because Prisma cannot express one, so it would have lived only in migration SQL and read as drift on every later migrate.

The GDPR export had the same hole: it filtered on the end-user as actor, so a subject-access request omitted everything operators and the system did to the person. It now reads by subject. The acting operator's id is still not exported, since it is staff data rather than the subject's.

Log retention and an optional archive

security_events, email_logs and webhook_deliveries only ever grew. LOG_RETENTION_DAYS (default 30) prunes them on the existing ten-minute prune timer in bounded batches.

The rules are narrower than "delete old rows". Deliveries are pruned only when SUCCEEDED or FAILED and age by updated_at, because a pending delivery is live retry state and a failed one an operator redelivered yesterday is not stale just because it was created last month; the delete re-applies the predicate so a redeliver that lands mid-sweep wins. usage_records is deliberately out of scope: it looks like a log but is a billing ledger whose idempotency keys are what make a replayed usage event a no-op.

With LOG_ARCHIVE_S3_* set, each batch is written to any S3-compatible store (R2 or S3) as gzipped NDJSON under Hive-style date partitions before it is deleted, and a failed upload leaves the batch in place for the next sweep. The settings are all or nothing and a partial set refuses to boot, because a bucket configured without keys must not quietly mean pruning without archiving. Object keys are a digest of the batch's ids, so a retried upload overwrites its own object, and uploads carry Content-MD5 so they work against Object Lock buckets. Signing uses aws4fetch rather than the AWS SDK, since one PUT is the whole surface. docs/data-erasure.md now records that the audit trail is retained but bounded, and that erasure does not reach an archive copy.

Panel freshness and speed

Mutations landed and the page did not change. Server actions end in redirect('/...?done=1'), and the Router Cache is keyed by segment path rather than search params, so the redirect resolved to the entry the operator already had: the flash rendered, the data did not move. Five of 45 action files called revalidatePath; the end-user surface called it nowhere. api() now invalidates the tree after every non-GET, which covers every mutation including ones not written yet.

Getting the cache settings right took two wrong turns, which the panel README now records so nobody repeats them. staleTimes.dynamic: 0, which is also Next's default, does not mean fresh. The router still issues every viewport and hover prefetch and discards the result, so a client navigation is slower than opening the same URL in a new tab; one tab click measured three to four full server renders about a second apart. static: 0 additionally re-fetched the shared layout on every click. The settings are now { dynamic: 30, static: 180 }; neither window can show an operator a stale write, because the write invalidates.

Smaller fixes

The self-host compose file was missing the four token-lifetime settings from its environment allowlist, so setting them in a hosting dashboard did nothing and the compose-env CI check was failing on main. .env.bak* is now ignored.

Testing

The full API suite ran on the tip of this branch: 1907 of 1908 cases passed across 182 files. The one failure was a setup step in end-user-destructive-role-floors.test.ts, where a workspace sign-up returned 500 after 7.7 seconds late in a 32-minute run. The same file passes 7 of 7 on its own against the same tree, so it looks load-related rather than caused by these changes, but the server-side error was not captured.

Coverage from the earlier revision includes route classification completeness, cross-tenant isolation probes for every new route group, the scope gate over REST and MCP, the email token-withholding rule, and the import decision table. This revision adds log retention and archive (10 cases), the subject filter (6), and a subject-access export test that now asserts operator actions on the person are included and the operator's id is not. The two new migrations were generated from the schema diff and apply cleanly on a migrated database.

libworky and others added 30 commits September 3, 2026 22:05
…k events

A device becomes a row of its own — (application, end-user, fingerprint) —
instead of a fact about a license. `devices` carries a status (ACTIVE |
RELEASED | BLOCKED), a label and first/last-seen timestamps; unique per
end-user, so the same fingerprint under two accounts is two devices.

`refresh_tokens.device_id` records the device a session was minted on
(nullable, SET NULL). `license_activations` gains `application_id`
(backfilled from the license, then NOT NULL — the one domain table that did
not carry its application), `device_id` and `released_at`, so a seat can be
given back and reactivated in place instead of counting forever.

The devices service owns the ACTIVE count: `touch` registers, refreshes or
reactivates under a per-(application, end-user) advisory lock, enforcing the
`max_devices` FEATURE entitlement resolved through the ordinary plan union.
No entitlement means no cap, so an Application that has not configured one
sees no change. `release`, `block` and `unblock` revoke the device's sessions
in the same transaction.

Five webhook events (device.registered/released/blocked/unblocked/
limit_reached) in both registries, six operator activity-feed entries, the
Device DTOs and OpenAPI components, and docs/devices.md.

Nothing calls `touch` from a request path yet; session binding, the
management routes and the licence-verify integration follow in this series.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…g policy

Every session-minting endpoint (sign-in, sign-up, mfa-verify, refresh, the
OAuth callback, magic-link verify, passkey complete) accepts an optional
`device: { fingerprint, label }`. The device is registered through the
devices service before any token is issued, so a machine over `max_devices`
never gets a session; the access token then carries a `dev` claim, the
refresh chain records the device and carries it across rotations, and
GET /auth/sessions lists it. Clients that send nothing see no change.

All of it sits at the one chokepoint (`issuePair`) every flow already
funnels through, so the limit is enforced in exactly one place.

Refresh refuses a different fingerprint on a bound chain
(REFRESH_TOKEN_DEVICE_MISMATCH, family revoked), binds an unbound chain on
the first refresh that identifies itself, and never gates on
`authConfig.deviceBinding` — that setting (`optional` | `required`) governs
primary sign-in only, so flipping it signs nobody out. It is exposed on the
PATCH route and the update_auth_config MCP tool, which the parity test
enforces.

The error envelope gains an optional `details` object; DEVICE_LIMIT_REACHED
uses it to list the active devices a client can offer to release.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…s; licence seat release

Three surfaces over the devices service: the end-user's own devices at
/users/me/devices (redacted), a secret-key surface at /devices for the
customer's backend, and an operator surface under each end-user with
release, block and unblock. Release revokes every session on the device in
the same transaction; block and unblock are operator-only and the reason
never reaches the end-user.

License seats can be given back: POST /licenses/deactivate on the same
publishable-key surface as verify (deterministic body, idempotent), an
operator activation list and per-activation release, and verify now links
the activation to the holder's Device for the same fingerprint. One new
event, license.deactivated.

SDK: licenses.deactivate, devices.list, devices.release.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
PERPETUAL and TIMED licences accepted unlimited activations ("no upper
bound today"), so one perpetual key could be pasted onto any number of
machines. The holder's max_devices entitlement — the same number that
bounds their sessions — now caps live activations on those kinds, answered
with the existing seats_exhausted. SEATS licences keep seatsAllowed as what
was bought; holders without the entitlement and org-pooled licences are
unchanged. The cap is resolved before the seat transaction and compared
inside it, under the existing FOR UPDATE on the licence row.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A backend that holds a secret key but not the user's token — a licence
server, a support tool, a migration script — had no way to answer "who is
this id / email" or "what does this user get" without minting or holding
that user's JWT.

GET /users?email= (exact, case-insensitive) and GET /users/:id return the
same PublicEndUser shape as /users/me, scoped to the calling Application,
secret key only: an exact-match lookup in a browser's hands is an
enumeration oracle, so the publishable key is refused outright.
GET /billing/entitlements/for-user?endUserId= returns the same union as
/billing/entitlements for the named user.

SDK: users.getByEmail, users.get, billing.getEntitlementsFor.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…-and-rehash

Moving a user base onto Rekey meant a forced password reset for everyone:
EndUser.passwordHash was argon2id only and nothing could create a user with
a hash you already held.

verifyPassword now recognises bcrypt hashes ($2a$/$2b$/$2y$) and checks them
with bcryptjs; hashPassword stays argon2id, and a bcrypt hash that just
verified is re-hashed to argon2id in the same sign-in — best-effort, so a
failed upgrade never turns a correct password into a failed sign-in.

POST /users/import (secret key, auth:write) takes up to 500 users per call
with hash, verification state, role, metadata and OAuth identities.
Existing addresses are skipped, never updated; the batch is validated
before any row is written; the workspace quota is checked per row.

SDK: users.import.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…rints, MCP tools

POST /licenses/verify and /deactivate take the publishable key, so the global
limiter keyed them by IP: an office behind one NAT shared a bucket while a
botnet enumerating a leaked key got a fresh one per exit node. They now
bucket on (application, sha256(key), sha256(fingerprint)) — hashed, because
a raw licence key is a credential and must never sit in a rate-limit store.

A machine fingerprint the person supplied is personal data and survived
GDPR erasure. Devices are now hard-deleted; retained license activations
have the fingerprint tombstoned per row and label/device cleared.

Operator MCP: list_devices, release_device, block_device, unblock_device.
End-user MCP: list_my_devices.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ted hash cost

- Refresh decides the device (mismatch check, limit, block) and the
  verification gates BEFORE rotating the presented token. A refusal no
  longer spends the token, so the client's retry is a plain refresh instead
  of a replay that burned every session.
- deviceBinding=required is checked before an account is created on
  sign-up, first OAuth login and first magic-link login, so a client that
  forgot `device` retries into a sign-up rather than EMAIL_ALREADY_EXISTS.
- Imported password hashes must be a full argon2id PHC string or a bcrypt
  hash at cost 14 or below; verify refuses higher costs too. A tenant could
  otherwise import cost-31 hashes and turn sign-in into a CPU sink.
- Import defaults emailVerified to false, reports OAuth identities it could
  not link (`unlinked`), and treats a concurrent sign-up as `already_exists`.
- Licence verify/deactivate throttle keys on (application, IP) at 60/min
  and fails open when the limiter store is down; the previous per-key bucket
  bounded nothing, since every guess opened a new bucket.
- Server-side release is attributed to `server` (API key id), not the user.
- DSAR export includes devices and licence activations, which erasure
  already treats as personal data.
- Docs and comments corrected: licence verification links a device but
  never registers one; the limit-reached repair path; access-token lifetime
  after release; the errors catalogue gains the device and import codes.
- Tests for the two ordering bugs, the hash ceiling and the new key.
… your own

A fourth provider module, `external`, for sales that happen somewhere Rekey
never calls: the operator's own billing stack, an invoicing tool, a
marketplace, a previous platform mid-migration. That system posts signed
events to POST /api/v1/webhooks/billing/external/<slug> and Rekey activates,
renews, recovers and cancels subscriptions from them, provisions entitlements,
records payments and emits its usual outbound webhooks. A sale for an address
Rekey has not met creates the end-user (no password, default role), recorded
as a security event and announced with user.created.

Provider contract: `capabilities.checkout` (required) marks whether buyers can
be sent to a module; the router, the public provider list and
getProviderForApplication honour it. A new normalized event,
subscription.granted, is the activate-or-create for a sale with no checkout
session; its applier goes through grantSubscription with the sender's
subscription id bound onto the row so the status mirror finds it later.

Signature and envelope mirror the webhooks Rekey sends (X-Rekey-Signature,
t + HMAC-SHA256 over the raw body, five-minute window, eventId idempotency).
Envelope and data are validated before the receipt row, so a malformed event
is a 400 stored nowhere. Cancellation through Rekey is refused for these rows
because the money lives elsewhere.

Panel: inbound-only providers get an ingress card (endpoint, signing recipe,
ping curl, event list) instead of the dashboard webhook flow, routing fields
are hidden, and an Application selling only externally gets a note instead
of the "plans not registered" banner. The MCP configure_billing_provider tool
is now registry-driven and no longer wipes omitted fields on an edit.

Docs: docs/external-billing.md, plus billing, billing-providers,
billing-architecture and errors.
…ebind a hosted row

- A scheduled cancellation no longer invents an ACTIVE status. A
  SubscriptionStatusEvent may now omit `status`, meaning "mirror the
  timestamps only"; a PAST_DUE row keeps its status and its dunning case
  when a future effectiveAt arrives.
- An activation for a subscriber whose live subscription to the plan was
  created by a hosted provider is recorded on the row (metadata.refusedGrants)
  and otherwise ignored, instead of rebinding the row and orphaning that
  provider's events.
- The period end only ever moves forward on a live row; an older-but-future
  value or an explicit null changes nothing.
- The organization preconditions are checked before the subscriber is
  created or a previous holder of the subscription id is retired.
- grantSubscription rethrows a P2002 on (applicationId, providerSubId)
  instead of reporting it as the benign idempotent race.
- An erased end-user named by id is refused (END_USER_ERASED).
- Plan readiness skips the trial check for an inbound-only module.
- Tests for each; the registry test pins `refunds` on every hosted module
  again; provider-refunds covers the external outbound class.
- Docs: lapsed periods, scheduled cancellations, hosted rows, applier retries.
…externally managed subscriptions

- release, block and unblock now take the same per-user advisory lock that
  touch holds and re-read the row under it, so a block cannot be overwritten
  by a concurrent reactivation and a release cannot be lost to a touch that
  read a moment earlier.
- The licence-activation operator route imports the licences service
  statically.
- The create-time billingProvider hint enumerates hosted providers from the
  registry instead of a hand-written list.
- Portal: a subscription created by the inbound-only provider shows where it
  is managed instead of a cancel button that would be refused, and the 409 is
  mapped to readable copy.
Security
- Imported argon2id hashes are bounded (m <= 256 MiB, t <= 10, p <= 8) and
  bcrypt cost is capped at 12; verify refuses hashes outside the budget. An
  unbounded import let one tenant make sign-in allocate gigabytes per attempt.
- Events from a billing provider act only on subscriptions that provider
  created: the granted applier refuses to retire a hosted row whose id a
  sender reused, and the status and payment appliers ignore rows another
  provider owns. Hosted ids are printed on invoices; they are not secrets.
- Inbound webhook receipts are pruned after WEBHOOK_EVENT_RETENTION_DAYS
  (default 90); nothing bounded the table before, and an external sender is
  arbitrary software.
- The external module looks up event types with Object.hasOwn.
- A pipeline retry records the body it actually applies.

Correctness
- A refresh or organization switch on a session whose device was released
  or blocked is refused (SESSION_DEVICE_RELEASED / DEVICE_BLOCKED) instead
  of walking the device back to ACTIVE through touch.
- The granted applier provisions on every pass and announces after the
  entitlements exist, so a retry after a failed provision completes it.
- A late payment.succeeded no longer clears a scheduled external
  cancellation; the "paid since cancellation" heuristic is for processors.
- A timestamp-only status event on a subscription that has ended is ignored.
- Erasure and hard delete end an inbound-only provider's subscription
  locally (the sender learns from user.erased/deleted) instead of failing on
  a provider that cannot be asked; tombstoned licence activations release
  their seats, and activations of the person's machines on org-pooled
  licences are tombstoned too.
- max_devices declared as a numeric STRING is honoured as the number.
- Licence routes no longer carry the fail-closed auth ceiling, so they fail
  open as documented.
- pickProvider names BILLING_PROVIDER_INBOUND_ONLY when a caller asks for
  the external provider explicitly.

Parity
- license_activations.application_id is a relation with a cascading FK
  (new migration).
- SDK: device on verifyMagicLink, verifyPasskeyAuthentication and
  completeOAuth; ImportUsersResult.unlinked; README sections for devices,
  users and deactivate.
- user.created carries via: "import" like the billing-created case.
- Docs: devices (re-mint rule, limit 0, deactivate by a co-holder, no
  ceiling), webhooks (count, via, server), errors, billing readiness codes,
  external-billing (one-time plans, ownership, retention), mcp tool tables,
  README rows, migration lock-profile note. @types/bcryptjs dropped (bcryptjs
  ships its own types).
Security
- OIDC discovery, token and userinfo fetches connect only to the addresses
  the SSRF guard approved, through the same pinned dispatcher the webhook
  sender uses (now shared from lib/ssrf-guard.ts). Resolve-then-refetch left
  a rebinding window on issuer URLs an operator controls.
- OIDC ID tokens are checked for a non-empty sub, the configured issuer,
  this client id in aud and a future exp before their subject names an
  account (OAUTH_ID_TOKEN_INVALID).
- End-user sign-in verifies against a decoy hash for unknown addresses and
  password-less accounts, so response time no longer says which accounts
  exist; the operator sign-in already did this.
- The 16 KB metadata ceiling now applies to organization create and update
  (end-user reachable) and to plan and coupon writers, from one lib module.
- A personal access token minted by an operator since demoted to MEMBER no
  longer lists every Application in the workspace.

Quality
- Licence-activation operator routes live in the licences module.
- One shared "end-user belongs to this Application" guard (lib/end-users.ts)
  replaces three copies.
- release/block return only the documented fields; docblocks sit on the
  functions they describe; comments state facts rather than change history;
  the external verifier refuses an empty secret outright; the licence
  verify link is described accurately; one licence rate-limit config.
- Route docs name SESSION_DEVICE_RELEASED where the path raises it.
- Tests: shared device fixtures (test/device-fixtures.ts), delivery polling
  instead of fixed sleeps, a precise cross-tenant assertion, typed licence
  issue helper, headers that describe the suite rather than a PR sequence.
Second review round over the security pass, plus the gaps an outside
reviewer would hit first.

OIDC: the ID token `iss` check compared the token against `doc.issuer`
and nothing checked `doc.issuer` against the configured issuer, so the
check proved nothing a hostile discovery document could not also supply.
Discovery now refuses a document that names another issuer, resolving
Microsoft's `{tenantid}` template against the token's own `tid` so a
multi-tenant configuration still pins exactly. `azp` is required when
the token carries several audiences, `exp` gets a minute of skew and is
accepted as a string, and the empty-`sub` identity fold is closed on the
userinfo path as well as the ID token path.

`pinnedFetchInit` handed back an agent nothing could destroy, so undici
held each socket for as long as the remote asked, up to ten minutes: one
per webhook delivery, three per OIDC sign-in. It now drops the socket
with the response.

The metadata ceiling reached some writers and not others while claiming
to reach all of them; usage records, credit drawdowns and licence create
are the ones it missed, two of them append-only and the highest-volume
tables in the schema.

Also: `TENANT_ROLE_INSUFFICIENT` documented on the route that raises it,
the organization metadata assert moved out of the Prisma catch, the
`ARCHITECTURE.md` module tree brought in line with the directories that
exist, `PROVIDER_INBOUND_ONLY` documented next to its near-twin, and
`WEBHOOK_EVENT_RETENTION_DAYS` added to the production compose file and
`.env.example`, which the config guard had been failing on since the
external provider landed.

Two test bugs that fail for anyone off UTC or on Windows are fixed with
them: a row aged through `$executeRaw` with a JS Date lands one session
offset in the future, and a path compared against a hardcoded posix
string silently disabled its own exclusion list.

Tests: 13 new cases over the OIDC claim checks, which had none. Full
apps/api suite green, 168 files, 1779 cases.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…k events

A device becomes a row of its own — (application, end-user, fingerprint) —
instead of a fact about a license. `devices` carries a status (ACTIVE |
RELEASED | BLOCKED), a label and first/last-seen timestamps; unique per
end-user, so the same fingerprint under two accounts is two devices.

`refresh_tokens.device_id` records the device a session was minted on
(nullable, SET NULL). `license_activations` gains `application_id`
(backfilled from the license, then NOT NULL — the one domain table that did
not carry its application), `device_id` and `released_at`, so a seat can be
given back and reactivated in place instead of counting forever.

The devices service owns the ACTIVE count: `touch` registers, refreshes or
reactivates under a per-(application, end-user) advisory lock, enforcing the
`max_devices` FEATURE entitlement resolved through the ordinary plan union.
No entitlement means no cap, so an Application that has not configured one
sees no change. `release`, `block` and `unblock` revoke the device's sessions
in the same transaction.

Five webhook events (device.registered/released/blocked/unblocked/
limit_reached) in both registries, six operator activity-feed entries, the
Device DTOs and OpenAPI components, and docs/devices.md.

Nothing calls `touch` from a request path yet; session binding, the
management routes and the licence-verify integration follow in this series.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…g policy

Every session-minting endpoint (sign-in, sign-up, mfa-verify, refresh, the
OAuth callback, magic-link verify, passkey complete) accepts an optional
`device: { fingerprint, label }`. The device is registered through the
devices service before any token is issued, so a machine over `max_devices`
never gets a session; the access token then carries a `dev` claim, the
refresh chain records the device and carries it across rotations, and
GET /auth/sessions lists it. Clients that send nothing see no change.

All of it sits at the one chokepoint (`issuePair`) every flow already
funnels through, so the limit is enforced in exactly one place.

Refresh refuses a different fingerprint on a bound chain
(REFRESH_TOKEN_DEVICE_MISMATCH, family revoked), binds an unbound chain on
the first refresh that identifies itself, and never gates on
`authConfig.deviceBinding` — that setting (`optional` | `required`) governs
primary sign-in only, so flipping it signs nobody out. It is exposed on the
PATCH route and the update_auth_config MCP tool, which the parity test
enforces.

The error envelope gains an optional `details` object; DEVICE_LIMIT_REACHED
uses it to list the active devices a client can offer to release.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…s; licence seat release

Three surfaces over the devices service: the end-user's own devices at
/users/me/devices (redacted), a secret-key surface at /devices for the
customer's backend, and an operator surface under each end-user with
release, block and unblock. Release revokes every session on the device in
the same transaction; block and unblock are operator-only and the reason
never reaches the end-user.

License seats can be given back: POST /licenses/deactivate on the same
publishable-key surface as verify (deterministic body, idempotent), an
operator activation list and per-activation release, and verify now links
the activation to the holder's Device for the same fingerprint. One new
event, license.deactivated.

SDK: licenses.deactivate, devices.list, devices.release.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
PERPETUAL and TIMED licences accepted unlimited activations ("no upper
bound today"), so one perpetual key could be pasted onto any number of
machines. The holder's max_devices entitlement — the same number that
bounds their sessions — now caps live activations on those kinds, answered
with the existing seats_exhausted. SEATS licences keep seatsAllowed as what
was bought; holders without the entitlement and org-pooled licences are
unchanged. The cap is resolved before the seat transaction and compared
inside it, under the existing FOR UPDATE on the licence row.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A backend that holds a secret key but not the user's token — a licence
server, a support tool, a migration script — had no way to answer "who is
this id / email" or "what does this user get" without minting or holding
that user's JWT.

GET /users?email= (exact, case-insensitive) and GET /users/:id return the
same PublicEndUser shape as /users/me, scoped to the calling Application,
secret key only: an exact-match lookup in a browser's hands is an
enumeration oracle, so the publishable key is refused outright.
GET /billing/entitlements/for-user?endUserId= returns the same union as
/billing/entitlements for the named user.

SDK: users.getByEmail, users.get, billing.getEntitlementsFor.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…-and-rehash

Moving a user base onto Rekey meant a forced password reset for everyone:
EndUser.passwordHash was argon2id only and nothing could create a user with
a hash you already held.

verifyPassword now recognises bcrypt hashes ($2a$/$2b$/$2y$) and checks them
with bcryptjs; hashPassword stays argon2id, and a bcrypt hash that just
verified is re-hashed to argon2id in the same sign-in — best-effort, so a
failed upgrade never turns a correct password into a failed sign-in.

POST /users/import (secret key, auth:write) takes up to 500 users per call
with hash, verification state, role, metadata and OAuth identities.
Existing addresses are skipped, never updated; the batch is validated
before any row is written; the workspace quota is checked per row.

SDK: users.import.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…rints, MCP tools

POST /licenses/verify and /deactivate take the publishable key, so the global
limiter keyed them by IP: an office behind one NAT shared a bucket while a
botnet enumerating a leaked key got a fresh one per exit node. They now
bucket on (application, sha256(key), sha256(fingerprint)) — hashed, because
a raw licence key is a credential and must never sit in a rate-limit store.

A machine fingerprint the person supplied is personal data and survived
GDPR erasure. Devices are now hard-deleted; retained license activations
have the fingerprint tombstoned per row and label/device cleared.

Operator MCP: list_devices, release_device, block_device, unblock_device.
End-user MCP: list_my_devices.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
- Refresh decides the device (mismatch check, limit, block) and the
  verification gates BEFORE rotating the presented token. A refusal no
  longer spends the token, so the client's retry is a plain refresh instead
  of a replay that burned every session.
- deviceBinding=required is checked before an account is created on
  sign-up, first OAuth login and first magic-link login, so a client that
  forgot `device` retries into a sign-up rather than EMAIL_ALREADY_EXISTS.
- Imported password hashes must be a full argon2id PHC string or a bcrypt
  hash at cost 14 or below; verify refuses higher costs too. A tenant could
  otherwise import cost-31 hashes and turn sign-in into a CPU sink.
- Import defaults emailVerified to false, reports OAuth identities it could
  not link (`unlinked`), and treats a concurrent sign-up as `already_exists`.
- Licence verify/deactivate throttle keys on (application, IP) at 60/min
  and fails open when the limiter store is down; the previous per-key bucket
  bounded nothing, since every guess opened a new bucket.
- Server-side release is attributed to `server` (API key id), not the user.
- DSAR export includes devices and licence activations, which erasure
  already treats as personal data.
- Docs and comments corrected: licence verification links a device but
  never registers one; the limit-reached repair path; access-token lifetime
  after release; the errors catalogue gains the device and import codes.
- Tests for the two ordering bugs, the hash ceiling and the new key.
… your own

A fourth provider module, `external`, for sales that happen somewhere Rekey
never calls: the operator's own billing stack, an invoicing tool, a
marketplace, a previous platform mid-migration. That system posts signed
events to POST /api/v1/webhooks/billing/external/<slug> and Rekey activates,
renews, recovers and cancels subscriptions from them, provisions entitlements,
records payments and emits its usual outbound webhooks. A sale for an address
Rekey has not met creates the end-user (no password, default role), recorded
as a security event and announced with user.created.

Provider contract: `capabilities.checkout` (required) marks whether buyers can
be sent to a module; the router, the public provider list and
getProviderForApplication honour it. A new normalized event,
subscription.granted, is the activate-or-create for a sale with no checkout
session; its applier goes through grantSubscription with the sender's
subscription id bound onto the row so the status mirror finds it later.

Signature and envelope mirror the webhooks Rekey sends (X-Rekey-Signature,
t + HMAC-SHA256 over the raw body, five-minute window, eventId idempotency).
Envelope and data are validated before the receipt row, so a malformed event
is a 400 stored nowhere. Cancellation through Rekey is refused for these rows
because the money lives elsewhere.

Panel: inbound-only providers get an ingress card (endpoint, signing recipe,
ping curl, event list) instead of the dashboard webhook flow, routing fields
are hidden, and an Application selling only externally gets a note instead
of the "plans not registered" banner. The MCP configure_billing_provider tool
is now registry-driven and no longer wipes omitted fields on an edit.

Docs: docs/external-billing.md, plus billing, billing-providers,
billing-architecture and errors.
- A scheduled cancellation no longer invents an ACTIVE status. A
  SubscriptionStatusEvent may now omit `status`, meaning "mirror the
  timestamps only"; a PAST_DUE row keeps its status and its dunning case
  when a future effectiveAt arrives.
- An activation for a subscriber whose live subscription to the plan was
  created by a hosted provider is recorded on the row (metadata.refusedGrants)
  and otherwise ignored, instead of rebinding the row and orphaning that
  provider's events.
- The period end only ever moves forward on a live row; an older-but-future
  value or an explicit null changes nothing.
- The organization preconditions are checked before the subscriber is
  created or a previous holder of the subscription id is retired.
- grantSubscription rethrows a P2002 on (applicationId, providerSubId)
  instead of reporting it as the benign idempotent race.
- An erased end-user named by id is refused (END_USER_ERASED).
- Plan readiness skips the trial check for an inbound-only module.
- Tests for each; the registry test pins `refunds` on every hosted module
  again; provider-refunds covers the external outbound class.
- Docs: lapsed periods, scheduled cancellations, hosted rows, applier retries.
First-class devices for the auth and billing layers: device model and limits, device-bound sessions, management routes, licence seat release, secret-key lookups, user import, hardening and MCP tools.
Inbound-only external billing provider: a billing system of the operator's own posts signed subscription and payment events; Rekey activates, renews and cancels from them and creates unknown subscribers.
…externally managed subscriptions

- release, block and unblock now take the same per-user advisory lock that
  touch holds and re-read the row under it, so a block cannot be overwritten
  by a concurrent reactivation and a release cannot be lost to a touch that
  read a moment earlier.
- The licence-activation operator route imports the licences service
  statically.
- The create-time billingProvider hint enumerates hosted providers from the
  registry instead of a hand-written list.
- Portal: a subscription created by the inbound-only provider shows where it
  is managed instead of a cancel button that would be refused, and the 409 is
  mapped to readable copy.
Security
- Imported argon2id hashes are bounded (m <= 256 MiB, t <= 10, p <= 8) and
  bcrypt cost is capped at 12; verify refuses hashes outside the budget. An
  unbounded import let one tenant make sign-in allocate gigabytes per attempt.
- Events from a billing provider act only on subscriptions that provider
  created: the granted applier refuses to retire a hosted row whose id a
  sender reused, and the status and payment appliers ignore rows another
  provider owns. Hosted ids are printed on invoices; they are not secrets.
- Inbound webhook receipts are pruned after WEBHOOK_EVENT_RETENTION_DAYS
  (default 90); nothing bounded the table before, and an external sender is
  arbitrary software.
- The external module looks up event types with Object.hasOwn.
- A pipeline retry records the body it actually applies.

Correctness
- A refresh or organization switch on a session whose device was released
  or blocked is refused (SESSION_DEVICE_RELEASED / DEVICE_BLOCKED) instead
  of walking the device back to ACTIVE through touch.
- The granted applier provisions on every pass and announces after the
  entitlements exist, so a retry after a failed provision completes it.
- A late payment.succeeded no longer clears a scheduled external
  cancellation; the "paid since cancellation" heuristic is for processors.
- A timestamp-only status event on a subscription that has ended is ignored.
- Erasure and hard delete end an inbound-only provider's subscription
  locally (the sender learns from user.erased/deleted) instead of failing on
  a provider that cannot be asked; tombstoned licence activations release
  their seats, and activations of the person's machines on org-pooled
  licences are tombstoned too.
- max_devices declared as a numeric STRING is honoured as the number.
- Licence routes no longer carry the fail-closed auth ceiling, so they fail
  open as documented.
- pickProvider names BILLING_PROVIDER_INBOUND_ONLY when a caller asks for
  the external provider explicitly.

Parity
- license_activations.application_id is a relation with a cascading FK
  (new migration).
- SDK: device on verifyMagicLink, verifyPasskeyAuthentication and
  completeOAuth; ImportUsersResult.unlinked; README sections for devices,
  users and deactivate.
- user.created carries via: "import" like the billing-created case.
- Docs: devices (re-mint rule, limit 0, deactivate by a co-holder, no
  ceiling), webhooks (count, via, server), errors, billing readiness codes,
  external-billing (one-time plans, ownership, retention), mcp tool tables,
  README rows, migration lock-profile note. @types/bcryptjs dropped (bcryptjs
  ships its own types).
libworky and others added 23 commits September 7, 2026 20:24
Document scopes and the verified defaults for a new member
Four lifetimes were hard-coded: the end-user and operator access tokens at
15 minutes, and both refresh tokens at 30 days. A desktop application that
opens once a month wants a longer refresh window, and an operator on a shift
on a trusted machine wants a longer access token. Both are now environment
settings with the old values as defaults and ceilings of 24 hours, 12 hours
and 365 days.

Refresh tokens are sliding, so the window only has to cover the longest gap
between uses; the access lifetime bounds how long a revoked session keeps
acting before its next renewal re-checks it, which is why the defaults stay
short. The settings, the env example and both auth docs say so.

The panel's session cookies followed constants of their own. They now take
their lifetimes from the expiries every auth response already returns, with
the token's own expiry as the fallback, so a longer configured lifetime is
not cut short by a cookie.
…erywhere

A longer access lifetime must not extend a session somebody ended, and until
now nothing could end one early: password change, sign-out everywhere,
single-session revoke and device release only revoked refresh tokens, which
was fine at fifteen minutes and is not at eight hours. Both user models gain
sessionsInvalidBefore; every one of those paths stamps it, and the two
session middlewares refuse an access token whose iat falls before the stamp
on its next use. Compared at second granularity so the pair a password
change hands back survives; other sessions renew silently from their live
refresh tokens, so the stamp being per user costs them one transparent
renewal, and the revoked session, whose refresh is gone, stops.

The panel's magic-link, OAuth and cloud sign-in routes and the end-user
portal still set cookies from constants of their own, so a longer refresh
lifetime was cut to thirty days on those paths. They take the lifetimes from
the auth response now, through one shared helper, and the response type
requires the expiries so a call site that lacks them fails to compile.

Every statement of "15 minutes" or "30 days" that would now disagree with a
configured value is reworded, in the OpenAPI descriptions, the docs and the
code comments. The lifetime test passes the real third argument, pins both
refresh lifetimes, and proves the kill switch over HTTP.
… SDKs

The browser SDK's user read, GET /auth/me, resolved the user without the
stamp check, so a revoked token could still read the holder's record there.
It takes the same door now, through a shared helper in lib/session-stamp.ts
that both session middlewares use as well.

The end-user password reset revoked refresh tokens inside its own
transaction without stamping, which is the compromise-recovery path where
the attacker's access token most needs to stop; it stamps. The operator
single-session revoke stamps too, mirroring the end-user one.

The stamp was leaving the server on every user DTO, since redaction only
stripped the password hash. Both public types omit it now.

The refusal reuses USER_TOKEN_INVALID with a message naming the cause,
because that is the code shipped SDKs already treat as "refresh me": a
session that was not ended renews silently, and only the ended one, whose
refresh is gone, lands on sign-in. The panel's password change signs the
operator out cleanly with a banner instead of surfacing as "expired".

Three tests reused a sibling session's pre-stamp token; they renew it first,
as a real client does. A new test proves the end-user kill switch over HTTP
on both session routes and that the stamp never appears in a response.
…fail closed

The org-switch test's outcome depended on whether a second boundary fell
between sign-in and release: same second, the device check answered; a
boundary crossed, the stamp check answered first. It now steps past the
sign-in's second and asserts the stamp over HTTP, and proves the device
check through the service, where nothing sits in front of it.

Passkey sign-in had its own redaction that stripped only the password hash,
so the session stamp was leaving in that one response; it strips both now.
A token with no issue time is refused once a stamp exists rather than
admitted, and the panel's unreachable password-changed banner is gone.
Make token lifetimes deployment settings, with a kill switch for long access tokens
The four token-lifetime settings landed in env.ts but never reached the
compose environment block, so a deployment that set any of them in its
hosting dashboard got the hard-coded defaults and no indication why. This is
the exact failure the block's own header documents, and the CI check catches
it - it was failing on main.
Let the self-host compose pass the token lifetimes through
…ments

Mutations landed in the API and then appeared to do nothing. An action ends
with `redirect('/...?granted=1')`, the Router Cache is keyed by segment path
rather than by search params, so the redirect resolves to the entry the
operator already had: the flash message rendered from the query string while
the data underneath did not move. It got worse the longer a session ran,
because more segments were cached.

Of the 45 files with `'use server'`, five called `revalidatePath`. The
end-user surface - credits, erase, impersonate, support flags, session revoke,
device actions, subscription grant and cancel - called it nowhere, which is
why the end-user and subscription tabs were where this showed up.

Fixed centrally in `api()` rather than at the call sites. Every mutation goes
through it, so every mutation is covered including the ones not written yet;
the alternative was around 250 call sites and a rule that has to be remembered
forever. GET is excluded because that is the render path and `revalidatePath`
throws during render.

Also turns the Router Cache off with `staleTimes: { dynamic: 0, static: 0 }`.
Invalidating on write only helps when the operator caused the change; it does
nothing when a second operator, a webhook or an end-user did. Next's defaults
still served a revisited segment from a payload fetched minutes earlier, and a
support console showing a stale device list is worse than one that takes an
extra moment. This console has no anonymous traffic and every page is already
dynamic and no-store, so the cache was saving little and costing a lot.

Note that the extra reads this produces are sized against the API's
RATE_LIMIT_MAX; the two settings have to move together.
…ions

Invalidate the router cache after a write, and stop reusing stale segments
`staleTimes.static: 0` made every tab click re-render and re-fetch the shared
layout as well as the page. On the end-user screen that shell is the identity
header and the tab strip, and its overview and security tabs each pull three
200-row security-events scans to show twenty rows, so the change roughly
doubled the work per click and took the tab strip's instant render with it.
Measured: the same browsing went from about 50 requests a minute to 210.

`dynamic: 0` stays, so a revisited tab never renders stale page data, and
`lib/api.ts` still invalidates the whole tree on every write, so an operator
action drops the layout immediately. 180 seconds of shell reuse only means
moving between tabs of one record does not re-fetch that record's header each
time.

The real cost underneath is the three 200-row scans, which exist because
GET /tenant/security-events has no actorId filter and the panel narrows in
memory. That wants fixing in the API.
…ions

Give the shell back its cache; keep page data uncached
`staleTimes.dynamic: 0` is Next's default and it is harmful for this console.
The router still issues every prefetch a <Link> in the viewport asks for, then
throws the result away, so the click starts from nothing and a client-side
navigation ends up slower than opening the same URL in a new tab. That was the
reported symptom exactly: the end-user overview and subscriptions tabs hung on
a click and loaded instantly in a fresh tab.

Measured on the bench: with dynamic 0, one tab click rendered the page three to
four times about a second apart - viewport prefetch, hover prefetch, then the
real navigation - each a full server render, none reusable. Every individual
API call was 5-32ms, so nothing was slow; the same work was simply done three
times and discarded twice.

Neither window can show an operator a stale write. `lib/api.ts` calls
`revalidatePath('/', 'layout')` after every non-GET, so a mutation drops the
whole tree. These windows only cover changes somebody else made.

Documents all of it in the panel README, including the tell (fresh tab fast,
click slow), the rule that mutations must go through `api()`, and the three
200-row security-events scans the end-user tabs pay per render.
Let prefetch actually work, and write down why zero is the wrong value
The panel's end-user screen asks one question - what happened to this person -
and the log stored the answer in two columns. Their own events name them in
actor_id; anything done TO them by an operator or the system names them in
metadata.endUserId with somebody else as the actor.

Nothing could express "either", so the panel read the application's latest 200
events three times over, once per actor type, and matched both fields in
memory. That is 600 rows fetched to render twenty, on every view of the
overview and security tabs, and it was wrong for a quiet user on a busy
application: their events simply were not in anyone's most recent 200.

recordSecurityEvent now derives subject_end_user_id once, at the single place
an event is written, so no emit site changes. GET /tenant/security-events takes
?endUserId= and the read is an indexed equality against
(application_id, subject_end_user_id, created_at). The migration backfills with
the same rule, so existing history is reachable and not just events written
from now on.

The explicit subject wins over the actor. The two have never disagreed in
recorded data - checked against every row on the bench - and if they ever did,
an event is about whoever it says it is about.

A JSONB expression index was the alternative and was rejected: Prisma cannot
express one, so it would live only in migration SQL and read as drift on every
subsequent migrate.
A `.env.bak-preflight` sitting in the repo root is one `git add -A` away from
being committed with every secret in it. Nearly was.
Filter security events by the end-user they are about
…/R2 first

security_events, email_logs and webhook_deliveries only ever grew; nothing had
deleted a row from any of them. LOG_RETENTION_DAYS (default 30) prunes them on
the existing 10-minute timer in bounded batches, and the optional
LOG_ARCHIVE_S3_* settings write each batch to Cloudflare R2 or AWS S3 before it
is deleted. A failed upload leaves the batch in place for the next sweep.

usage_records is deliberately out of scope. It is a billing ledger: its
idempotency keys are what make a replayed usage event a no-op, and period
totals are summed from it.

Deliveries age by updated_at and only SUCCEEDED/FAILED are pruned. PENDING is
live retry state, and a FAILED delivery an operator redelivered yesterday is
not stale because it was first created last month. The delete re-applies the
predicate, so a redeliver that lands mid-sweep wins.

The archive config is all-or-nothing and a partial set refuses to boot: a
bucket without keys must not quietly mean pruning without archiving. Keys are
Hive-style date partitions named by a digest of the batch's ids, so a retried
upload overwrites its own object. Uploads carry Content-MD5 so they work
against Object Lock buckets. Signed with aws4fetch rather than the AWS SDK.

Adds created_at / (status, updated_at) indexes so the sweep is a range scan
rather than a sequential scan per tick, and marks SecurityEvent as retained
but bounded in docs/data-erasure.md, including that erasure does not reach
the archive.
Prune the log tables after 30 days, and optionally archive them to S3/R2 first
The GDPR subject-access export read security events with
actorType=end_user and actorId=the user, which returns only the person's own
actions. Everything done TO them - an operator blocking their device, the
billing webhook creating their account - is recorded with somebody else as
the actor, and the export silently left all of it out.

It now reads by subjectEndUserId, derived at write time and indexed. The
select is unchanged: the acting operator's id is staff data, not the
subject's, and the response shape does not move.

The existing test seeded its event with prisma.securityEvent.create, which no
production path uses and which skips the subject derivation. It now seeds
through recordSecurityEvent, and asserts that an operator action about the
subject is exported, another user's event is not, and the operator's id never
appears.
Export everything done to a data subject, not only what they did
@libworky
libworky force-pushed the contrib/rekey-2.2-features branch from 19deaf9 to 55525b2 Compare September 14, 2026 17:59
@Evy04

Evy04 commented Sep 22, 2026

Copy link
Copy Markdown
Contributor

Thank you for this, it was a large and careful piece of work. This repository is generated from our private monorepo on each release, so pull requests here can't be merged directly; instead the whole change was ported there as one commit authored to you. Everything landed: devices and fingerprints, password-hash import (argon2id and bcrypt, rehashed on first sign-in), the external billing provider, the operator console, and operator permissions. A few things changed on the way in: imported and external trials now go through the one-trial-per-buyer ledger, log retention defaults to unset instead of 30 and 90 days so upgrades never prune history, and erasure no longer releases another user's activation on a shared fingerprint. It shipped in 2.2.0-rc.1. Closing this now that it is in the release.

@Evy04 Evy04 closed this Sep 22, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants