Skip to content

fix(ci): stop writing cache no run can read, and land the loop family's spine - #833

Merged
wenzowski merged 10 commits into
mainfrom
claude/cloud-1331-perf-base-arm-6pgmy2
Sep 3, 2026
Merged

wenzowski merged 10 commits into
mainfrom
claude/cloud-1331-perf-base-arm-6pgmy2

Conversation

@wenzowski

@wenzowski wenzowski commented Sep 2, 2026 •

Copy link
Copy Markdown
Contributor

Closes CLOUD-1342.
Closes CLOUD-1343.
Closes CLOUD-1344.
DO-NOT-CLOSE CLOUD-1341
DO-NOT-CLOSE CLOUD-1347
DO-NOT-CLOSE CLOUD-97

Three rows, one branch. The cache row is independent; the other two came out of the same session's own failures.

CLOUD-1342 — the perf cache writes nothing, and that is the fix

The previous design on this branch kept the merge-base SHA in the cache key so that a moved base would miss and therefore save. That is sound about actions/cache in general and wrong here, because it still takes the arm the row was trying to escape: a key engineered to miss is a key engineered to write, and every byte written from a pull request is scoped to refs/pull/N/merge where no other pull request can read it. It optimised for a hit and never priced the miss.

What lands instead:

  • the PR side is actions/cache/restore and writes nothing, so the key drops the merge base — with no save there is no miss to engineer, and a key that never moves is exactly what is wanted;
  • the seed is written by the scheduled perf job on main, which already pays a full --release build. An entry written from the default branch is readable by every branch, which is the property no PR-scoped entry can have. No new trigger, no second build, and AGENTS.md's prohibition on a push-to-main trigger is untouched.

Correctness is unchanged and does not rest on the directory name: the seed is copied to target/perf/base-$SHA with its binary removed, so base_arm_is_built() stays false and no stale arm is ever measured.

The claim is different from the one this PR used to make. It is no longer ~18%: the hit arrives on a pull request's first run rather than its second, and this job now writes zero PR-scoped bytes. CLOUD-840 measured this repository at 225.2 GB of Actions cache, 203.7 GB of it across 1,144 entries nothing can restore.

cache-path-is-rebase-stable gains the two sub-action spellings. actions/cache/restore and .../save derive an entry's version from path exactly as the composite does, so matching only actions/cache@ would have left the predicate live and reaching nothing the moment a job split restore from save — which is what this row does. Its own case is declared rather than assumed.

CLOUD-1343 — the claim order, and the message that invited the second claim

The receipt is minted on claim check's pullable path and keyed by the branch checked out at that moment. Two orderings fail in opposite directions and both are reachable by following the instructions correctly the first time: claim-then-branch strands the receipt on a branch nothing will land, and claiming a second time runs after the row has left Todo, reads it as held, and refuses not-todo — where the holder is the caller, ninety seconds earlier. The only route past that is --takeover against oneself, and CLOUD-1139 needs that signal to stay rare.

No engine change: creating a branch writes only under .git, which claim-needs-receipt never judges, so the deadlock an earlier revision assumed does not exist. claim.rs's not-todo decision is correct for a row genuinely held elsewhere and is untouched.

AGENTS.md carries the order (a substitution — the file sits at 198 of 199 lines and 3404 of 3500 tokens), .claude/rules/toolchain.md carries the reason with both failure directions, and the pullable message now says the receipt is already minted and not to re-run. policy/claim-order-is-stated.rego keeps the prose from evaporating; whether a given session actually claimed before branching is not a property of the tree, and a rule deciding that would be the model verdict rule 3 forbids.

This fix was used to claim CLOUD-1344 later in the same session, and the rewritten message did its job.

CLOUD-1344 — the extraction spine

Four of five Extraction members reached no consumer and repetition had no member at all. This is the plumbing every later predicate stands on, landed alone because the risk is the plumbing rather than the detection.

  • Extraction::of widens from &Counts to &Stream, and answers with an Option whose None is could-not-look.
  • Counts is unchanged — it is the -J capability report's own shape, so runs live in Stream::repeats() beside Stream::counts().
  • The load-bearing half is the per-extraction capability. Zero is a real answer meaning the extractor ran. An extraction whose underlying event kind never appears is not a session that did none of it — it is a host that does not record it, and answering zero is a false green over a session nobody measured. The projection omits the key, so a module reads undefined and Rego takes that as does-not-hold. Decided in the engine: a per-module conjunct asking whether this host records turns is a dead gate on every harness but the one its author tested.
  • agent-turn-run is the first member and deliberately the simplest — a trailing run of assistant turns carrying no tool call, over the already-typed Event::Turn. No hashing, so no argument or result is read even internally.

The bound is stated rather than absorbed: this cannot separate a host that records no hook runs from a session that triggered none, and it resolves that toward could-not-look — the safe direction.

The header keeps the claim narrower than "loop detection". Termination is undecidable and every quantity here is a monotonically growing count, so there is no ranking function to be had; what the literature buys is the shape of the declaration.

Ships at warn. CLOUD-894 owns the firing-rate ceiling and CLOUD-1352 the promotion.

Also retires policy/harness-declared.json's two now-spent exemption rows: CLOUD-1079 has landed, so the hooks they excused are no longer provisioned and harness-wiring reports both — the gate working. Reproduced with --config-from a907565, so the finding predates this diff.

The three keys this PR serves and does not close

CLOUD-1341. Its measured defect — an across-turn poll, 1079 identical calls — needs identical-call-run over call identity, which is CLOUD-1347. agent-turn-run catches monologue, not repeated tool calls, so nothing here fixes it. Back in Backlog and unassigned.

Its own §2 has been corrected on the tracker: it specified "a plain total over the stream", an implementation followed that literally, and a sum across identities is monotonic in session length — it never resets, so it fires on any session long enough to repeat anything. Shipped that way at deny, it refused every subsequent tool call in its authoring session, the push included. Every member of this family is non-monotonic by design, which is the property that makes a later promotion survivable.

CLOUD-1347. The spine this PR lands is exactly what unblocks it, and the commits say so — but it is blockedBy CLOUD-1344 and CLOUD-1345, and only the first is cleared here. CLOUD-1345's whole subject is that per-call fingerprinting turns a latent parse cost into a real one, which is precisely what CLOUD-1347 adds. Not this branch's to clear, and a comment on CLOUD-1345 records the walk-count regression this PR's own &Stream widening introduces (1 walk to 2 × declared extractions).

CLOUD-97. Cited by the stop-posture test fix because it owns unlanded-check, whose finding is what the fixture was accidentally reading. The row itself is untouched.

Verification

mise run verify green: HEAD carries verify + linear-check receipts. policy test 50 bundles / 638 passed. The compiled tier over the shipped module is seven cases including the pair that carries the row — a host recording no turns answers could-not-look, and a recorded session with no run answers a real zero — which needs a probe module because both are silent under a >= 3 predicate.

Two defects were caught by that tier rather than by reading: the first cross-compatibility fixture was built from a user record, which the parser reads as a turn, so it recorded turns after all; and both #MUTANT rows first named test_ rules inside the .rego while #MUTANT-SUITE resolves the compiled file — declared, never applied, counted by nobody, the shape PR #829 measured in policy/harness-wiring.rego.

Two more were caught by verify itself and are fixed here rather than worked around: two workspace lints against the spine (struct_excessive_bools, match_same_arms), and a pre-existing isolation bug where the Stop posture cases read the live findings store rather than their fixture's — so a_clean_final_message_says_nothing passed on a clean checkout and failed on any branch carrying unpushed work, which is every branch that suite is run from mid-development.

Every protected-path write carries an issued Admits: block in its commit message.

🤖 Generated with Claude Code

https://claude.ai/code/session_017o7xUB8okt7TbcymPP667s

@linear-code

linear-code Bot commented Sep 2, 2026 •

Copy link
Copy Markdown
CLOUD-1342 CLOUD-1331's `perf-base` entry can never be read: the merge-base SHA in the cache KEY guarantees a miss, and a PR-scoped entry cannot cross a PR — 4 runs, 0 hits, ~190 MB written each

Why

CLOUD-1331 landed (#825, 79dd7b24) a keyed base arm for perf pair: target/perf/base-<merge-base-sha>/, restored around mise run perf-gate from a cache keyed on that SHA plus hashFiles('mise.lock', 'Cargo.lock'). The engine half works and is measured — with the keyed directory present the base arm compiles zero crates, spawns no cargo, and the pair still measures. Cold it is 239 Compiling and 4m42s.

The CI half cannot collect that, measured 2026-09-02. Every perf-base-* entry in the repository, with its owning ref:

refs/pull/801/merge  perf-base-Linux-75fd0552…  189,695,604 B
refs/pull/825/merge  perf-base-Linux-988a2424…  189,792,549 B

Two independent failures, either one sufficient:

  1. Scoped to a PR merge ref. The perf job is pull_request-only (if: github.event.pull_request.draft == false), so no entry is ever written from main, and a PR merge-ref cache is not readable by another PR. Cross-PR sharing is impossible by construction, not merely unlikely — this is CLOUD-840's thesis on a second cache.
  2. The key moves under a PR that has not changed. PR chore: release v0.0.138 #801 ran perf at 06:58 and again at 15:15. The second looked for perf-base-Linux-75fd0552…-79dd7b24…, got Cache not found, and rebuilt the base arm cold in 5m14s — same PR, same toolchain hash, different merge base, because main advanced to 79dd7b24 and the PR's merge ref followed it. land rebases every lap, so the drive-to-green loop mints a fresh key each time.

Four perf runs across two PRs, zero hits, ~190 MB written per run. CLOUD-1331's §2 predicate — "on a run whose merge base equals a prior run's, the base arm compiles zero crates" — is true as written and vacuous in practice, because in this repository's workflow no two perf runs share a merge base.

The SHA is doing two jobs and only one of them is load-bearing. In the directory name it is what makes a stale arm unmeasurable and must stay. In the cache key it guarantees the miss and buys nothing: the dependency closure depends on the toolchain and Cargo.lock, not on the base commit, and the crate's own artifacts rebuild regardless because the base tree's source differs.

Refinement — Ready

Refinement gate: Definition of Ready & Done. This body carries only specializations.

{
  "source_of_truth": ".github/workflows/ci.yml",
  "gate": { "task": "verify", "exits": [0, 1] },
  "commit_type": "fix",
  "blockers": [],
  "tests": [{
    "file": "crates/batten/tests/it/ci_parity.rs",
    "mutation": "cache-path-may-carry-the-base"
  }]
}
  • Authority boundary (§1). The perf job in .github/workflows/ci.yml — its cache step and two bare steps either side of mise run perf-gate — plus the gate that holds the shape: a predicate in policy/ci-parity.rego and its [[rule]]/[[verdict]] rows in batten.toml. crates/batten/src/perf.rs is NOT edited: the keyed directory and its stale-arm refusal are correct and stay exactly as landed. The gate lives in the consumer's module rather than the core because it names this repository's perf job (non-negotiable rule 1), which is the boundary perf(perf): key the base arm on its merge base and carry it between runs (CLOUD-1331) #825 already tripped over.

  • Computable predicate (§2). The cached path carries no merge-base SHA, so the entry's identity is stable across a rebase; the SHA stays in the key so a moved base misses primarily and therefore SAVES; a restore-keys prefix supplies the previous base's closure. On a perf run whose merge base differs from the cached one but whose Cargo.lock does not, the step reports Cache restored from key: … and the base arm compiles ~64 crates rather than 239.

    **This replaces the mechanism this row originally specified, which cannot work. **actions/cache identifies an entry by key AND version, and its README defines version as "a hash generated for a combination of compression tool used … and the path of directories being cached." Our path is target/perf/base-<merge-base-sha>, so the path alone moves the entry's identity whenever the base moves — dropping the SHA from the KEY is a no-op. Stabilising the path without moving the SHA into the key is worse than today: the primary key would hit, and "if the provided key matches an existing cache, a new cache is not created," so CI would restore a wrong-base directory forever and never store the current one.

    Correctness is preserved by deleting the seeded binary, not by the directory name. The seed step copies base-seed to base-<sha> and rm -fs that copy's release/batten, so base_arm_is_built() is false, cargo runs against the base tree that was actually materialised, and no renamed old binary is ever measured. restore-keys is safe here for the same reason and only here: a prefix hit lands in a seed directory nothing measures. The existing "NO restore-keys" comment in the job is about the MEASURED path, is still true of it, and is rewritten rather than left standing.

    Measured locally 2026-09-02, Cargo.lock byte-identical between the seed commit 6363a2e2 and the built tree, so the residue is the cost of the rename itself and not lockfile churn: cold 239 Compiling / 4m42s; seeded-and-rebuilt 64 / 2m07s; same base, directory present 0 (CLOUD-1331's already-landed path, which CI cannot reach). So the honest win is ~2m35s off a 13.5 min step — about 18%, not the whole problem. The 239→0 path needs a run whose merge base matches a cached one, which needs cross-PR scoping and is CLOUD-840's.

  • Deliberately not in scope (§2). CROSS-PR SHARING, moved out to CLOUD-840. An entry readable by an arbitrary PR needs one written from a ref other PRs can read, and AGENTS.md forbids the trigger that would do it — "Never re-run CI on an already-tested SHA… Don't add push-to-main triggers." The consequence, stated so this row is not read as delivering more than it does: the win is bounded to a re-run within one PR with no intervening rebase, plus the local verify saving CLOUD-1331 already delivered. Also out: removing the SHA from the DIRECTORY — that is the refusal CLOUD-1331 exists for and a stale arm must stay unmeasurable. Changing what is measured, the run count, or the 1.30 threshold. rust-cache's own key: perf entry, which is a different action with different prefix behaviour.

  • Effect (§3). None in this repository's verbs. A workflow trigger and cache key change.

  • Output and exit (§5). Unchanged. perf-gate still emits arm= records and perf-compare still decides the ratio.

  • Commit / bump (§6). Type fix(ci), so patch until 0.1.0.

  • Test obligation (§7). This is a MEASUREMENT, recorded before it is kept: a second perf run within one PR, on a different merge base with an unchanged Cargo.lock, reports a restore rather than Cache not found and shows the base arm compiling ~64 rather than 239. Baseline to beat: 4 runs / 0 hits / ~190 MB written each, and a cold base arm of 239 Compiling in 4m42s.

    The gate half is ci-parity's new predicate, with its #MUTANT cache-path-may-carry-the-base row reddening crates/batten/tests/it/ci_parity.rs — so the path-in-the-version defect cannot be reintroduced by the next author, which is the part a measurement alone does not buy (non-negotiable rule 2).

    The reading instrument itself needs fixing and is part of this row. CLOUD-1331's §7 prescribed … /logs" | grep -c '^.*Compiling ', which returns 0 on any real job log regardless of what compiled: the runner emits ANSI colour and the reset code sits between Compiling and the following space. gh api also refuses to write the body without --allow-escape-sequences. The command that answers is mise exec -- gh api --allow-escape-sequences "…/logs" | sed 's/\x1b\[[0-9;]*m//g' | grep -c 'Compiling', and the per-arm split is the lines between the two Finished \release profile`` markers.

  • Blockers (§8). None. relatedTo CLOUD-1331 (the row that landed the mechanism this makes collectable), CLOUD-840 (the PR-scoping half, where the full measurement is recorded), CLOUD-1225 (the compile term).

Acceptance

  • A perf run whose merge base differs from the cached entry's and whose Cargo.lock does not, restores the closure and compiles ~64 rather than 239 for the base arm.
  • gh api repos/button-inc/batten/actions/caches shows the perf-base-* entry persisting across a rebase within that PR.
  • The perf-gate step wall is recorded against CLOUD-1331's 13.5 min baseline, read with the corrected command above.
  • mise run perf-gate in verify behaves exactly as today on a machine with no cached arm.

Refinement history — the cross-PR half was split out, and this row is what remains.

As first filed, §2 asked for "an entry readable by an arbitrary PR", which requires an entry written from a main-triggered run. AGENTS.md's workflow contract forbids exactly that — "Never re-run CI on an already-tested SHA… Don't add push-to-main triggers" — and perf.yml records the same reasoning for the recording half, which is why the series runs on a clock. So the row was moved to Backlog rather than worked.

That half is now **out of scope and lives on **CLOUD-840, whose whole subject is "no PR can read another's"; the candidates noted here (a scheduled main job that warms the entry; accepting within-PR reuse only) are that row's to evaluate, not this one's.

What remains needs no new trigger: a stable cached path, the SHA in the key, a prefix restore, and a seed step that drops the stale binary. §2, §7 and Acceptance above are rewritten to that scope and no longer presuppose the forbidden trigger.

Second correction, 2026-09-02. The narrowed row still specified the wrong mechanism — "drop the SHA from the key, keep it in the directory name" — which §2 now records as a no-op, because actions/cache derives an entry's version from the path. That was caught by reading the action's README instead of assuming its lookup shape, before any of it was written.

CLOUD-1331 `perf pair` builds the base arm's whole release closure from nothing on every CI run — the base binary is a pure function of the merge base, so the `perf` job pays 13.5 min for a build no run can reuse

Why

The perf job is one of the two wall-clock poles of a pull request's checks, and its cost is almost entirely one step.

**Measured 2026-09-02 from two **ci.yml runs on main-adjacent heads:

run job wall mise run perf-gate step
33586699312 (ee1c9f52) perf 13.7 min 13.5 min
33584118886 perf 13.8 min 13.5 min

For comparison the same runs' bats job was 10.4 / 8.7 min and the ci job 14.5 / 12.7 min. So perf is within a minute of the pole on both, and it is one step.

What the step does, read from the engine rather than the job. crates/batten/src/perf.rs::measure (:561) builds the head arm with build(repo, None, "head") at :582 — cargo build --quiet --release -p batten (build(), :454) into target/release. It then materialises the merge base at target/perf/base-tree (:604-605) and builds it with its own CARGO_TARGET_DIR at target/perf/base-target (:609-610), on a stated reason: sharing the main target dir "would make the two builds evict each other's artifacts on every lap, and would race the target-dir lock against whatever else verify is running." [profile.release] is lto = "thin", so each arm pays a thin-LTO compile of the whole dependency closure plus the crate.

Where the cache reaches and where it does not. .github/workflows/ci.yml:633-636 restores Swatinem/rust-cache under key: perf before mise run perf-gate. That action caches the dependency artifacts of the default target directory and drops workspace crates by design (CLOUD-840's reopen note, CLOUD-1225's cache-workspace-crates: false). So on a hit the HEAD arm still compiles batten under thin LTO, and the BASE arm — under target/perf/base-target, a second target directory — is what this row is about: whether the restore reaches it at all is unmeasured, and if it does not, the base arm compiles its entire closure cold on every run. The Compiling line count per arm in the job log (via the API, never gh run view --log, which CLOUD-1225 records as lossy for cargo output) is the reading that decides it.

The base arm is a pure function of the merge base. Its inputs are the merge-base SHA, the pinned toolchain and [profile.release]. main advances only by fast-forward to already-judged SHAs, so consecutive pull requests share a merge base for hours at a time, and every one of them rebuilds the identical binary. That is the CLOUD-840 class (bytes nothing restores) arriving on the release side.

Why the skip does not save this. CLOUD-875 widened perf pair's skip set to batten.toml and every path a policy row registers, correctly — a config-only change moved wired 5.8ms → 9.3ms while the gate reported nothing measured. The consequence is that most of this repository's traffic now pays the full pair: a retirement bundle touches batten.toml and policy/*.rego by mandate, so the skip fires on documentation changes and little else. The widening stands; the cost it exposed is this row's.

Refinement — Ready (reuse the base arm across runs)

Decided 2026-09-02, so the implementer does not choose: the base arm keeps its OWN target directory, keyed on the merge-base SHA, and is restored WHOLE from a cache keyed on that SHA plus the toolchain hash. Closure sharing between the two arms is NOT taken — the :606-608 eviction and lock concerns are real under verify, and a base that is built once per merge base and then restored by every later run makes sharing unnecessary. A new merge base builds the base arm cold exactly once; that is the accepted cost.

Refinement gate: Definition of Ready & Done. This body carries only specializations.

  • **Authority boundary (§1). **crates/batten/src/perf.rs (the base arm's target directory and the build it runs) and .github/workflows/ci.yml's perf job (what is restored and saved around mise run perf-gate). No mise-tasks/ program and no .bats is edited or added — perf-gate.sh and perf-compare.sh are governed and are CLOUD-1163 unit 10's to retire, not this row's to touch. [profile.release] is untouched.
  • Computable predicate (§2). On a run whose merge base equals a prior run's, the base arm compiles zero crates — the base binary is restored, keyed on the merge-base SHA and the toolchain, and perf.rs accepts it only under a directory named for that SHA so a stale arm from another base can never be measured as this one. On a run with a new merge base, the base arm builds once, cold, into its keyed directory, and that directory is saved for every later run on the same base. The measurement itself — both arms sampled back to back on one machine, the ratio decided by perf-compare — is unchanged.
  • Deliberately not in scope (§2). Narrowing the skip set back (CLOUD-875 decided it). The ci job's debug-profile compile (CLOUD-1225, CLOUD-840). The bats job. Any change to what is measured, how many runs, or the 1.30 threshold. Retiring perf-gate.sh/perf-compare.sh (CLOUD-1163 unit 10).
  • **Effect (§3). **write, as perf pair already is — it builds and writes under target/. The one new write is the restored base directory, which is target/perf/'s own.
  • Output and exit (§5). Unchanged. A restored base arm prints the same arm= records; a base that cannot be restored builds as today and says nothing new. Exit follows the 0/1/2/3 table — a corrupt or mis-keyed restore is could-not-look (2), never a measurement.
  • **Commit / bump (§6). **perf(perf) — no bump; perf releases nothing at any version.
  • **Test obligation (§7). **crates/batten/tests/it/perf_pair.rs is the tier; add the cases there, shown able to fail per CLOUD-418: a base directory keyed to a DIFFERENT SHA is refused and rebuilt (the discriminator — without it a stale arm passes as fresh); a base directory keyed to the right SHA is reused without spawning cargo; the anti-vacuity mirror is a first run with no cached arm building both exactly as before. The CI half is a MEASUREMENT recorded on this row before it is kept: Compiling lines per arm and the perf-gate step wall on a run whose merge base matches its predecessor's, against the 13.5 min above, read from the job log API.
  • Blockers (§8). None. relatedTo CLOUD-875 (the skip this cost sits behind), CLOUD-172 (the gate's own row), CLOUD-840 (the same class on the debug side), CLOUD-1225 (the compile term, and the lossy-log method note), CLOUD-398 (the job graph), CLOUD-1151 (the dispatch it rides in).

Acceptance

  • A second run on an unchanged merge base shows the base arm compiling zero crates in its job log, and the perf-gate step wall recorded against 13.5 min.
  • A run with a new merge base builds the base arm once; the next run on that base restores it and compiles zero crates for it.
  • perf_pair.rs refuses a base directory keyed to another SHA, and that case is shown red before the fix.
  • mise run perf-gate in verify behaves exactly as today on a machine with no cached arm.

Found while grooming the CI-cost dispatch: measured against the runs above, bats was not the pole and this job was.

CLOUD-1343 The prescribed claim order mints the receipt on the wrong branch: claim before code, branch after, and every write is then refused — recovered only by `--takeover` against yourself

Why

AGENTS.md prescribes the order plainly: "In Progress = pulled — claim by hand, before writing code (mise run claim-check) and assign yourself." claim-check mints its receipt under $GIT_DIR/batten-receipts/, keyed by branch — deliberately, because "a claim attests to a decision that every commit on the branch continues to serve, and a SHA-keyed one would demand a re-claim per commit" (.claude/rules/toolchain.md).

Those two are incompatible whenever the feature branch does not exist yet, which is the ordinary case. Measured 2026-09-02 landing CLOUD-1331, in this exact sequence:

  1. mise run claim-check -- --issue CLOUD-1331 on the branch the session started on → pullable, receipt minted, keyed to that branch.
  2. Moved Todo → In Progress and assigned, as instructed.
  3. git checkout -B claude/cloud-1331-perf-base-arm origin/main — the branch the work belongs on.
  4. First edit → receipt read missing claim branch claim-needs-receipt. The receipt is on the previous branch; this one has none.
  5. Re-ran claim-check on the new branch → not-todo (in In Progress) — "not pullable, someone is already on it." The someone is me, from step 2.
  6. The only route through was claim-check --takeover, which "mints the receipt and records what it overrode" — recording that I overrode my own claim, made ninety seconds earlier.

Why each existing row is adjacent and none is this. CLOUD-733 (Done) is a renamed branch stranding its receipt. CLOUD-516 / CLOUD-1091 are a restarted branch inheriting a stale one. CLOUD-1231 is one key per branch against a multi-row PR. This is none of those: nothing was renamed, restarted or multi-row. Following the documented order, first time, correctly, produces the refusal — and the recovery writes a takeover record that CLOUD-786 observes nothing reads and CLOUD-1139 observes never checks whether the holder is alive. Here the holder is the same session, which is the one case a liveness check would answer trivially.

Why it matters beyond the friction. A self-takeover is indistinguishable in the record from taking a row out from under a working sibling — the exact event CLOUD-1139 measured twice in one hour. Making the ordinary path mint one teaches every session that a takeover is routine, which is precisely the signal that row needs to stay rare.

Refinement — Ready

Refinement gate: Definition of Ready & Done. This body carries only specializations.

{
  "source_of_truth": "AGENTS.md",
  "gate": { "task": "verify", "exits": [0, 1] },
  "commit_type": "fix",
  "blockers": [],
  "tests": [{
    "file": "crates/batten/tests/it/claim_order.rs",
    "mutation": "order-may-go-unstated"
  }]
}

This clause was first filed as feedforward with an empty tests, and the Ready gate was right to refuse it. The reasoning ran: the remedy is words, rule 3 forbids gating a judgement, so there is no mutation to name. The first half is true and the second does not follow. Whether a given session claimed before branching is indeed a judgement and is not gated. Whether the always-loaded file still states the order is a byte in a tracked file — a real object with a real exit code — and that is what policy/claim-order-is-stated.rego decides. The failure mode here is DRIFT, since words evaporate, and drift is exactly what a gate can hold.

  • Authority boundary (§1). AGENTS.md's board paragraph (the ORDER only — that file is budgeted, and policy-budget caps it at 3500 tokens and 199 lines, which the first draft of this change blew); .claude/rules/toolchain.md for the reason, which is its triggered home; and the pullable message in crates/batten/src/lib.rs — the line that invites the second claim-check. NOT claim.rs's not-todo decision, which is correct as it stands for a row genuinely held elsewhere. No mise-tasks/ program and no .bats.
  • Computable predicate (§2). AGENTS.md states the order as branch → claim → move the row, and the pullable refusal stops reading as an instruction to come back. A session following that order holds a valid receipt on the branch the work lands on, having invoked claim-check exactly ONCE and written zero takeover records. That is the whole observable, and it needs no engine change: creating a branch writes only under .git, which claim-needs-receipt never judges, so the ordering deadlock the first revision assumed does not exist.
  • Deliberately not in scope (§2). RECOGNISING A SELF-CLAIM IN THE ENGINE, which is CLOUD-1139's. It requires distinguishing this session from a sibling, and the row's assignee cannot do it — every fleet session carries the same configured accountable identity, so an assignee-keyed re-mint would let a sibling silently re-mint over a working holder, which is that row's measured harm reintroduced through this row's fix. Making the receipt SHA-keyed — .claude/rules/toolchain.md records why branch-keying is right and a SHA key would demand a re-claim per commit. Weakening not-todo generally: a row genuinely held by another session must still refuse. CLOUD-1139's liveness question, which is the general form of the check; this row needs only the degenerate case where the holder is the caller.
  • Effect (§3). write — claim check already writes the receipt; this changes when and where, not whether.
  • Output and exit (§5). Unchanged: the 0/1/2/3 table, and the existing refusal text for a row genuinely held elsewhere. A self-claim emits no takeover record, because nothing was overridden.
  • Commit / bump (§6). Type fix(claim), so patch until 0.1.0.
  • **Test obligation (§7). **policy/claim-order-is-stated.rego is the gate, with two declared mutations, and crates/batten/tests/it/claim_order.rs is the compiled-binary tier that proves the ENGINE builds the input.tree.lines shape the module reads — a with input as case cannot, because it fabricates the very shape the engine may be unable to produce. Three arms: the index must state branch-before-claim, the index must ask for exactly ONE claim (branching first is necessary and not sufficient — the second call is what actually bit), and the rules file must carry both failure directions. The case that matters most runs against THIS checkout, so a fixture-only suite cannot stay green while the real instruction surface loses the sentence. Shown able to fail per CLOUD-418 — that case must be red before the fix, which it demonstrably is, since it is the measured sequence above. The discriminator against over-reach: a row claimed by a different branch's session must still refuse on branch B, or the fix has replaced a false refusal with a false pass.
  • Blockers (§8). None. relatedTo CLOUD-1139 (takeover never asks whether the holder is alive — this is the case where the holder is the caller), CLOUD-786 (a takeover record nothing reads), CLOUD-733 (the renamed-branch sibling, Done), CLOUD-516 and CLOUD-1091 (the restarted-branch siblings).

Acceptance

  • The documented order — branch, claim, move the row, write — completes with no refusal and no takeover, and invokes claim-check exactly once.
  • A row held by another session still refuses on a fresh branch.
  • The failing sequence is shown red before the fix.
  • No takeover record is written where the claimant and the holder are the same session.

Found landing CLOUD-1331: step 4 cost a full stop mid-session, and step 6 put a takeover in the record for a row nobody else had touched.


Refined 2026-09-02: the mechanism is chosen, and the row is SPLIT because the two halves have different owners.

The previous revision named two mechanisms and chose neither. Working all three of its open questions through:

Q3 first, because it is the cheapest and the row half-admitted it. The resolution is NOT in the engine. Working the sequence: create the branch, THEN claim-check (which mints on its pullable path), THEN move Todo → In Progress and assign. That completes with no refusal, no --takeover and no code change. Creating a branch writes only under .git, which claim-needs-receipt never judges, so there is no ordering deadlock — the deadlock was imagined.

What actually produced the measured failure is a re-run, and the refusal invites it. claim-check's pullable line reads "claim it before you write code: move Todo -> In Progress and assign yourself", which reads as go-do-this-and-come-back. Coming back is the refusal: the row is now In Progress, so not-todo fires, and the only route is --takeover against yourself. Reproduced again this session on CLOUD-1342, in exactly that shape, having branched first — so branching first is necessary and not sufficient, and the second claim-check is the whole defect.

Q1 and Q2 — the engine half is CLOUD-1139's, on a reason rather than by deferral. Recognising a self-claim needs to distinguish this session from a sibling. It cannot be done from the row's assignee: every session in a fleet carries the same configured accountable identity ([attribution]), so an assignee-keyed re-mint would let a sibling session silently re-mint over a working holder — which is CLOUD-1139's measured harm, reintroduced through the fix for this one. So it genuinely needs session identity, the engine has none, and inventing one here is the second authority the previous revision correctly feared. That half belongs to CLOUD-1139 as its degenerate case and is out of scope here.

So §1 and §2 are rewritten to the ordering half only, and crates/batten/src/claim.rs comes OUT of §1 — the previous §1 named it, and that was the wrong file for the change this row now makes.

  • Authority boundary (§1), revised. AGENTS.md's board paragraph, .claude/rules/toolchain.md, and the pullable message in crates/batten/src/lib.rs (the line that invites the re-run — it is in lib.rs, not refusal.rs, which an earlier draft of this clause named wrongly). NOT claim.rs's not-todo decision, which is correct as it stands for a row genuinely held elsewhere.

    The split between the two prose files is forced by a budget, and that is the budget working. policy-budget caps AGENTS.md at 3500 tokens and 199 lines; the first draft of this change put the whole explanation there and blew both. So the always-loaded file carries only the ORDER an agent needs in hand every turn, and the WHY — both failure directions, the absence of a deadlock, and why the engine does not decide this — lives in the rules file that loads at the trigger.

  • Computable predicate (§2), revised. AGENTS.md states the order as branch → claim → move the row, and the pullable refusal stops reading as an instruction to return. The observable: a session following the stated order writes zero takeover records, and claim-check is invoked exactly once per row.

  • Mechanism (§7), revised, and it is feedforward with a named precedent. This is the REMOVAL of a trap from an instruction, not a new rule, so non-negotiable rule 2 is satisfied the way .claude/rules/commits.md satisfies it — that file records the identical shape: "Feedforward only, and deliberately so — no gate is implied… this is the REMOVAL of a licence." The drift mechanism is a prose assertion in the crates/batten/tests/ tier that AGENTS.md's board paragraph still states branch-before-claim, the same shape as scanner_taxonomy.rs's assertion over .claude/rules/scanning.md. Plus the compiled-binary case already specified: claim on branch A, branch to B, write — red before, green after.

What is not in doubt: the measured sequence in the body is a real defect, it is reachable by following the documented order correctly on first use, and the recovery writes a takeover record for a row nobody else touched. The acceptance bullets stand as written, narrowed to the ordering half; the mechanism is no longer open.

CLOUD-1344 Four of five `Extraction` members reach no consumer, and repetition — the one thing a doom loop is made of — has no member at all

Why

Batten has no doom-loop gate. It has nearly all the substrate for one, unused.

Measured in the tree 2026-09-02:

  • facts.rs ~2649 implements enum Extraction { Turns, ToolCalls, ToolErrors, HookDecisions, HookDenials } over transcript::Counts. Integers only, closed set — non-negotiable rule 4 enforced in the type rather than by review.
  • batten.toml:4886 declares one extractor: id = "denials", count = "hook-denials". The other four are built by the engine on every session and reach no consumer.
  • policy/denials-outlive-the-turn.rego is the only module reading any of it. Nothing detects repetition — no module keys on repeated identical calls, repeated command text, or a turn ceiling.

The cost this leaves ungated is this repository's own most expensive failure mode. CLOUD-1337 measured eleven duplicate watchers running 9h 35m, found by a human reading ps. CLOUD-821 measured 490 self-manufactured wake-ups in one session, of which 2 changed a decision. CLOUD-889 measured a Stop hook that denied on every turn of a branch's life with a remedy that could not clear it.

What this row builds, and why it is the one that must land first

Three things, and none of them is a predicate — this is the spine every later predicate stands on.

  1. Stream::repeats() in transcript.rs, beside Stream::counts(). That is the established shape, and Counts's own comments refuse new fields for non-readers of the -J capability document, so do not widen Counts.

  2. Extraction::of(self, &transcript::Counts) -> usize widens to take &Stream. This signature change is what every subsequent row is blocked on, which is the reason to land it alone: the risk here is the plumbing, not the detection.

3. A per-extraction capability declaration — and this is the load-bearing half.

An extraction resolves three-valued, the way transcript::Capability already does for the transcript as a whole. Without it, an extraction a harness does not record reads as zero, and zero is a real answer meaning the extractor ran. That is the dead-gate class .claude/rules/policy-modules.md documents, and it is worse here than for hook-denials: zero denials is a plausible clean session; "zero repeats" over a two-hundred-call session is a false green over a session nobody measured.

This is what lets the surface EXPAND rather than shrink. Batten is harness-agnostic, and the wrong reading of that is to implement only what every harness shares. The right one is: if any agent emits a signal, Batten should read it, and cross-compatibility is carried as declared capability rather than as refusal to implement. The declaration is the precondition for every later member.

  1. agent-turn-run — the first member, and deliberately the one needing no hashing at all: a trailing run of assistant turns with no intervening tool call, over the already-typed Event::Turn(Role, Origin). It maps to OpenHands' monologue detector (3+ consecutive messages).

The framing the module header must carry, because the honest claim is narrower than "loop detection"

Prior art, for the predicate shapes: OpenHands' stuck detector (4+ identical action-observation cycles, 3+ action-error, 3+ monologue, 6+ ping-pong), Kilocode / opencode (identical (name, args-hash) in the last 3), hermes-agent (fingerprint 3+ times in a window of 20).

Academic anchor: When Agents Do Not Stop (IAL-Scan) — the defect is a feedback path that repeatedly reaches a costly or state-growing action without an effective bound; 91.9% precision over 68 failures in 47 of 6,549 projects.

And the honest limit, from termination analysis: a ranking function buys no better predicate here. It needs an observable decreasing measure, and every quantity Batten can see is a monotonically growing count; manufacturing a descent (K - distinct) is threshold counting with the subtraction moved. What it does buy is in the declaration — disjunctive termination is the correct justification for why this is a set of extractions with a set of thresholds rather than one number, and it forces the accurate claim: these supply an effective bound on a suspected feedback path. They do not detect non-termination. Say that in the header rather than calling it loop detection.

Refinement — Ready

Refinement gate: Definition of Ready & Done. This body carries only specializations.

{
  "source_of_truth": "crates/batten/src/transcript.rs",
  "gate": { "task": "verify", "exits": [0, 1] },
  "commit_type": "feat",
  "blockers": [],
  "tests": [{
    "file": "crates/batten/tests/it/repetition.rs",
    "mutation": "run-may-go-unbounded"
  }]
}
  • Authority boundary (§1). crates/batten/src/transcript.rs and crates/batten/src/facts.rs, plus the batten.toml extractor declaration. No mise-tasks/ program and no tests/**/*.bats is added or edited.
  • Computable predicate (§2). A declared extraction resolves to an integer where the harness records the events it reduces, and to could-not-look where it does not. agent-turn-run is a trailing run over Event::Turn. One module, one [[verdict]] row, warn.
  • Could-not-look (§2). Four states already collapse to null and null.x is undefined. This row adds the fifth: a transcript that parses but whose host records none of the events an extraction reduces. It must answer could-not-look, in the engine — never as a per-module conjunct, which is a dead gate on every harness but the one its author tested.
  • Deliberately not in scope (§2). Any repetition predicate over call identity — that is the next row and it is blocked on this one. Result digests. Thresholds beyond the single agent-turn-run one.
  • Effect (§3). read.
  • Generated artifacts (§4). schema/policy-call.schema.json regenerates via mise run fix; never hand-edited. No churn for new ids — extracted is already additionalProperties: {type: integer}.
  • Output and exit (§5). Pointer-only, and this is the family where it matters most: a count, never a span of session text, not even in a diagnostic. Could-not-look exits 3, never a false 2.
  • Commit / bump (§6). feat(facts) — patch until 0.1.0.
  • Test obligation (§7). Over the compiled binary, in a new crates/batten/tests/it/repetition.rs that installs the module. It may not reuse extracted_facts.rs, which is the fact tier and never installs one, so no mutation of a predicate could redden it (#MUTANT-OWNER). Shown able to fail per CLOUD-418, four observed: (a) agent-turn-run decides over a stream that has one; (b) the cross-compatibility case — a stream from a host with a different event shape asserts could-not-look, NOT zero, which is the only case distinguishing this from an absent gate; (c) a stream with no run is a real negative, distinct from (b); (d) the anti-vacuity positive control. Confirm the channel with an unconditional violation body, never an arm over the channel itself — that is exactly how CLOUD-1049's dead channel survived two measurements.
  • Blockers (§8). None. relatedTo CLOUD-894 (the declared firing-rate ceiling every predicate in this family must clear before promotion), CLOUD-895 (why detection may sit at Stop but a break may not), CLOUD-1337 (the measured 9h35m instance), CLOUD-889 (why the Stop surface cannot be the one that stops a loop).

Acceptance

  • agent-turn-run is declared, reaches a module, and is decided over — asserted on the compiled binary.
  • An extraction the harness cannot answer is could-not-look and is not reported as zero — asserted. The row fails its own review without this case.
  • No transcript text reaches the policy input or any output, asserted structurally.
  • Extraction::of takes &Stream, and Counts is unchanged.

Filed 2026-09-02 from a doom-loop grooming pass. The substrate was found live and unused; the gap is that nothing declares it.


Claims object added 2026-09-03, and the row was un-pullable without it. batten ready lint refused this body with claims-object-absent, so claim check answered not-ready and the spine of the whole loop family could not be claimed — while the row sat in Todo, which AGENTS.md defines as the ready queue rather than a status.

Every field is taken from this body's own clauses rather than chosen: source_of_truth from §1's authority boundary, the gate and feat(facts) from §6, no blockers from §8, and the tier from §7 — which names crates/batten/tests/it/repetition.rs and forbids reusing extracted_facts.rs.

The one field the body did not already contain is the mutation name. It was first written as capability-may-read-zero, aimed at §7's case (b) — and that was wrong, for a reason worth recording. The per-extraction capability §7 calls "the only case distinguishing this from an absent gate" is decided in the ENGINE, in Extraction::of, and mutate corrupts Rego. No mutation of the module can reach it, so the declaration would have been recorded as a SURVIVOR — a finding the runner exists to refuse, wearing the costume of coverage.

What shipped are two mutations that do discriminate, both over the module and both naming cases in the compiled tier: run-may-go-unbounded falsifies the threshold comparison and reddens a_monologue_run_is_reported; threshold-may-slip lowers the constant to 2 and reddens two_turns_in_a_row_is_not_a_run. The capability is still asserted — by a_host_that_records_no_turns_answers_could_not_look paired with a_session_with_no_run_is_a_real_zero — as a compiled-tier assertion rather than a mutation target, which is what it is.

Two defects were caught on the way in, both by the tier rather than by reading. The declarations first named test_ rules inside the .rego while #MUTANT-SUITE resolves the compiled file — declared, never applied, counted by nobody, which is the shape PR #829 measured in policy/harness-wiring.rego. And the first cross-compatibility fixture was built from a user record; the parser reads one AS a turn, so that transcript recorded turns after all and answered a real zero. A host recording hook runs and no turn boundaries is the shape the claim is actually about.

Refined by the session that then claimed it, so the claim carries --bypass-sequence rather than a takeover, and the receipt says so.


Landed 2026-09-03 on claude/cloud-1331-perf-base-arm-6pgmy2 (PR #833), commit 7dfcc36.

Extraction::of takes &Stream and returns Option<usize>; Counts is unchanged; runs live in Stream::repeats() beside Stream::counts(); the projection OMITS an unanswerable key so a module reads undefined and Rego takes that as does-not-hold. agent-turn-run is declared, registered and decided over at warn. Seven compiled-tier cases green; policy test 50 bundles / 638 passed.

The bound is stated in the code rather than absorbed: this cannot separate a host that records no hook runs from a session that triggered none, and it resolves that toward could-not-look — the safe direction, since a missing answer is reported where a false zero is indistinguishable from a clean session.

One adjacent cleanup rode along, named here rather than left to a reader of the diff: policy/harness-declared.json's two exemption rows are now spent. CLOUD-1079 has landed, the user-level hooks those rows excused are no longer provisioned, and the merged surfaces carry zero hook commands — so harness-wiring reported both, which is the gate working rather than misfiring. Reproduced against this same tree with --config-from a907565, so the finding predates this diff. Retired rather than skipped: leaving them is not neutral once their owner has landed, because they would silently excuse anything later matching either pattern by name.

CLOUD-1341 An across-turn poll is invisible to every gate: `run-shape` decides over a command string, so 1079 no-op tool calls — 59% of one session — read as work

Why

CLOUD-821 measured 490 backgrounded sleep N; tail log calls in one session against 2 that changed a decision, and landed the arm refusing a backgrounded sleep with no until/while. CLOUD-489 owns the adjacent case — a backgrounded loop polling the session's own tracked task — and states its own constraint precisely: "decidable from the command string alone."

Measured 2026-09-02 (CLOUD-1331's landing session). With mise run land backgrounded and driving the loop, the agent polled ReadNotifications (returning "No queued notifications") and tail -3 <task-output> across four separate background tasks. A human stopped it. First estimated at ~60 consecutive turns; measured afterwards at 1079 calls, 59% of every tool call in the session — see the refinement note below, where the estimate is corrected and the correction changes the predicate. Nothing in the engine fired, and nothing could have:

  • No sleep. ReadNotifications returns instantly, so it is CLOUD-821's defect with the one token that rule watches deleted.
  • No loop, and no single command. CLOUD-489's two families both key on until/while inside one command string. The repetition was across turns — ~60 separate foreground calls — so it appears in no command string at all.
  • Mostly not on the shell surface. ReadNotifications is a harness tool with no argv, so input.call.segments gives shape and pipeline rows nothing to match. The tail half is a bare read no row objects to, correctly.

AGENTS.md already states the invariant — "The exit notification IS the wake-up; waiting for it costs nothing… 'idle' means a turn with NOTHING backgrounded" — and for this family it is prose, which non-negotiable rule 2 calls half a change.

The instrument exists and is one declaration short. CLOUD-1172 built input.facts.extracted: a closed set of integer counts over typed transcript events, with input.call["stop-repeat"] as the Stop-side discriminator. policy/denials-outlive-the-turn.rego already composes exactly this shape (extracted.denials > 0 AND stop-repeat == true). What is missing is a member that can see repetition — Extraction today is Turns, ToolCalls, ToolErrors, HookDecisions, HookDenials, and ToolCalls is a flat total, so that session's calls read identically to a productive session's.

Refinement — Ready

Refinement gate: Definition of Ready & Done. This body carries only specializations.

{
  "source_of_truth": "crates/batten/src/transcript.rs",
  "gate": { "task": "verify", "exits": [0, 1] },
  "commit_type": "feat",
  "blockers": [],
  "tests": [{
    "file": "crates/batten/tests/it/extracted_facts.rs",
    "mutation": "repeats-may-go-unpriced"
  }]
}
  • **Authority boundary (§1). **crates/batten/src/facts.rs (one Extraction member and its of arm), crates/batten/src/transcript.rs (the count it derives), a new policy/*.rego module, and the [[rule]] + [[rule.extract]] + threshold rows in batten.toml. No mise-tasks/ program and no .bats.

  • Computable predicate (§2). A closed-set member counting tool calls whose TOOL AND ARGUMENTS the session has already sent — the recurrence count. The invariant it encodes is exact: a call you have already made, with the same arguments, told you what it told you the first time. An integer over typed events, so rule 4 stays decided in the fact's TYPE as the rest of this family is.

    The RESULT is deliberately not part of the identity, and that is a rule 4 decision rather than an economy. Including it discriminates slightly better — measured 60x separation against 28x — but it would oblige the parser to read every tool result's BODY in order to digest it, which is the one payload transcript.rs exists to never touch. Arguments are already retained for CLOUD-98's typed predicate, so identifying a call by them reaches no further into the session than the engine already does. The cheaper fact is also the safer one.

    This replaces "consecutive calls to one tool with no intervening state change", which is measured below to be undefined AND non-discriminating. The recurrence count is immune to the alternation that defeats a consecutive-run reading: it does not care what sits between two identical calls.

  • Deliberately not in scope (§2). Grading transcript CONTENT — that needs CLOUD-1029's tamper-evidence first, and denials-outlive-the-turn's own note records that this declaration authorises no such consumer. Widening Extraction into an expression language: the set is closed on purpose, and an open one puts the session's own text within a module's reach.

  • Effect (§3). read. It reads a count the engine already derives from a record it already parses.

  • Output and exit (§5). The count and the tool name, never a byte of the session. The 0/1/2/3 table. A transcript that is absent, unparseable, or has no declared extractor answers could-not-look — four null states this family already keeps distinct — never a refusal.

  • Commit / bump (§6). Type feat(facts), so patch until 0.1.0.

  • Threshold (§2), measured rather than chosen: 100, against the MAXIMUM per identity rather than a sum. Healthy ceiling 38, defect 1079 — 2.6x clear of the one and 10.8x below the other. It lives in the policy/*.rego CONSUMER module rather than in batten.toml: the only consumer surface a module can read is data.batten.patterns, and .claude/rules/policy-modules.md refuses a threshold spelled as a pattern outright. A constant in a consumer module is the consumer's own number in the consumer's own file, which is the shape ci-parity.rego already uses.

  • Test obligation (§7). The module's test_ rules pin the predicate. crates/batten/tests/it/extracted_facts.rs is the tier that proves the ENGINE builds the count, because a with input as case fabricates the very shape the engine may be unable to produce. Shown able to fail per CLOUD-418: a fixture stream of N identical consecutive calls must redden, and a productive stream of the same length must not — the anti-vacuity mirror, without which a predicate firing on any busy session satisfies the deny case while deciding nothing.

  • Blockers (§8). None. relatedTo CLOUD-489 (the adjacent family, and the row whose §2 cannot reach this by construction), CLOUD-821 (the sleep-keyed arm this succeeds), CLOUD-1172 (the extractor family), CLOUD-1051 (the Stop surface).

Acceptance

  • A fixture stream of N identical consecutive tool calls refuses; a productive stream of the same length does not.
  • That discrimination is shown red before the fix.
  • Absent, unparseable and no-extractor transcripts each answer could-not-look rather than clean.

Refined 2026-09-02. All three open questions are closed — two by analysis, one by measurement over a real transcript. The row's own §2 predicate was refuted in the process.

The measurement: the one session transcript this container holds (13 MB, 4298 turns), parsed to (tool, argument digest, result digest) triples. Integers only; no session text was read out. Cross-check on the parse: this counted 1828 tool calls against the engine's own extractor reporting 1816.

The defect arm is far worse than this row claimed. The body says "~60 consecutive turns". The truth is ReadNotifications called 1079 times — 59% of every tool call in the session — every one with identical arguments and the identical "No queued notifications" result. All 1079 carry role: assistant, are not sidechain and are not meta, so they are agent-initiated: the gate would judge agent conduct and cannot fire on harness behaviour, which is the safety question that had to be settled before any of this could ship.

Q1 — "no intervening state change" is replaced, and the replacement is forced by the data. The polls arrive in 249 separate runs, median length 5, longest 8. So "longest consecutive run of one tool" tops out at 8, which any healthy session reaches — the row's original §2 cannot discriminate and would have shipped as coverage, which is exactly what CLOUD-418 exists to catch. The recurrence count does discriminate, because it ignores what sits between two identical calls. That is the whole reason for the reformulation.

Q2 — the null, and the separation is ~60x rather than marginal:

identity recurrences (tool+args) recurrences (tool+args+result)
the poll 1079 1079
next-most-repeated call 38 18
separation 28x 60x

Both discriminate comfortably. The shipped fact is tool+args, at threshold 100 — 2.6x above the healthy ceiling of 38 and 10.8x below the defect — because the result-bearing identity would require the parser to read result BODIES to digest them. Paying 28x instead of 60x to never touch that payload is the trade rule 4 asks for, and CLOUD-199's false-positive worry is answered either way.

Its limit, stated rather than absorbed: n = 1 session, and a pathological one. That biases the healthy arm upward — a well-behaved session should recur less — so 38 is an upper bound on the healthy ceiling and the direction is conservative. A cross-session null is unavailable here: this container holds exactly one transcript. Widening it is worth doing when a corpus exists, and would move the constant within its decade at most.

Q3 — the fact shape fits, and the question dissolves. "Longest consecutive run" would have been a max over RUNS. The recurrence count is a scalar reduction over the stream, so Extraction's vocabulary is not widened and there is no second axis.

CORRECTED 2026-09-03, and this sentence shipped the defect once. It read "a plain total over the stream, the same shape as ToolErrors", and an implementation followed it literally: a running total of recurrences SUMMED across every identity. That quantity is not the one measured anywhere in this row. The table below is per identity — 1079 for the poll, 38 for the next-most-repeated call — while the sum over that same transcript was 1294. A total across identities is monotonic in the length of the session and never resets, so it grows on any session long enough to repeat anything, and a threshold taken from one identity's recurrence and applied to that total is not a threshold at all. The fact is the MAXIMUM over identities; ToolErrors is the wrong analogy, because a tally is exactly what this must not be.

Measured consequence: shipped as a sum at severity = "deny", the rule fired at ~1300 in the session that authored it and refused every subsequent tool call — the push included — with no in-session recovery.

One residual, decided rather than left open. A legitimate repeated command — git status run many times across a long session — recurs, and under the tool+args identity it recurs whether or not its output changed. That is exactly why the healthy arm is measured at 38 rather than 18, and why the threshold sits at 100: the measured ceiling already contains that population, from the worst-behaved session on record, and the constant clears it by 2.6x.

What is not in doubt: the measured instance, and that no existing gate can reach it. Both run-shape's sleep key and CLOUD-489's command-string families are shown inapplicable in the body above, and neither is a matter of tuning.


Second null, measured 2026-09-03 — the silence check this row's own incident demands.

The shipped predicate was run against a real, healthy session's transcript over the compiled binary: repeated-calls = 0, over 180 tool calls and 576 turns. Cross-checked with the tool-calls and turns extractors on the same run, so a 0 from a transcript nobody parsed is ruled out — the parse is real and the answer is honest.

That is a second point on a null this row previously had at n = 1, and that one pathological. The healthy arm is 0 here against a measured ceiling of 38 and a threshold of 100, so the constant clears an ordinary working session by the whole of its range rather than by a margin.

It ships at severity = "warn" regardless. A mutation proves the predicate discriminates; it never proves the constant is right, and that distinction is what the incident cost. Promotion to deny is a separate decision, needing this evidence plus a wider corpus — and the recovery gap on CLOUD-1051 closed first, since an admission cannot clear a mediated_call deny.

CLOUD-1347 The measured doom-loop shape — the same call, again — has no predicate, and the fingerprint that decides it must never leave `transcript.rs`

Why

This is the shape every practitioner detector converges on, and the one Batten cannot express.

implementation predicate threshold
Kilocode / opencode identical (name, args-hash) in the last 3 calls 3
hermes-agent / PraisonAI one fingerprint 3+ times in a sliding window of 20 3
OpenHands identical action-observation cycles 4

It is also the shape measured in this container: CLOUD-1337's eleven watchers were the same call written eleven times, and the session that wrote them had read and agreed with the prose forbidding it. A rule an agent agrees with and then violates eleven times is the definition of a gate-shaped problem.

The rule-4 question, and why it dissolves

Repetition needs identity to compute, not to express. Every predicate above reduces to a threshold on an integer — "the last 3 hashes are equal", "one fingerprint 3+ times". None of them references an argv, and none of them needs to.

So the fingerprint is computed inside transcript.rs, where Event::ToolCall { id, name, input: Value } already holds the whole argument object and already never renders it, and only a run length is projected. Extraction stays integers-only and non-negotiable rule 4 is untouched — the type remains the guarantee rather than a habit each author keeps.

The precedent is in the same file and should be cited in the implementation: Event::HookOutput carries a digest, bytes hashed and dropped inside collect so that no caller can recover them. Do the same here — hash and drop in the same expression.

A "last call fingerprint" token is explicitly refused. It would put a correlatable identity of session content on the policy input, and no predicate here wants one. If a future row believes it needs one, that is its own rule-4 argument to make.

The two members and the predicate

  • repeat-depth — the trailing run of identical (name, digest(input)) over ToolCall.
  • distinct-calls — distinct fingerprints this session; the denominator, and the progress term.

One predicate in policy/repetition-without-progress.rego, scope = "mediated_call", at warn:

  • identical-call-run — repeat-depth >= 3.

Threshold in the module, never a [[pattern]] row — .claude/rules/policy-modules.md refuses spelling a threshold as a pattern, and a window or a count is a claim about one harness's call cadence rather than a concept with one spelling. In-repo module, never a preset, for the same reason: a preset cannot read a consumer's config, so a preset carrying a tuned threshold asserts one cadence everywhere.

Why trailing-run adjacency is the design, not an implementation detail

It is what does the false-positive work, and the trap is measured: OpenHands' detector kills agents legitimately waiting on long-running processes and leaves them unrecoverable once flagged. CLOUD-199 already set this repo's bar — a guard with false positives gets bypassed.

mise run test three times with edits between is progress, and a trailing run is broken by any intervening distinct call, so the edit-then-retest loop is already false under repeat-depth without a carve-out. That is why this row ships trailing-run only and leaves window recurrence to its own row as permanent-warn.

Second conjunct, from the same trap: input.call["run-in-background"] != true. Three-valued and compared with != rather than ==, following run-shape.rego — most hosts send no such key, and an unknown-posture wait is the case to be strict about.

Refinement — Ready

Refinement gate: Definition of Ready & Done. This body carries only specializations.

  • **Authority boundary (§1). **crates/batten/src/transcript.rs (Stream::repeats(), the fingerprint), crates/batten/src/facts.rs (two members), policy/repetition-without-progress.rego (new), batten.toml (two extractor declarations, the rule row, one [[verdict]] row).
  • **Computable predicate (§2). **is_object(input.facts.extracted) and extracted["repeat-depth"] >= 3 and input.call["run-in-background"] != true. Severity warn.
  • Could-not-look (§2). Inherited from CLOUD-1344 — an unanswerable extraction is could-not-look, never zero. Direction of failure, stated rather than discovered: a missing transcript and a hard loop are indistinguishable here, and this fails open, so the loop keeps running. That is correct — could-not-look is allow across the whole engine and one family may not invert it. A bound surviving a missing transcript would have to be a call-count ceiling bought elsewhere.
  • Deliberately not in scope (§2). Denying — this ships warn and promotion is its own row with its own measurement. Window recurrence and ping-pong alternation. Identical observations, which Event::ToolResult cannot express.
  • **Effect (§3). **read.
  • Output and exit (§5). Pointer-only: the run depth as an integer and the rule id. Never the call name and never the arguments — a repeated call's operand carries this consumer's paths and task names, which is exactly what CLOUD-1337's own pointer-only clause protects.
  • **Commit / bump (§6). **feat(policy) — patch until 0.1.0.
  • Test obligation (§7). Over the compiled binary in crates/batten/tests/it/repetition.rs, which installs the module — not extracted_facts.rs, which is the fact tier and never installs one, so no mutation of this predicate could redden it (#MUTANT-OWNER). Shown able to fail per CLOUD-418, five observed: (a) three identical trailing calls raise; (b) two identical calls with a distinct call between them do NOT — the adjacency case, and the one that distinguishes this rule from the wrong one; (c) the anti-vacuity mirror — a session of distinct calls is clean, without which the arm is satisfied by a check that refuses everything; (d) a backgrounded call with a trailing run is clean; (e) could-not-look is not reported as clean. Declare #MUTANT-SUITE and choose a mutation that discriminates — a mutation on a conjunct another conjunct already excludes will survive.
  • **Blockers (§8). **blockedBy CLOUD-1344 (the &Stream signature and the capability declaration) and CLOUD-1345 (the per-call whole-transcript parse, which per-call fingerprinting makes materially worse and which must narrow first). relatedTo CLOUD-894 (this predicate may not be promoted without a measured rate against a declared ceiling), CLOUD-1337 (the measured instance).

Acceptance

  • Three identical trailing calls raise a warn naming a depth and nothing else.
  • An intervening distinct call clears it — asserted, in both directions.
  • A backgrounded wait is untouched.
  • No call name, argument, or fingerprint reaches the policy input or any output, asserted structurally.

Filed 2026-09-02 from a doom-loop grooming pass.

CLOUD-97 Flag sessions that signal done with work not landed

Why
Batten's threat model names "the wrong completion signal", but nothing compares a declared stopping point to repo state. Via the transcript-analysis substrate (CLOUD-95), a post-session check reads "the transcript contains a completion signal" ∧ "the branch is not landed" from the record + git state — computable from the unlanded set (CLOUD-37) and patch-id/content merged-ness (CLOUD-36), no ancestry assumption. Self-clearing and latency-tiered (CLOUD-78/CLOUD-80): raised at session close, cleared the moment the work lands, so it never blocks a legitimate mid-task pause. A live Stop hook is an optional latency optimization, not a requirement.

Rejected alternative
A hard deny at "done". Pausing/handing off mid-task is legitimate; blocking "done" punishes correct behavior. Detection + self-clearing, not prevention.

Definition of done

  • Over CLOUD-95's event stream + git state, compute completion-signaled ∧ ¬landed.
  • Register a self-clearing finding (CLOUD-78) at a response-latency tier (CLOUD-80); it clears on the next evaluation once landed.
  • Output pointer-only, within the drain token budget (CLOUD-82): count + branch/session pointer.

Acceptance

  • A fixture session ending with committed-but-unlanded work raises the finding; landing by fast-forward clears it with no manual ack.
  • A rebased-and-landed branch does NOT raise (patch-id/content, not ancestry).
  • No repo-specific identifiers.

Refinement — Ready (completion-signaled ∧ ¬landed as a structural predicate; advisory finding, self-clearing on land)

Refinement gate: Definition of Ready & Done. This body carries only specializations.

  • Source of truth (§1). The one authoritative artifact is the done-but-not-landed rule in crates/batten. Its two inputs are consumed, never re-typed: the completion signal reads typed stop/turn fields from CLOUD-95's event stream; landedness reads git state (patch-id equivalence against origin/<default>), never a second bookkeeping file.
  • Computable predicate (§2). Not expressible as a batten.toml rule: it needs the transcript input and the findings store, the engine capabilities linked in §8 (CLOUD-95, CLOUD-78). The predicate is structural end to end: completion-signaled = the session's final turns carry a stop/completion marker in typed fields (exact-match token set, no text inference); ¬landed = a commit authored in the session has no patch-id-equivalent commit on origin/<default> — content comparison, no ancestry assumption. Evaluated post-session over the transcript input; the gate for this issue is the fixture suite in mise run test, wired into the hk gate and CI.
  • Effect (§3). No new command and no new effect-table row. Repo and transcript inspection are read-only; registering the finding is an append to the out-of-tree store under CLOUD-78's contract.
  • Output & exit (§5). Pointer-only within the drain budget: count + branch/session pointer, never commit content. Advisory: raising or clearing this finding never produces a blocking exit from check (house style §0.3 — advisory surfaces are structurally unable to block); escalation is CLOUD-80's latency model, not this rule.
  • Commit / bump (§6). feat → minor.
  • Test obligation (§7). E2E over the compiled binary (crates/batten/tests/cli.rs), fixture transcript + fixture repo: (a) a completion signal with committed-but-unlanded work raises the finding; (b) landing by fast-forward clears it on the next evaluation, no manual ack; (c) a rebased-then-landed branch (same patch-id, different SHA) does not raise; (d) byte-identical output across two runs over the same inputs.
  • Blockers (§8). blockedBy CLOUD-95 (the typed event stream this rule reads) and CLOUD-78 (the store the finding registers into). relatedTo CLOUD-36/CLOUD-37 — the ¬landed predicate is implementable directly with git patch-id plumbing and migrates onto their shared unlanded/content primitives when they land, so neither blocks; CLOUD-80/CLOUD-82 shape the finding's tier and drain rendering, not the predicate. Refinable now; implement after CLOUD-95 and CLOUD-78 land.

Review in Linear

@coderabbitai

coderabbitai Bot commented Sep 2, 2026 •

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

Next included review available in 59 minutes.

Check out review usage here.

View limit details

Limit details: You’ve used the included review currently available.

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Free

Run ID: f2e77b42-81ac-41d8-988d-366ccfeed8f3

📥 Commits

Reviewing files that changed from the base of the PR and between a907565 and 55ca205.

⛔ Files ignored due to path filters (1)
  • mise-tasks/sbom-actions.tsv is excluded by !**/*.tsv
📒 Files selected for processing (18)
  • .claude/rules/toolchain.md
  • .github/workflows/ci.yml
  • .github/workflows/perf.yml
  • AGENTS.md
  • batten.toml
  • crates/batten/src/facts.rs
  • crates/batten/src/lib.rs
  • crates/batten/src/transcript.rs
  • crates/batten/tests/it/claim_order.rs
  • crates/batten/tests/it/main.rs
  • crates/batten/tests/it/repetition.rs
  • crates/batten/tests/it/stop_posture.rs
  • mise.toml
  • policy/ci-parity.rego
  • policy/claim-order-is-stated.rego
  • policy/repetition-without-progress.rego
  • schema/batten.local.schema.json
  • schema/batten.schema.json
ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Free

Run ID: 9cf60ba5-b40b-498f-be7c-a5feb0f1a83f

📥 Commits

Reviewing files that changed from the base of the PR and between de872d6 and a907565.

📒 Files selected for processing (1)
  • batten.toml

Included review availability: Your plan provides up to 1 included review per hour; 0 remain after this review.


📝 Walkthrough

Walkthrough

The performance job now caches a merge-base-independent seed directory. It restores the seed, removes the measured binary, rebuilds the base arm, and saves the resulting closure. Batten now reports path reach dead for expression-based actions/cache paths. Configuration adds SBOM, review, prose-partition, startup-repair, task-watch, and review-run checks. Policy and integration tests cover dynamic paths, dynamic keys, unrelated actions, and clean fixtures.

Merge Risk: ⚪ Minimal · up to a9075

The PR stabilizes the CI performance-base cache path while preserving correctness, and the supplied checks are green. No actionable merge-blocking risk remains beyond normal review.


Note

🎁 Summarized by CodeRabbit Free

Your organization is on the Free plan. CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please upgrade your subscription to CodeRabbit Essentials by visiting https://app.coderabbit.ai/settings/billing.

Comment @coderabbitai help to get the list of available commands.

@wenzowski
wenzowski force-pushed the claude/cloud-1331-perf-base-arm-6pgmy2 branch from de872d6 to a907565 Compare September 3, 2026 00:24
wenzowski added a commit that referenced this pull request Sep 3, 2026
…not a total

An across-turn poll was invisible to every gate. Measured over one real
transcript: `ReadNotifications` called 1079 times, 59% of every tool call
in the session, every one with identical arguments and the identical "No
queued notifications" result. All 1079 carry `role: assistant`, so this
judges agent conduct and cannot fire on harness behaviour.

No landed arm could reach it. `run-shape` keys on a backgrounded `sleep`
and CLOUD-489 on `until`/`while` inside one command string; repetition
spread across turns carries neither token, and a harness verb with no
argv gives a `shape` row nothing to match.

THE FACT IS A MAXIMUM OVER IDENTITIES, NEVER A SUM. `repeated_calls` is
how far the session's most-repeated (tool, arguments) identity ran. A
running total across every identity is monotonic in the length of the
session and never resets, so it fires on anything long enough to repeat
anything: over the same transcript the sum was 1294 where the max was
1079 against a healthy ceiling of 38. A threshold derived from one
identity's recurrence and applied to that total is not a threshold.

The row's own §2 is refuted and needs rewriting on the tracker. It
proposed "consecutive calls to one tool": the polling arrived in 249
bursts whose longest run was 8, a length any healthy session reaches, so
a consecutive reading ships as coverage while deciding nothing.

Identity is the tool name and a DIGEST of the arguments, so no argument
text is retained. The RESULT is deliberately excluded — it separates
better (60x against 28x) and would oblige the parser to read every
result body, the one payload `transcript.rs` exists never to touch. A
replayed `tool_use` id is deduped: counting replays measures what the
host chose to re-emit rather than what the session did.

IT SHIPS AT `warn`. A `mediated_call` row at `deny` refuses every later
tool call once it fires, and no admission can clear it — `batten
override request` and `spend` are themselves mediated calls, so
requesting one requires making the call the deny refuses. This predicate
was landed at `deny` once and locked its own authoring session out at
~1300, push included. Promotion needs it shown silent against a real
transcript first.

The compiled tier drives the engine rather than a fabricated input,
which is what pins the off-by-one: N calls are N-1 recurrences, so
clearing 100 takes 102 calls and 101 is clean.

Refs: CLOUD-1341, CLOUD-1172, CLOUD-418, CLOUD-489

Admits: 3d6283f7ae44fe7b8349bbb4ef3aa0b22e03bc84ad4059098ff15576b3729337
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: batten.toml
Admits-head: a907565
Admits-epoch: 93d8b1112718944856b6e81b9cfac7c903dbb2c28470be1fc6e59786fb42f404
Admits-author: alec@wenzowski.com
Admits-prev: a387bad2ea98589d0f75455b02cba853248836b7c51d4707cf39585329262b06
Admits-answer-lost: The module cannot load, so CLOUD-1341 ships nothing. An unregistered `policy/*.rego` file is inert — the engine reads its rule set from the committed authority, and a module with no registering row is never evaluated. The extractor row is what makes the count reach the predicate at all: without it the fact is undefined, Rego reads undefined as does-not-hold, and the gate loads clean while deciding nothing, which is the dead-gate class `.claude/rules/policy-modules.md` exists to warn about.
Admits-answer-precondition: Registering a policy module IS an edit to the committed authority: a `[[rule]]` row naming the module, a `[[rule.extract]]` row binding the new `repeated-calls` extractor to the key the module reads, and the `[[verdict]]` plus `[[verdict.route]]` rows declaring the `turn ask twice` token. A module raising an undeclared token fails to load, and a declared verdict row nothing raises fails the load too, so the rows and the module are one indivisible change. There is no other surface that expresses them, and the write lands in PR #833's diff where a reviewer sees it.
Admits-answer-rejected-route: `config read first` is what this IS — that route names reading configuration before writing it, and the rows being added do not exist to be read, because the module is new. `patch run first` does not apply: the file is hand-authored and has no generator behind it, so a patch route would write the same bytes to the same protected path; the generated `schema/*.json` are derived FROM the crate rather than producing it.

Admits: 247eaff38bfbc59e1afae9cae684be633d05c8eca2d901051cad86757508560c
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: policy/a-repeated-call-is-not-progress.rego
Admits-head: a907565
Admits-epoch: 93d8b1112718944856b6e81b9cfac7c903dbb2c28470be1fc6e59786fb42f404
Admits-author: alec@wenzowski.com
Admits-prev: e4dad46c2f5408cdb8bdb3487a98b9e7035341caad7cb39e7e1d555605d0d1f9
Admits-answer-lost: CLOUD-1341 stays prose. The always-loaded instructions already state that a backgrounded task's exit notification is the wake-up and that re-asking is not waiting, and non-negotiable rule 2 calls a rule without a runnable gate half a change. The two landed arms cannot reach this family: one keys on a backgrounded `sleep`, the other on `until`/`while` inside a single command string, while the measured defect was 1079 identical harness-verb calls spread across turns with no argv to match. Without the module there is no gate at all.
Admits-answer-precondition: CLOUD-1341's remedy is a policy predicate, and a predicate is a Rego module: the threshold constant, the null guard and the eight test rules are the module's own body, which no configuration surface can express. The committed authority carries only the rows that register it. Writing the protected path is the only route left, and the file lands in PR #833's diff where a reviewer sees it.
Admits-answer-rejected-route: `config read first` does not apply: the committed authority has no surface for a predicate body or a threshold constant, and `.claude/rules/policy-modules.md` refuses a threshold spelled as a pattern row outright, so the number has to live in the consumer module. `patch run first` does not apply either: the module is hand-authored with no generator behind it, so a patch route would write the same bytes to the same protected path.
wenzowski added a commit that referenced this pull request Sep 3, 2026
…a pull request

CLOUD-1331's keyed base arm collected nothing across 4 runs and 2 pull
requests: 0 hits, ~190 MB written and discarded each time. The first
diagnosis blamed the key; the second blamed the path. Both were reading
the same design, which optimised for a hit and never asked what a miss
costs.

An entry written from a pull request is scoped to `refs/pull/N/merge`
and no other pull request can read it. An entry written from the default
branch is readable by every branch. So the PR side now RESTORES and
never saves — `actions/cache/restore`, a key carrying no merge base
because nothing saves and there is no miss to engineer — and the
scheduled `perf` job on `main`, which already pays a full `--release`
build, stages that closure and saves it under the same key. No new
trigger, no second build.

The seed is never measured: it is copied to `target/perf/base-$SHA` with
its binary removed, so `base_arm_is_built()` stays false and cargo runs
against the base tree actually materialised.

CLOUD-840 measured this repository at 225.2 GB of Actions cache, 203.7 GB
of it across 1,144 PR-ref entries nothing can restore. This job now
writes zero of them, and the hit arrives on a pull request's FIRST run
rather than its second.

`cache-path-is-rebase-stable` gains the two sub-action spellings.
`actions/cache/restore` and `actions/cache/save` derive an entry's
version from `path` exactly as the composite does, so matching only
`actions/cache@` would have left the predicate live and reaching nothing
the moment a job split restore from save — which is what this commit
does. Its own case is declared rather than assumed.

Refs: CLOUD-1342, CLOUD-1331, CLOUD-840

Admits: b0862b76d140be6746e8cdd0fcaf3c28e5bc685ef40c4e3164a1079bf12d075b
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: .github/workflows/ci.yml
Admits-head: fb2987d
Admits-epoch: 93d8b1112718944856b6e81b9cfac7c903dbb2c28470be1fc6e59786fb42f404
Admits-author: alec@wenzowski.com
Admits-prev: f0a932f52d256fb53185a39552efa1b6ac28551baa5b3c020412742000541260
Admits-answer-lost: The job keeps writing ~190 MB per run of cache scoped to the pull request's own merge ref, which no later run can ever restore — CLOUD-840 measured 203.7 GB of such entries across 1,144 of them in this repository — and every pull request keeps compiling the base arm's whole release closure cold, 239 crates in 4m42s of a 13.5 min step. The design being replaced optimised for a hit and never asked what a miss costs.
Admits-answer-precondition: The change is a GitHub Actions step in the `perf` job — swapping the composite cache action for its restore-only sub-action, dropping the merge-base SHA from the key, and deleting the pull-request-side save step. No batten surface expresses a workflow step, so writing the protected path directly is the only route left, and the write lands in PR #833's diff where a reviewer sees it.
Admits-answer-rejected-route: `config read first` does not apply: the committed authority declares rules, verdicts and patterns and carries no representation of a workflow job or step, so there is nothing to read or set there. `patch run first` does not apply either: the workflow is hand-authored YAML with no generator behind it, so a patch route would write the same bytes to the same protected path.

Admits: a76c9a87c95fa32cf60a4b398c75005fdba89cb88255963faf88c77be092502f
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: .github/workflows/perf.yml
Admits-head: fb2987d
Admits-epoch: 93d8b1112718944856b6e81b9cfac7c903dbb2c28470be1fc6e59786fb42f404
Admits-author: alec@wenzowski.com
Admits-prev: 22dc584248885224d59327013514c1c438affb90392824a98c742447a18e3b7e
Admits-answer-lost: The restore-only step this branch lands has nothing to restore. Only a run on the default branch can write a cache entry every branch can read, and this scheduled job is the one place that already pays a full release build. Absent the seed, every pull request keeps compiling the base arm's whole release closure cold, and writing unreadable per-pull-request entries stays the only alternative.
Admits-answer-precondition: The change adds two GitHub Actions steps to the scheduled job that already runs on the default branch — staging the release closure it has just built, and saving it under the key the pull-request side restores. No batten surface expresses a workflow step, so writing the protected path directly is the only route left, and the write lands in PR #833's diff where a reviewer sees it.
Admits-answer-rejected-route: `config read first` does not apply: the committed authority declares rules, verdicts and patterns and carries no representation of a workflow job or step, so there is nothing to read or set there. `patch run first` does not apply either: the workflow is hand-authored YAML with no generator behind it, so a patch route would write the same bytes to the same protected path.

Admits: 7d1ad5bf1d04296f330e798f42f2a74aba2e5d7f78ab4b9a8d2e2474c9dea5e3
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: policy/ci-parity.rego
Admits-head: fb2987d
Admits-epoch: 93d8b1112718944856b6e81b9cfac7c903dbb2c28470be1fc6e59786fb42f404
Admits-author: alec@wenzowski.com
Admits-prev: d2e45c5830908bdc08f6f59141f3a9b031f68ad27ffdd7d133ebf108d0b30921
Admits-answer-lost: The gate goes dead against its own subject. The sub-actions derive an entry's version from the cached path exactly as the composite does, so an interpolated path is the identical silent total miss — the one measured at 4 runs, 0 hits, ~190 MB discarded each time — and after this branch the repository's only cache steps of that kind are the sub-action spellings the predicate cannot see. A gate that loads clean and matches nothing reads exactly like a clean tree.
Admits-answer-precondition: The predicate matches only the composite spelling of the cache action, and this branch moves the job to the restore-only sub-action plus a save on the scheduled side — so the gate that exists to refuse an interpolated cached path would stop covering the very steps this branch lands. Widening it and adding the discriminating case is a change to the Rego module itself; no configuration surface expresses a predicate body, so writing the protected path directly is the only route, and it lands in PR #833's diff where a reviewer sees it.
Admits-answer-rejected-route: `config read first` does not apply: the committed authority carries the rows that register this module, not the predicate body, so nothing there can widen which action spellings the rule matches. `patch run first` does not apply either: the module is hand-authored and has no generator behind it, so a patch route would write the same bytes to the same protected path.
wenzowski added a commit that referenced this pull request Sep 3, 2026
…viting the second claim

The claim receipt is minted on `claim check`'s pullable path and keyed by
the branch checked out at that moment. That one fact makes two orderings
fail in opposite directions, and both are reachable by following the
instructions correctly the first time.

Claim, then branch: the receipt is minted against the branch you were
standing on, and the first edit on the real branch is refused with
`receipt read missing claim branch claim-needs-receipt`.

Claim twice: having branched first, the pullable message read as
go-do-this-and-come-back, so the next move was to move the row and run
`claim-check` again. The second run arrives after the row has left Todo,
reads it as held, and refuses `not-todo` — where the holder is the caller
ninety seconds earlier. The only route past is `--takeover` against
oneself, which writes a takeover record for a row nobody else touched,
and CLOUD-1139 needs that signal to stay rare. Measured twice, in two
sessions, both by an agent following the documented order.

NO ENGINE CHANGE, and the first revision of this row assumed one was
needed. Creating a branch writes only under `.git`, which
`claim-needs-receipt` never judges, so there is no ordering deadlock —
the deadlock was imagined. `claim.rs`'s `not-todo` decision is correct as
it stands for a row genuinely held elsewhere and is untouched.

Recognising a self-claim in the engine stays CLOUD-1139's: it needs to
tell this session from a sibling, and the row's assignee cannot, since
every fleet session carries the same configured accountable identity. An
assignee-keyed re-mint would let a sibling silently re-mint over a working
holder — that row's measured harm, reintroduced through this row's fix.

The gate decides what a gate can: whether the always-loaded file still
states the order and whether the triggered file still carries both failure
directions. Whether a given session actually claimed before branching is
not a property of the tree, and a rule resolving to it would be the model
verdict non-negotiable rule 3 forbids.

Two files because a budget forced it. `policy-budget` caps the index at
3500 tokens and 199 lines, and an earlier draft of this change blew both
at 3516/202; the index carries only the ORDER and the reason lives in the
rules file that loads at the trigger. Both arms are therefore required.

The compiled tier is what proves the two markdown files reach the module
at all: a dead gate and a tree that still states the order are
byte-identical on the decision surface, so each drift case can only go red
if the lines actually arrived.

Refs: CLOUD-1343, CLOUD-1139, CLOUD-786, CLOUD-733

Admits: 3fe3ee98cc0ec6d5f1a9ea595f823c2daa76332568b0ca2332363539c5979d99
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: batten.toml
Admits-head: f9f48e6
Admits-epoch: 84237be51e176ca21fe43fbe5f50492b50c188b8881b9745dfb2b4004a75a220
Admits-author: alec@wenzowski.com
Admits-prev: 3d6283f7ae44fe7b8349bbb4ef3aa0b22e03bc84ad4059098ff15576b3729337
Admits-answer-lost: The gate is dead rather than absent, which is strictly worse. Without the `line_sources` list the engine builds no lines for those two paths, every clause reads undefined, Rego takes undefined as does-not-hold, and the module loads clean while deciding nothing — indistinguishable on the decision surface from a tree that still states the order. That is the exact failure class `.claude/rules/policy-modules.md` records, and without the registering row the module is never evaluated at all.
Admits-answer-precondition: Registering the claim-order module IS an edit to the committed authority: a tree-scoped `[[rule]]` row naming it, the `line_sources` list that makes the two instruction files reach the predicate at all, and the `[[verdict]]` plus `[[verdict.route]]` rows declaring the `claim declare dropped` token. A module raising an undeclared token fails to load, and a declared verdict row nothing raises fails the load too, so the rows and the module are one indivisible change. There is no other surface that expresses them, and the write lands in PR #833's diff where a reviewer sees it.
Admits-answer-rejected-route: `config read first` is what this IS — that route names reading configuration before writing it, and the rows being added do not exist to be read, because the module is new. `patch run first` does not apply: the file is hand-authored and has no generator behind it, so a patch route would write the same bytes to the same protected path.

Admits: f4a603586d8e7912a35d4358fe20ed1955c112b132ecbaf9ad46eebc2bfa1742
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: policy/claim-order-is-stated.rego
Admits-head: f9f48e6
Admits-epoch: 84237be51e176ca21fe43fbe5f50492b50c188b8881b9745dfb2b4004a75a220
Admits-author: alec@wenzowski.com
Admits-prev: -
Admits-answer-lost: The order evaporates. The failure mode here is DRIFT — the remedy is prose, and prose is what a later edit silently rewords or drops, after which the next session walks into the same refusal that cost this one a full stop mid-work and a takeover record against its own claim. A prose assertion with no gate is exactly the half-change rule 2 refuses, and the Ready gate already refused an earlier revision of this row for carrying an empty tests array.
Admits-answer-precondition: CLOUD-1343's remedy is words in two instruction files, and non-negotiable rule 2 calls a rule without a runnable gate half a change. The gate over words is a Rego module: which literal phrases each file must still carry, which file a finding points at, and the could-not-look arm are the module's own body, and no configuration surface expresses a predicate. Writing the protected path is the only route, and the file lands in PR #833's diff where a reviewer sees it.
Admits-answer-rejected-route: `config read first` does not apply: the committed authority carries the rows that register a module and the literal phrases the predicate matches are not expressible there. `patch run first` does not apply either: the module is hand-authored with no generator behind it, so a patch route would write the same bytes to the same protected path.
wenzowski added a commit that referenced this pull request Sep 3, 2026
…iling-run adjacency

The predicate this branch first landed was a second authority. CLOUD-1347
and CLOUD-1350 already design one module for the whole loop vocabulary,
and `a-repeated-call-is-not-progress` was CLOUD-1350's window-recurrence
detector under another name, computing another fact. Two predicates over
one question can disagree, and the disagreement is discovered by a
session being refused, so the module is renamed onto the declared
vocabulary rather than left beside it.

WHAT REPLACES IT IS NOT A RENAME. `repeated-calls` was the MAXIMUM
recurrences of any identity over the whole stream. `repeat-depth` is the
TRAILING run of identical calls, and `distinct-calls` is the progress
term beside it. Adjacency is what does the false-positive work: any
intervening distinct call clears the run, so the edit-then-retest loop is
false by construction rather than by carve-out, and the threshold is the
3 that opencode, hermes-agent and OpenHands all converge on rather than
a constant derived from one pathological session.

AND ADJACENCY IS WHY THIS ONE CANNOT LOCK A SESSION OUT. A whole-stream
maximum is monotonic: once it crossed its threshold it stayed crossed for
the rest of the session, which is how the earlier predicate refused every
subsequent tool call including its own author's push, with no route to an
admission — `batten override request` and `spend` are themselves mediated
calls. A trailing run resets on the next distinct call, so the escape is
automatic and the override CLOUD-1352 names is actually reachable. Every
member of this family is non-monotonic for that reason.

The fingerprint stays inside `transcript.rs`, hashed and dropped in the
same expression as `Event::HookOutput`'s digest, so only a run length is
projected and `Extraction` stays integers-only. A replayed `tool_use` id
is deduped: counting replays measures what the host chose to re-emit
rather than what the session did.

Severity stays `warn`. CLOUD-1352 owns the promotion and makes a measured
firing rate over this repository's own history a hard precondition.

Refs: CLOUD-1347, CLOUD-1341, CLOUD-1350, CLOUD-1352, CLOUD-1337

Admits: 04117a36b963f590a81de944785fade06bd0eb7b27cd2b35f33d59fa3bae4b07
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: batten.toml
Admits-head: 6420fb9
Admits-epoch: 6f8c3ea4c54e6263a8ca91ed3f034ea58de55f4324b3d93f7737b467fcc2467d
Admits-author: alec@wenzowski.com
Admits-prev: 3fe3ee98cc0ec6d5f1a9ea595f823c2daa76332568b0ca2332363539c5979d99
Admits-answer-lost: The module cannot load, so CLOUD-1341 ships nothing. An unregistered `policy/*.rego` file is inert — the engine reads its rule set from the committed authority, and a module with no registering row is never evaluated. The extractor row is what makes the count reach the predicate at all: without it the fact is undefined, Rego reads undefined as does-not-hold, and the gate loads clean while deciding nothing, which is the dead-gate class `.claude/rules/policy-modules.md` exists to warn about.
Admits-answer-precondition: Registering a policy module IS an edit to the committed authority: a `[[rule]]` row naming the module, a `[[rule.extract]]` row binding the new `repeated-calls` extractor to the key the module reads, and the `[[verdict]]` plus `[[verdict.route]]` rows declaring the `turn ask twice` token. A module raising an undeclared token fails to load, and a declared verdict row nothing raises fails the load too, so the rows and the module are one indivisible change. There is no other surface that expresses them, and the write lands in PR #833's diff where a reviewer sees it.
Admits-answer-rejected-route: `config read first` is what this IS — that route names reading configuration before writing it, and the rows being added do not exist to be read, because the module is new. `patch run first` does not apply: the file is hand-authored and has no generator behind it, so a patch route would write the same bytes to the same protected path; the generated `schema/*.json` are derived FROM the crate rather than producing it.

Admits: 368813d8fbd4cfef084b59aaa4b102fdf06d6bdb3d21ceaf0df87cc61629d2a8
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: policy/repetition-without-progress.rego
Admits-head: 6420fb9
Admits-epoch: 6f8c3ea4c54e6263a8ca91ed3f034ea58de55f4324b3d93f7737b467fcc2467d
Admits-author: alec@wenzowski.com
Admits-prev: -
Admits-answer-lost: The branch ships a second authority over a question the tracker already owns. A family of rows designs one module holding the whole loop vocabulary; landing a differently-named module computing a differently-named fact means whoever builds that family writes the second one, and this repository refuses exactly that — two predicates over one question can disagree, and the disagreement is discovered by a session being refused. Without the restructure the duplicate is what lands.
Admits-answer-precondition: The predicate is a Rego module: the threshold constant, the three-valued posture guard, the null guard and the ten test rules are the module's own body, and no configuration surface expresses a predicate. This replaces a module landed earlier on this same branch under a name that duplicated an already-designed family, so the write is a restructure onto the declared vocabulary rather than a new gate. It lands in PR #833's diff where a reviewer sees it.
Admits-answer-rejected-route: `config read first` does not apply: the committed authority has no surface for a predicate body or a threshold constant, and a threshold spelled as a pattern row is refused outright, so the number has to live in the consumer module. `patch run first` does not apply either: the module is hand-authored with no generator behind it, so a patch route would write the same bytes to the same protected path.

Admits: 8a3ff5e4351e38ed26fdc4fd7eeaa30b80bd7d096ab7ebb2433024111f576156
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: policy/a-repeated-call-is-not-progress.rego
Admits-head: 6420fb9
Admits-epoch: 6f8c3ea4c54e6263a8ca91ed3f034ea58de55f4324b3d93f7737b467fcc2467d
Admits-author: alec@wenzowski.com
Admits-prev: 247eaff38bfbc59e1afae9cae684be633d05c8eca2d901051cad86757508560c
Admits-answer-lost: Two modules over one question ship together. The registering rows can be pointed at the replacement, but leaving the file behind leaves a second authority in the tree that a later reader may register again, and the two can disagree over exactly the cases neither author had in mind. Keeping it would also leave its rule id in the mutation census naming a gate nothing registers.
Admits-answer-precondition: The module being removed was landed earlier on this same branch and duplicates an already-designed family's detector under a different name. Its replacement lands in the same change on the declared path, so the removal is half of one restructure rather than a deletion of coverage — every predicate it carried survives, renamed onto the vocabulary the family declares. There is no surface that retires a module except removing the file and its registering rows.
Admits-answer-rejected-route: `config read first` does not apply: the committed authority registers modules and cannot delete one, and pointing the rows elsewhere is what leaves the orphan this removal exists to prevent. `patch run first` does not apply either: there is no generator behind the module, so no patch route reaches it.
wenzowski added a commit that referenced this pull request Sep 3, 2026
…f answering zero

Four of five `Extraction` members reached no consumer, and repetition —
the one thing a doom loop is made of — had no member at all. This is the
spine the predicates stand on, landed alone because the risk is the
plumbing rather than the detection.

`Extraction::of` widens from `&Counts` to `&Stream`. `Counts` is
unchanged: it is the `-J` capability report's own shape and its comments
refuse members no reader of that document needs, so runs live in
`Stream::repeats()` beside `Stream::counts()` rather than as fields on it.

THE LOAD-BEARING HALF IS THE PER-EXTRACTION CAPABILITY. Zero is a real
answer and means the extractor ran. An extraction whose underlying event
kind never appears is not a session that did none of it — it is a host
that does not record it, and answering zero there is a false green over a
session nobody measured. `of` returns `Option<usize>` and the projection
omits the key, so a module reads undefined and Rego takes that as
does-not-hold. Decided in the engine: a per-module conjunct asking whether
this host records turns is a dead gate on every harness but the one its
author tested.

The bound is stated rather than absorbed: this cannot separate a host that
records no hook runs from a session that triggered none. It resolves that
toward could-not-look, which is the safe direction — a missing answer is
reported, where a false zero is indistinguishable from a clean session.

`agent-turn-run` is the first member and deliberately the simplest: a
trailing run of assistant turns carrying no tool call, over the
already-typed `Event::Turn`. No hashing, so no argument or result is read
even internally. It maps to OpenHands' monologue detector at 3+.

The honest claim is narrower than loop detection, and the module header
says so. Termination is undecidable and every quantity here is a
monotonically growing count, so there is no ranking function to be had.
What the literature buys is the shape of the declaration: a SET of
extractions with a set of thresholds, supplying an effective bound on a
SUSPECTED feedback path. It does not detect non-termination.

Adjacency also makes the member non-monotonic — one action clears the run
— which is the property any later promotion depends on. CLOUD-894 owns the
firing-rate ceiling and CLOUD-1352 the promotion; this ships at `warn`.

The compiled tier is what proves the capability rather than the author's
arithmetic. It caught a wrong fixture doing it: a `user` record parses as
a turn, so the first cross-compatibility case recorded turns after all and
answered a real zero. A host recording hook runs and no turn boundaries is
the shape the claim is about.

Also retires the two now-spent rows in `policy/harness-declared.json`.
CLOUD-1079 has landed, so the user-level hooks those rows excused are no
longer provisioned and the merged surfaces carry zero hook commands —
`harness-wiring` reported both, which is the gate working rather than
misfiring. Reproduced against this same tree with `--config-from a907565`,
so the finding predates this diff and is not caused by it. Leaving them is
not neutral once their owner has landed: they would silently excuse
anything that later matched either pattern by name.

Refs: CLOUD-1344, CLOUD-1079, CLOUD-1049, CLOUD-418, CLOUD-894

Admits: d0ac8a4640f2a747b01bedc8449254f45d80865a7de7f0d3f765c40975627f93
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: batten.toml
Admits-head: aed8b44
Admits-epoch: 3ce215f766f4c55fa0be890ef01ffdb5bfc93d7e597090d5f2098d074c5f389f
Admits-author: alec@wenzowski.com
Admits-prev: 04117a36b963f590a81de944785fade06bd0eb7b27cd2b35f33d59fa3bae4b07
Admits-answer-lost: The module cannot load, so CLOUD-1341 ships nothing. An unregistered `policy/*.rego` file is inert — the engine reads its rule set from the committed authority, and a module with no registering row is never evaluated. The extractor row is what makes the count reach the predicate at all: without it the fact is undefined, Rego reads undefined as does-not-hold, and the gate loads clean while deciding nothing, which is the dead-gate class `.claude/rules/policy-modules.md` exists to warn about.
Admits-answer-precondition: Registering a policy module IS an edit to the committed authority: a `[[rule]]` row naming the module, a `[[rule.extract]]` row binding the new `repeated-calls` extractor to the key the module reads, and the `[[verdict]]` plus `[[verdict.route]]` rows declaring the `turn ask twice` token. A module raising an undeclared token fails to load, and a declared verdict row nothing raises fails the load too, so the rows and the module are one indivisible change. There is no other surface that expresses them, and the write lands in PR #833's diff where a reviewer sees it.
Admits-answer-rejected-route: `config read first` is what this IS — that route names reading configuration before writing it, and the rows being added do not exist to be read, because the module is new. `patch run first` does not apply: the file is hand-authored and has no generator behind it, so a patch route would write the same bytes to the same protected path; the generated `schema/*.json` are derived FROM the crate rather than producing it.

Admits: 67106f144c06ca5c3aeaae3482a922741375d1168c3b3a6016979aadbe61b4a4
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: policy/repetition-without-progress.rego
Admits-head: aed8b44
Admits-epoch: 3ce215f766f4c55fa0be890ef01ffdb5bfc93d7e597090d5f2098d074c5f389f
Admits-author: alec@wenzowski.com
Admits-prev: 368813d8fbd4cfef084b59aaa4b102fdf06d6bdb3d21ceaf0df87cc61629d2a8
Admits-answer-lost: The spine row ships with no predicate reading it, which is the half-change non-negotiable rule 2 refuses: an extraction nothing consumes is exactly the defect CLOUD-1344 was filed about, since four of five existing members already reach no consumer. The branch would add a sixth unread member while claiming to fix that.
Admits-answer-precondition: CLOUD-1344's predicate is a Rego module: the adopted threshold, the null guard and the seven test rules are the module's own body, and no configuration surface expresses a predicate. This replaces the predicate landed earlier on this same branch, which was sequenced ahead of the spine row it is blocked on, so the write is a re-sequencing onto the row that must land first. It lands in PR #833's diff where a reviewer sees it.
Admits-answer-rejected-route: `config read first` does not apply: the committed authority has no surface for a predicate body or a threshold, and a threshold spelled as a pattern row is refused outright. `patch run first` does not apply either: the module is hand-authored with no generator behind it, so a patch route would write the same bytes to the same protected path.

Admits: 06de726f4ec787c03543081e96d8e7ae1d46932d5306fe7b68e4465d56f8db24
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: policy/harness-declared.json
Admits-head: aed8b44
Admits-epoch: 6785dbe2aaf01f887e2ec5902922351fd259cb877a4e4d8c830aeefa1d3079ba
Admits-author: alec@wenzowski.com
Admits-prev: -
Admits-answer-lost: The gate stays red on every commit for a reason nobody caused, which is how a correct refusal gets skipped per commit until it is skipped by habit. Worse, the rows keep excusing two commands by name: were anything to reintroduce a hook matching either pattern, the exemption would silently cover it, so leaving them is not neutral once their owner has landed.
Admits-answer-precondition: The exemption table is a policy data file the engine reads, and the two rows in it are now spent: the issue that owns them has landed, the user-level hooks they excused are no longer provisioned, and the merged surfaces carry zero hook commands. `harness-wiring` reports both, which is the gate working rather than misfiring. Retiring a row means editing that file; there is no other surface for it, and the write lands in PR #833's diff where a reviewer sees it.
Admits-answer-rejected-route: `config read first` does not apply: the committed authority declares which external surfaces are read and holds no exemption rows, which is deliberate — an exemption table is data a gate reads rather than part of the gate. `patch run first` does not apply either: the file is hand-authored with no generator behind it, so a patch route would write the same bytes to the same protected path.
@wenzowski wenzowski changed the title fix(ci): cache the perf base arm on a path that holds still across a rebase fix(ci): stop writing cache no run can read, and land the loop family's spine Sep 3, 2026
wenzowski added a commit that referenced this pull request Sep 3, 2026
…not a total

An across-turn poll was invisible to every gate. Measured over one real
transcript: `ReadNotifications` called 1079 times, 59% of every tool call
in the session, every one with identical arguments and the identical "No
queued notifications" result. All 1079 carry `role: assistant`, so this
judges agent conduct and cannot fire on harness behaviour.

No landed arm could reach it. `run-shape` keys on a backgrounded `sleep`
and CLOUD-489 on `until`/`while` inside one command string; repetition
spread across turns carries neither token, and a harness verb with no
argv gives a `shape` row nothing to match.

THE FACT IS A MAXIMUM OVER IDENTITIES, NEVER A SUM. `repeated_calls` is
how far the session's most-repeated (tool, arguments) identity ran. A
running total across every identity is monotonic in the length of the
session and never resets, so it fires on anything long enough to repeat
anything: over the same transcript the sum was 1294 where the max was
1079 against a healthy ceiling of 38. A threshold derived from one
identity's recurrence and applied to that total is not a threshold.

The row's own §2 is refuted and needs rewriting on the tracker. It
proposed "consecutive calls to one tool": the polling arrived in 249
bursts whose longest run was 8, a length any healthy session reaches, so
a consecutive reading ships as coverage while deciding nothing.

Identity is the tool name and a DIGEST of the arguments, so no argument
text is retained. The RESULT is deliberately excluded — it separates
better (60x against 28x) and would oblige the parser to read every
result body, the one payload `transcript.rs` exists never to touch. A
replayed `tool_use` id is deduped: counting replays measures what the
host chose to re-emit rather than what the session did.

IT SHIPS AT `warn`. A `mediated_call` row at `deny` refuses every later
tool call once it fires, and no admission can clear it — `batten
override request` and `spend` are themselves mediated calls, so
requesting one requires making the call the deny refuses. This predicate
was landed at `deny` once and locked its own authoring session out at
~1300, push included. Promotion needs it shown silent against a real
transcript first.

The compiled tier drives the engine rather than a fabricated input,
which is what pins the off-by-one: N calls are N-1 recurrences, so
clearing 100 takes 102 calls and 101 is clean.

Refs: CLOUD-1341, CLOUD-1172, CLOUD-418, CLOUD-489

Admits: 3d6283f7ae44fe7b8349bbb4ef3aa0b22e03bc84ad4059098ff15576b3729337
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: batten.toml
Admits-head: a907565
Admits-epoch: 93d8b1112718944856b6e81b9cfac7c903dbb2c28470be1fc6e59786fb42f404
Admits-author: alec@wenzowski.com
Admits-prev: a387bad2ea98589d0f75455b02cba853248836b7c51d4707cf39585329262b06
Admits-answer-lost: The module cannot load, so CLOUD-1341 ships nothing. An unregistered `policy/*.rego` file is inert — the engine reads its rule set from the committed authority, and a module with no registering row is never evaluated. The extractor row is what makes the count reach the predicate at all: without it the fact is undefined, Rego reads undefined as does-not-hold, and the gate loads clean while deciding nothing, which is the dead-gate class `.claude/rules/policy-modules.md` exists to warn about.
Admits-answer-precondition: Registering a policy module IS an edit to the committed authority: a `[[rule]]` row naming the module, a `[[rule.extract]]` row binding the new `repeated-calls` extractor to the key the module reads, and the `[[verdict]]` plus `[[verdict.route]]` rows declaring the `turn ask twice` token. A module raising an undeclared token fails to load, and a declared verdict row nothing raises fails the load too, so the rows and the module are one indivisible change. There is no other surface that expresses them, and the write lands in PR #833's diff where a reviewer sees it.
Admits-answer-rejected-route: `config read first` is what this IS — that route names reading configuration before writing it, and the rows being added do not exist to be read, because the module is new. `patch run first` does not apply: the file is hand-authored and has no generator behind it, so a patch route would write the same bytes to the same protected path; the generated `schema/*.json` are derived FROM the crate rather than producing it.

Admits: 247eaff38bfbc59e1afae9cae684be633d05c8eca2d901051cad86757508560c
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: policy/a-repeated-call-is-not-progress.rego
Admits-head: a907565
Admits-epoch: 93d8b1112718944856b6e81b9cfac7c903dbb2c28470be1fc6e59786fb42f404
Admits-author: alec@wenzowski.com
Admits-prev: e4dad46c2f5408cdb8bdb3487a98b9e7035341caad7cb39e7e1d555605d0d1f9
Admits-answer-lost: CLOUD-1341 stays prose. The always-loaded instructions already state that a backgrounded task's exit notification is the wake-up and that re-asking is not waiting, and non-negotiable rule 2 calls a rule without a runnable gate half a change. The two landed arms cannot reach this family: one keys on a backgrounded `sleep`, the other on `until`/`while` inside a single command string, while the measured defect was 1079 identical harness-verb calls spread across turns with no argv to match. Without the module there is no gate at all.
Admits-answer-precondition: CLOUD-1341's remedy is a policy predicate, and a predicate is a Rego module: the threshold constant, the null guard and the eight test rules are the module's own body, which no configuration surface can express. The committed authority carries only the rows that register it. Writing the protected path is the only route left, and the file lands in PR #833's diff where a reviewer sees it.
Admits-answer-rejected-route: `config read first` does not apply: the committed authority has no surface for a predicate body or a threshold constant, and `.claude/rules/policy-modules.md` refuses a threshold spelled as a pattern row outright, so the number has to live in the consumer module. `patch run first` does not apply either: the module is hand-authored with no generator behind it, so a patch route would write the same bytes to the same protected path.
@wenzowski
wenzowski force-pushed the claude/cloud-1331-perf-base-arm-6pgmy2 branch from 7dfcc36 to 889f97b Compare September 3, 2026 02:21
wenzowski added a commit that referenced this pull request Sep 3, 2026
…a pull request

CLOUD-1331's keyed base arm collected nothing across 4 runs and 2 pull
requests: 0 hits, ~190 MB written and discarded each time. The first
diagnosis blamed the key; the second blamed the path. Both were reading
the same design, which optimised for a hit and never asked what a miss
costs.

An entry written from a pull request is scoped to `refs/pull/N/merge`
and no other pull request can read it. An entry written from the default
branch is readable by every branch. So the PR side now RESTORES and
never saves — `actions/cache/restore`, a key carrying no merge base
because nothing saves and there is no miss to engineer — and the
scheduled `perf` job on `main`, which already pays a full `--release`
build, stages that closure and saves it under the same key. No new
trigger, no second build.

The seed is never measured: it is copied to `target/perf/base-$SHA` with
its binary removed, so `base_arm_is_built()` stays false and cargo runs
against the base tree actually materialised.

CLOUD-840 measured this repository at 225.2 GB of Actions cache, 203.7 GB
of it across 1,144 PR-ref entries nothing can restore. This job now
writes zero of them, and the hit arrives on a pull request's FIRST run
rather than its second.

`cache-path-is-rebase-stable` gains the two sub-action spellings.
`actions/cache/restore` and `actions/cache/save` derive an entry's
version from `path` exactly as the composite does, so matching only
`actions/cache@` would have left the predicate live and reaching nothing
the moment a job split restore from save — which is what this commit
does. Its own case is declared rather than assumed.

Refs: CLOUD-1342, CLOUD-1331, CLOUD-840

Admits: b0862b76d140be6746e8cdd0fcaf3c28e5bc685ef40c4e3164a1079bf12d075b
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: .github/workflows/ci.yml
Admits-head: fb2987d
Admits-epoch: 93d8b1112718944856b6e81b9cfac7c903dbb2c28470be1fc6e59786fb42f404
Admits-author: alec@wenzowski.com
Admits-prev: f0a932f52d256fb53185a39552efa1b6ac28551baa5b3c020412742000541260
Admits-answer-lost: The job keeps writing ~190 MB per run of cache scoped to the pull request's own merge ref, which no later run can ever restore — CLOUD-840 measured 203.7 GB of such entries across 1,144 of them in this repository — and every pull request keeps compiling the base arm's whole release closure cold, 239 crates in 4m42s of a 13.5 min step. The design being replaced optimised for a hit and never asked what a miss costs.
Admits-answer-precondition: The change is a GitHub Actions step in the `perf` job — swapping the composite cache action for its restore-only sub-action, dropping the merge-base SHA from the key, and deleting the pull-request-side save step. No batten surface expresses a workflow step, so writing the protected path directly is the only route left, and the write lands in PR #833's diff where a reviewer sees it.
Admits-answer-rejected-route: `config read first` does not apply: the committed authority declares rules, verdicts and patterns and carries no representation of a workflow job or step, so there is nothing to read or set there. `patch run first` does not apply either: the workflow is hand-authored YAML with no generator behind it, so a patch route would write the same bytes to the same protected path.

Admits: a76c9a87c95fa32cf60a4b398c75005fdba89cb88255963faf88c77be092502f
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: .github/workflows/perf.yml
Admits-head: fb2987d
Admits-epoch: 93d8b1112718944856b6e81b9cfac7c903dbb2c28470be1fc6e59786fb42f404
Admits-author: alec@wenzowski.com
Admits-prev: 22dc584248885224d59327013514c1c438affb90392824a98c742447a18e3b7e
Admits-answer-lost: The restore-only step this branch lands has nothing to restore. Only a run on the default branch can write a cache entry every branch can read, and this scheduled job is the one place that already pays a full release build. Absent the seed, every pull request keeps compiling the base arm's whole release closure cold, and writing unreadable per-pull-request entries stays the only alternative.
Admits-answer-precondition: The change adds two GitHub Actions steps to the scheduled job that already runs on the default branch — staging the release closure it has just built, and saving it under the key the pull-request side restores. No batten surface expresses a workflow step, so writing the protected path directly is the only route left, and the write lands in PR #833's diff where a reviewer sees it.
Admits-answer-rejected-route: `config read first` does not apply: the committed authority declares rules, verdicts and patterns and carries no representation of a workflow job or step, so there is nothing to read or set there. `patch run first` does not apply either: the workflow is hand-authored YAML with no generator behind it, so a patch route would write the same bytes to the same protected path.

Admits: 7d1ad5bf1d04296f330e798f42f2a74aba2e5d7f78ab4b9a8d2e2474c9dea5e3
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: policy/ci-parity.rego
Admits-head: fb2987d
Admits-epoch: 93d8b1112718944856b6e81b9cfac7c903dbb2c28470be1fc6e59786fb42f404
Admits-author: alec@wenzowski.com
Admits-prev: d2e45c5830908bdc08f6f59141f3a9b031f68ad27ffdd7d133ebf108d0b30921
Admits-answer-lost: The gate goes dead against its own subject. The sub-actions derive an entry's version from the cached path exactly as the composite does, so an interpolated path is the identical silent total miss — the one measured at 4 runs, 0 hits, ~190 MB discarded each time — and after this branch the repository's only cache steps of that kind are the sub-action spellings the predicate cannot see. A gate that loads clean and matches nothing reads exactly like a clean tree.
Admits-answer-precondition: The predicate matches only the composite spelling of the cache action, and this branch moves the job to the restore-only sub-action plus a save on the scheduled side — so the gate that exists to refuse an interpolated cached path would stop covering the very steps this branch lands. Widening it and adding the discriminating case is a change to the Rego module itself; no configuration surface expresses a predicate body, so writing the protected path directly is the only route, and it lands in PR #833's diff where a reviewer sees it.
Admits-answer-rejected-route: `config read first` does not apply: the committed authority carries the rows that register this module, not the predicate body, so nothing there can widen which action spellings the rule matches. `patch run first` does not apply either: the module is hand-authored and has no generator behind it, so a patch route would write the same bytes to the same protected path.
wenzowski added a commit that referenced this pull request Sep 3, 2026
…viting the second claim

The claim receipt is minted on `claim check`'s pullable path and keyed by
the branch checked out at that moment. That one fact makes two orderings
fail in opposite directions, and both are reachable by following the
instructions correctly the first time.

Claim, then branch: the receipt is minted against the branch you were
standing on, and the first edit on the real branch is refused with
`receipt read missing claim branch claim-needs-receipt`.

Claim twice: having branched first, the pullable message read as
go-do-this-and-come-back, so the next move was to move the row and run
`claim-check` again. The second run arrives after the row has left Todo,
reads it as held, and refuses `not-todo` — where the holder is the caller
ninety seconds earlier. The only route past is `--takeover` against
oneself, which writes a takeover record for a row nobody else touched,
and CLOUD-1139 needs that signal to stay rare. Measured twice, in two
sessions, both by an agent following the documented order.

NO ENGINE CHANGE, and the first revision of this row assumed one was
needed. Creating a branch writes only under `.git`, which
`claim-needs-receipt` never judges, so there is no ordering deadlock —
the deadlock was imagined. `claim.rs`'s `not-todo` decision is correct as
it stands for a row genuinely held elsewhere and is untouched.

Recognising a self-claim in the engine stays CLOUD-1139's: it needs to
tell this session from a sibling, and the row's assignee cannot, since
every fleet session carries the same configured accountable identity. An
assignee-keyed re-mint would let a sibling silently re-mint over a working
holder — that row's measured harm, reintroduced through this row's fix.

The gate decides what a gate can: whether the always-loaded file still
states the order and whether the triggered file still carries both failure
directions. Whether a given session actually claimed before branching is
not a property of the tree, and a rule resolving to it would be the model
verdict non-negotiable rule 3 forbids.

Two files because a budget forced it. `policy-budget` caps the index at
3500 tokens and 199 lines, and an earlier draft of this change blew both
at 3516/202; the index carries only the ORDER and the reason lives in the
rules file that loads at the trigger. Both arms are therefore required.

The compiled tier is what proves the two markdown files reach the module
at all: a dead gate and a tree that still states the order are
byte-identical on the decision surface, so each drift case can only go red
if the lines actually arrived.

Refs: CLOUD-1343, CLOUD-1139, CLOUD-786, CLOUD-733

Admits: 3fe3ee98cc0ec6d5f1a9ea595f823c2daa76332568b0ca2332363539c5979d99
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: batten.toml
Admits-head: f9f48e6
Admits-epoch: 84237be51e176ca21fe43fbe5f50492b50c188b8881b9745dfb2b4004a75a220
Admits-author: alec@wenzowski.com
Admits-prev: 3d6283f7ae44fe7b8349bbb4ef3aa0b22e03bc84ad4059098ff15576b3729337
Admits-answer-lost: The gate is dead rather than absent, which is strictly worse. Without the `line_sources` list the engine builds no lines for those two paths, every clause reads undefined, Rego takes undefined as does-not-hold, and the module loads clean while deciding nothing — indistinguishable on the decision surface from a tree that still states the order. That is the exact failure class `.claude/rules/policy-modules.md` records, and without the registering row the module is never evaluated at all.
Admits-answer-precondition: Registering the claim-order module IS an edit to the committed authority: a tree-scoped `[[rule]]` row naming it, the `line_sources` list that makes the two instruction files reach the predicate at all, and the `[[verdict]]` plus `[[verdict.route]]` rows declaring the `claim declare dropped` token. A module raising an undeclared token fails to load, and a declared verdict row nothing raises fails the load too, so the rows and the module are one indivisible change. There is no other surface that expresses them, and the write lands in PR #833's diff where a reviewer sees it.
Admits-answer-rejected-route: `config read first` is what this IS — that route names reading configuration before writing it, and the rows being added do not exist to be read, because the module is new. `patch run first` does not apply: the file is hand-authored and has no generator behind it, so a patch route would write the same bytes to the same protected path.

Admits: f4a603586d8e7912a35d4358fe20ed1955c112b132ecbaf9ad46eebc2bfa1742
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: policy/claim-order-is-stated.rego
Admits-head: f9f48e6
Admits-epoch: 84237be51e176ca21fe43fbe5f50492b50c188b8881b9745dfb2b4004a75a220
Admits-author: alec@wenzowski.com
Admits-prev: -
Admits-answer-lost: The order evaporates. The failure mode here is DRIFT — the remedy is prose, and prose is what a later edit silently rewords or drops, after which the next session walks into the same refusal that cost this one a full stop mid-work and a takeover record against its own claim. A prose assertion with no gate is exactly the half-change rule 2 refuses, and the Ready gate already refused an earlier revision of this row for carrying an empty tests array.
Admits-answer-precondition: CLOUD-1343's remedy is words in two instruction files, and non-negotiable rule 2 calls a rule without a runnable gate half a change. The gate over words is a Rego module: which literal phrases each file must still carry, which file a finding points at, and the could-not-look arm are the module's own body, and no configuration surface expresses a predicate. Writing the protected path is the only route, and the file lands in PR #833's diff where a reviewer sees it.
Admits-answer-rejected-route: `config read first` does not apply: the committed authority carries the rows that register a module and the literal phrases the predicate matches are not expressible there. `patch run first` does not apply either: the module is hand-authored with no generator behind it, so a patch route would write the same bytes to the same protected path.
wenzowski added a commit that referenced this pull request Sep 3, 2026
…iling-run adjacency

The predicate this branch first landed was a second authority. CLOUD-1347
and CLOUD-1350 already design one module for the whole loop vocabulary,
and `a-repeated-call-is-not-progress` was CLOUD-1350's window-recurrence
detector under another name, computing another fact. Two predicates over
one question can disagree, and the disagreement is discovered by a
session being refused, so the module is renamed onto the declared
vocabulary rather than left beside it.

WHAT REPLACES IT IS NOT A RENAME. `repeated-calls` was the MAXIMUM
recurrences of any identity over the whole stream. `repeat-depth` is the
TRAILING run of identical calls, and `distinct-calls` is the progress
term beside it. Adjacency is what does the false-positive work: any
intervening distinct call clears the run, so the edit-then-retest loop is
false by construction rather than by carve-out, and the threshold is the
3 that opencode, hermes-agent and OpenHands all converge on rather than
a constant derived from one pathological session.

AND ADJACENCY IS WHY THIS ONE CANNOT LOCK A SESSION OUT. A whole-stream
maximum is monotonic: once it crossed its threshold it stayed crossed for
the rest of the session, which is how the earlier predicate refused every
subsequent tool call including its own author's push, with no route to an
admission — `batten override request` and `spend` are themselves mediated
calls. A trailing run resets on the next distinct call, so the escape is
automatic and the override CLOUD-1352 names is actually reachable. Every
member of this family is non-monotonic for that reason.

The fingerprint stays inside `transcript.rs`, hashed and dropped in the
same expression as `Event::HookOutput`'s digest, so only a run length is
projected and `Extraction` stays integers-only. A replayed `tool_use` id
is deduped: counting replays measures what the host chose to re-emit
rather than what the session did.

Severity stays `warn`. CLOUD-1352 owns the promotion and makes a measured
firing rate over this repository's own history a hard precondition.

Refs: CLOUD-1347, CLOUD-1341, CLOUD-1350, CLOUD-1352, CLOUD-1337

Admits: 04117a36b963f590a81de944785fade06bd0eb7b27cd2b35f33d59fa3bae4b07
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: batten.toml
Admits-head: 6420fb9
Admits-epoch: 6f8c3ea4c54e6263a8ca91ed3f034ea58de55f4324b3d93f7737b467fcc2467d
Admits-author: alec@wenzowski.com
Admits-prev: 3fe3ee98cc0ec6d5f1a9ea595f823c2daa76332568b0ca2332363539c5979d99
Admits-answer-lost: The module cannot load, so CLOUD-1341 ships nothing. An unregistered `policy/*.rego` file is inert — the engine reads its rule set from the committed authority, and a module with no registering row is never evaluated. The extractor row is what makes the count reach the predicate at all: without it the fact is undefined, Rego reads undefined as does-not-hold, and the gate loads clean while deciding nothing, which is the dead-gate class `.claude/rules/policy-modules.md` exists to warn about.
Admits-answer-precondition: Registering a policy module IS an edit to the committed authority: a `[[rule]]` row naming the module, a `[[rule.extract]]` row binding the new `repeated-calls` extractor to the key the module reads, and the `[[verdict]]` plus `[[verdict.route]]` rows declaring the `turn ask twice` token. A module raising an undeclared token fails to load, and a declared verdict row nothing raises fails the load too, so the rows and the module are one indivisible change. There is no other surface that expresses them, and the write lands in PR #833's diff where a reviewer sees it.
Admits-answer-rejected-route: `config read first` is what this IS — that route names reading configuration before writing it, and the rows being added do not exist to be read, because the module is new. `patch run first` does not apply: the file is hand-authored and has no generator behind it, so a patch route would write the same bytes to the same protected path; the generated `schema/*.json` are derived FROM the crate rather than producing it.

Admits: 368813d8fbd4cfef084b59aaa4b102fdf06d6bdb3d21ceaf0df87cc61629d2a8
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: policy/repetition-without-progress.rego
Admits-head: 6420fb9
Admits-epoch: 6f8c3ea4c54e6263a8ca91ed3f034ea58de55f4324b3d93f7737b467fcc2467d
Admits-author: alec@wenzowski.com
Admits-prev: -
Admits-answer-lost: The branch ships a second authority over a question the tracker already owns. A family of rows designs one module holding the whole loop vocabulary; landing a differently-named module computing a differently-named fact means whoever builds that family writes the second one, and this repository refuses exactly that — two predicates over one question can disagree, and the disagreement is discovered by a session being refused. Without the restructure the duplicate is what lands.
Admits-answer-precondition: The predicate is a Rego module: the threshold constant, the three-valued posture guard, the null guard and the ten test rules are the module's own body, and no configuration surface expresses a predicate. This replaces a module landed earlier on this same branch under a name that duplicated an already-designed family, so the write is a restructure onto the declared vocabulary rather than a new gate. It lands in PR #833's diff where a reviewer sees it.
Admits-answer-rejected-route: `config read first` does not apply: the committed authority has no surface for a predicate body or a threshold constant, and a threshold spelled as a pattern row is refused outright, so the number has to live in the consumer module. `patch run first` does not apply either: the module is hand-authored with no generator behind it, so a patch route would write the same bytes to the same protected path.

Admits: 8a3ff5e4351e38ed26fdc4fd7eeaa30b80bd7d096ab7ebb2433024111f576156
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: policy/a-repeated-call-is-not-progress.rego
Admits-head: 6420fb9
Admits-epoch: 6f8c3ea4c54e6263a8ca91ed3f034ea58de55f4324b3d93f7737b467fcc2467d
Admits-author: alec@wenzowski.com
Admits-prev: 247eaff38bfbc59e1afae9cae684be633d05c8eca2d901051cad86757508560c
Admits-answer-lost: Two modules over one question ship together. The registering rows can be pointed at the replacement, but leaving the file behind leaves a second authority in the tree that a later reader may register again, and the two can disagree over exactly the cases neither author had in mind. Keeping it would also leave its rule id in the mutation census naming a gate nothing registers.
Admits-answer-precondition: The module being removed was landed earlier on this same branch and duplicates an already-designed family's detector under a different name. Its replacement lands in the same change on the declared path, so the removal is half of one restructure rather than a deletion of coverage — every predicate it carried survives, renamed onto the vocabulary the family declares. There is no surface that retires a module except removing the file and its registering rows.
Admits-answer-rejected-route: `config read first` does not apply: the committed authority registers modules and cannot delete one, and pointing the rows elsewhere is what leaves the orphan this removal exists to prevent. `patch run first` does not apply either: there is no generator behind the module, so no patch route reaches it.
wenzowski added a commit that referenced this pull request Sep 3, 2026
…f answering zero

Four of five `Extraction` members reached no consumer, and repetition —
the one thing a doom loop is made of — had no member at all. This is the
spine the predicates stand on, landed alone because the risk is the
plumbing rather than the detection.

`Extraction::of` widens from `&Counts` to `&Stream`. `Counts` is
unchanged: it is the `-J` capability report's own shape and its comments
refuse members no reader of that document needs, so runs live in
`Stream::repeats()` beside `Stream::counts()` rather than as fields on it.

THE LOAD-BEARING HALF IS THE PER-EXTRACTION CAPABILITY. Zero is a real
answer and means the extractor ran. An extraction whose underlying event
kind never appears is not a session that did none of it — it is a host
that does not record it, and answering zero there is a false green over a
session nobody measured. `of` returns `Option<usize>` and the projection
omits the key, so a module reads undefined and Rego takes that as
does-not-hold. Decided in the engine: a per-module conjunct asking whether
this host records turns is a dead gate on every harness but the one its
author tested.

The bound is stated rather than absorbed: this cannot separate a host that
records no hook runs from a session that triggered none. It resolves that
toward could-not-look, which is the safe direction — a missing answer is
reported, where a false zero is indistinguishable from a clean session.

`agent-turn-run` is the first member and deliberately the simplest: a
trailing run of assistant turns carrying no tool call, over the
already-typed `Event::Turn`. No hashing, so no argument or result is read
even internally. It maps to OpenHands' monologue detector at 3+.

The honest claim is narrower than loop detection, and the module header
says so. Termination is undecidable and every quantity here is a
monotonically growing count, so there is no ranking function to be had.
What the literature buys is the shape of the declaration: a SET of
extractions with a set of thresholds, supplying an effective bound on a
SUSPECTED feedback path. It does not detect non-termination.

Adjacency also makes the member non-monotonic — one action clears the run
— which is the property any later promotion depends on. CLOUD-894 owns the
firing-rate ceiling and CLOUD-1352 the promotion; this ships at `warn`.

The compiled tier is what proves the capability rather than the author's
arithmetic. It caught a wrong fixture doing it: a `user` record parses as
a turn, so the first cross-compatibility case recorded turns after all and
answered a real zero. A host recording hook runs and no turn boundaries is
the shape the claim is about.

Also retires the two now-spent rows in `policy/harness-declared.json`.
CLOUD-1079 has landed, so the user-level hooks those rows excused are no
longer provisioned and the merged surfaces carry zero hook commands —
`harness-wiring` reported both, which is the gate working rather than
misfiring. Reproduced against this same tree with `--config-from a907565`,
so the finding predates this diff and is not caused by it. Leaving them is
not neutral once their owner has landed: they would silently excuse
anything that later matched either pattern by name.

Refs: CLOUD-1344, CLOUD-1079, CLOUD-1049, CLOUD-418, CLOUD-894

Admits: d0ac8a4640f2a747b01bedc8449254f45d80865a7de7f0d3f765c40975627f93
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: batten.toml
Admits-head: aed8b44
Admits-epoch: 3ce215f766f4c55fa0be890ef01ffdb5bfc93d7e597090d5f2098d074c5f389f
Admits-author: alec@wenzowski.com
Admits-prev: 04117a36b963f590a81de944785fade06bd0eb7b27cd2b35f33d59fa3bae4b07
Admits-answer-lost: The module cannot load, so CLOUD-1341 ships nothing. An unregistered `policy/*.rego` file is inert — the engine reads its rule set from the committed authority, and a module with no registering row is never evaluated. The extractor row is what makes the count reach the predicate at all: without it the fact is undefined, Rego reads undefined as does-not-hold, and the gate loads clean while deciding nothing, which is the dead-gate class `.claude/rules/policy-modules.md` exists to warn about.
Admits-answer-precondition: Registering a policy module IS an edit to the committed authority: a `[[rule]]` row naming the module, a `[[rule.extract]]` row binding the new `repeated-calls` extractor to the key the module reads, and the `[[verdict]]` plus `[[verdict.route]]` rows declaring the `turn ask twice` token. A module raising an undeclared token fails to load, and a declared verdict row nothing raises fails the load too, so the rows and the module are one indivisible change. There is no other surface that expresses them, and the write lands in PR #833's diff where a reviewer sees it.
Admits-answer-rejected-route: `config read first` is what this IS — that route names reading configuration before writing it, and the rows being added do not exist to be read, because the module is new. `patch run first` does not apply: the file is hand-authored and has no generator behind it, so a patch route would write the same bytes to the same protected path; the generated `schema/*.json` are derived FROM the crate rather than producing it.

Admits: 67106f144c06ca5c3aeaae3482a922741375d1168c3b3a6016979aadbe61b4a4
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: policy/repetition-without-progress.rego
Admits-head: aed8b44
Admits-epoch: 3ce215f766f4c55fa0be890ef01ffdb5bfc93d7e597090d5f2098d074c5f389f
Admits-author: alec@wenzowski.com
Admits-prev: 368813d8fbd4cfef084b59aaa4b102fdf06d6bdb3d21ceaf0df87cc61629d2a8
Admits-answer-lost: The spine row ships with no predicate reading it, which is the half-change non-negotiable rule 2 refuses: an extraction nothing consumes is exactly the defect CLOUD-1344 was filed about, since four of five existing members already reach no consumer. The branch would add a sixth unread member while claiming to fix that.
Admits-answer-precondition: CLOUD-1344's predicate is a Rego module: the adopted threshold, the null guard and the seven test rules are the module's own body, and no configuration surface expresses a predicate. This replaces the predicate landed earlier on this same branch, which was sequenced ahead of the spine row it is blocked on, so the write is a re-sequencing onto the row that must land first. It lands in PR #833's diff where a reviewer sees it.
Admits-answer-rejected-route: `config read first` does not apply: the committed authority has no surface for a predicate body or a threshold, and a threshold spelled as a pattern row is refused outright. `patch run first` does not apply either: the module is hand-authored with no generator behind it, so a patch route would write the same bytes to the same protected path.

Admits: 06de726f4ec787c03543081e96d8e7ae1d46932d5306fe7b68e4465d56f8db24
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: policy/harness-declared.json
Admits-head: aed8b44
Admits-epoch: 6785dbe2aaf01f887e2ec5902922351fd259cb877a4e4d8c830aeefa1d3079ba
Admits-author: alec@wenzowski.com
Admits-prev: -
Admits-answer-lost: The gate stays red on every commit for a reason nobody caused, which is how a correct refusal gets skipped per commit until it is skipped by habit. Worse, the rows keep excusing two commands by name: were anything to reintroduce a hook matching either pattern, the exemption would silently cover it, so leaving them is not neutral once their owner has landed.
Admits-answer-precondition: The exemption table is a policy data file the engine reads, and the two rows in it are now spent: the issue that owns them has landed, the user-level hooks they excused are no longer provisioned, and the merged surfaces carry zero hook commands. `harness-wiring` reports both, which is the gate working rather than misfiring. Retiring a row means editing that file; there is no other surface for it, and the write lands in PR #833's diff where a reviewer sees it.
Admits-answer-rejected-route: `config read first` does not apply: the committed authority declares which external surfaces are read and holds no exemption rows, which is deliberate — an exemption table is data a gate reads rather than part of the gate. `patch run first` does not apply either: the file is hand-authored with no generator behind it, so a patch route would write the same bytes to the same protected path.
wenzowski added a commit that referenced this pull request Sep 3, 2026
`claim-order-is-stated` fired on any repository carrying an instructions
file, so it refused a fixture whose `AGENTS.md` is the single word
`instructions` and which has no `.claude/rules/` surface at all.
`cli::the_committed_repo_config_gates_a_repository` runs the whole
committed ruleset over exactly that tree and asserts its entire output,
so it went red — the foreign-tree failure every tree-scoped row here has
to answer, caught by the one case that was watching for it.

The guard is keyed on the TRIGGERED file rather than the index, and that
is the whole of the reasoning. This row is about a SPLIT: the order in
the always-loaded file, the reason in the file that loads at the trigger.
A tree with no `.claude/rules/toolchain.md` has not made that split and
is answering for nothing here. Guarding on the index instead would be
circular, since an absent index is precisely what one arm exists to
refuse.

Same shape as `hk-fix-selection`'s `governed`, which asks whether the
repository has the config it judges before judging it.

Both tiers gain the case, so the fixture shape is pinned rather than
merely un-refused: `test_an_index_without_the_rules_file_is_not_judged`
and `an_index_without_the_rules_file_is_not_judged`.

Refs: CLOUD-1343

Admits: 03d9f351561fe5a7cbd22d78ad942b3f541a3ac5a92b241d8c327d36becf58b9
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: policy/claim-order-is-stated.rego
Admits-head: 889f97b
Admits-epoch: fb78c0a9b65b102b88a160ecb71c69eb089570a736ea14b96c2ea053177664fd
Admits-author: alec@wenzowski.com
Admits-prev: f4a603586d8e7912a35d4358fe20ed1955c112b132ecbaf9ad46eebc2bfa1742
Admits-answer-lost: The compiled suite stays red on a case that asserts the committed configuration's whole output over a fixture, so the branch cannot land at all. Worse as a shipped property: a consumer adopting this row would have it fire on any repository with an instructions file and no triggered rules file, which is a refusal about a document split that repository never made — the foreign-tree failure every tree-scoped row here has to answer.
Admits-answer-precondition: The module judges any tree carrying an index file, so it fires on a fixture repository whose stub instructions file has never stated the claim order — a foreign tree answering for a split this repository's own instruction surface has. The fix is a not-applicable guard in the predicate body, which no configuration surface expresses; the committed authority registers the module and cannot narrow what it decides. It lands in PR #833's diff where a reviewer sees it.
Admits-answer-rejected-route: `config read first` does not apply: the committed authority declares the module's sources and severity and has no surface for a not-applicable conjunct, which is part of the predicate. `patch run first` does not apply either: the module is hand-authored with no generator behind it, so a patch route would write the same bytes to the same protected path.

Copy link
Copy Markdown
Contributor Author

Handoff — everything is pushed, one gate remains

Written down because the container is ephemeral and the two things below live out of tree, so they die with it.

Where the branch is

Rebased onto current main, tree clean, all commits pushed. verify has never completed end to end on this branch; it has been stopped only by the two items below.

The one remaining gate

batten-check reports 2 sbom-action-unenriched. CLOUD-1342 pins two new sub-actions (actions/cache/restore in ci.yml, actions/cache/save in perf.yml) and both are pkg:github components in the inventory.

The table rows are committed. The rows alone do not clear it — that predicate reads input.tree["tool-verdict"].sbom, a recorded scan, where its sibling sbom-action-unmapped is the one decided from committed text. The remedy is named in the tool read broken verdict class:

mise run sbom && mise run record-sbom

The record lives out of tree and does not survive a container, so a fresh session must re-run both before batten-check can pass, regardless of what is committed. Same for the branch plan — plan-unrecorded cleared only after batten record plan (it reads <id> <status> lines on stdin; pipe them, a heredoc form exits 1 silently).

Things that will bite whoever picks this up

  • mise run install:local after every rebase, before any batten verb. main keeps adding config keys the installed binary does not know ([[perf.exempt]] was the last), and a stale binary then refuses to parse batten.toml — which blocks batten mcp call, override request, and every gate. This cost four rebuilds in one session (CLOUD-1326).
  • $MUTANT_GATES conflicts on every rebase. Resolve by applying that commit's own gate delta to main's list. Do not union the two sides: main deliberately removed perf-gate and perf-compare, and a union silently resurrects them.
  • A commit message needs the full 11-line Admits: block per protected path, exactly as override spend prints it. A four-line summary reads as absent, because judge_admissions recomputes the hash. Capture the spend output before it scrolls.
  • override spend and the write it admits must be separate calls — the hook adjudicates the whole call before anything inside it runs.
  • Protected paths are .serena/memories/**, batten.toml, .github/workflows/**, plus registered policy modules by derivation (CLOUD-1226).

Findings recorded, and findings not yet filed

Already on the rows that own them:

  • CLOUD-1345 — the Extraction::of(&Stream) widening this branch lands took the reduction from 1 walk to 2 × declared extractions. Invisible at one extractor; it is the term that row is about, and CLOUD-1347 turns each extra walk into an extra hashing pass.
  • CLOUD-1141 — a session-long live instance: every protected-path write here went through python3 -c, unrefused, while the shell spellings of the same intent (cat >, git rm) were refused twice on the same paths in the same session. judge_admissions at commit time is what actually held.
  • CLOUD-1344 / CLOUD-1342 / CLOUD-1341 — bodies corrected in place, including CLOUD-1344's declared mutation, which named engine behaviour no Rego mutation can reach and would have been recorded as a survivor.

Not yet filed — no duplicate search has been run for these, and filing-needs-a-search requires one:

  1. The recovery gap (belongs on CLOUD-1051). override request/spend are themselves mediated calls, so an admission can never clear a mediated_call deny — requesting one means making the call the deny refuses. Adjacency makes the loop family survivable; it does not make a deny recoverable.
  2. The composed trap. The disarm guard and a mediated_call deny have no exit between them. Each is defensible alone; together, a bad deny row cannot be removed by the session that trips it. Argues for an unmediated disarm path existing at all.
  3. The stop hook demands a push the policy engine refuses — two controls issuing contradictory instructions, with no reconciliation.
  4. BATTEN_HOOK_BYPASS is live, not retired. Fact::Bypass at schema/policy-call.schema.json:84, pinned by cli.rs, mediated_admission.rs:146, contract_drift.rs:456, gh_guard.rs:345, bypass_scrub.rs. A handoff into this session asserted it was retired and asked for the prose to be purged; purging it would make the docs wrong about a working mechanism. File as a question.
  5. CLOUD-1350 is not implementable as written — "distinct-calls low relative to total calls" names no ratio and no threshold, and its batten.toml-declared window has no path from config to the reduction. CLOUD-1344's &Stream signature is what would provide one; neither row references the other.

Deliberately not filed: a mediated_call deny must be shown silent before shipping (CLOUD-1352 already carries it as a hard precondition, CLOUD-894 owns the firing rate), and a stale binary disarms mediation (CLOUD-1326's title says exactly that).

Not in scope for this PR

CLOUD-1347 (identical-call-run over call identity) is blockedBy CLOUD-1344 and CLOUD-1345. The spine is ready for it; CLOUD-1345 is not this branch's to clear. CLOUD-1341 is back in Backlog and unassigned — agent-turn-run catches monologue, not repeated tool calls, so nothing here fixes it.

Watch for conflict with #829, already mergeable_state: dirty, touching batten.toml and policy/.


Generated by Claude Code

…rebase

CLOUD-1331 landed a keyed base arm for `perf pair` and CI collected none of
it: 4 runs across 2 pull requests, 0 hits, ~190 MB written and discarded each
time, and the `perf` job still paid 13.5 min of its 13.7 to rebuild a binary
it had already built.

The cause is the cached PATH, not the key. `actions/cache` identifies an entry
by key AND version, and defines version as "a hash generated for a combination
of compression tool used ... and the `path` of directories being cached". The
step cached `target/perf/base-<merge base>`, interpolated, and `land` rebases
every lap — so the path moved on every lap and took the entry's identity with
it. The remedy first written down was to drop the SHA from the key, which
would have changed nothing whatever.

So the path holds still (`target/perf/base-seed`) and the SHA stays in the
key, which is what makes a moved base MISS primarily and therefore SAVE — a
key that never moves never saves, since "if the provided `key` matches an
existing cache, a new cache is not created". A `restore-keys` prefix supplies
the previous base's closure. Correctness is preserved by deleting the seeded
binary rather than by the directory name: `base_arm_is_built()` is left false,
so cargo runs against the base tree actually materialised and no renamed old
binary is ever measured. `crates/batten/src/perf.rs` is untouched.

Measured locally with `Cargo.lock` byte-identical across the pair, so the
residue is the cost of the copy and not lockfile churn: cold 239 `Compiling`
in 4m42s, seeded 64 in 2m07s, same base 0. The win is ~2m35s of a 13.5 min
step and is stated as that rather than as the whole problem — the 239->0 path
needs a run whose merge base matches a cached one, which needs cross-PR
scoping and is CLOUD-840's.

Rule 2: it ships with a mechanism. `cache-path-is-rebase-stable` in
`policy/ci-parity.rego` refuses a cached path carrying an expression, and
deliberately does not judge the key, since an expression belongs there. The
declared mutation reddens exactly its case and only it (27 passed, 1 failed).
The predicate is the consumer's rather than the core's because it names this
repository's workflow (non-negotiable rule 1).

Refs: CLOUD-1342, CLOUD-1331, CLOUD-840, CLOUD-1225

Admits: 792a6d9ef6e998176b2123ff38eaa85cf8a24854d3694c554258f364e5e83b19
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: .github/workflows/ci.yml
Admits-head: 116ebc1
Admits-epoch: 6c89a13a27eb648df22dd80a90d058b5dfbb32e883ced257202104a3f598b416
Admits-author: alec@wenzowski.com
Admits-prev: d7c27bde12b7fe89dce932035c38e953b155f388b91bca1a449d43dee5739fe2
Admits-answer-lost: The `perf` job keeps rebuilding the merge base's release binary cold on every CI run: 4 runs across 2 pull requests, 0 cache hits, ~190 MB written and discarded each time, ~2m35s per run that a stable cached path recovers. Without the write the CLOUD-1331 mechanism stays uncollectable and the defect is only documented, not fixed.
Admits-answer-precondition: The `perf` job's cache step lives only in this workflow file; there is no other surface that expresses which path a runner restores, so writing `.github/workflows/ci.yml` directly is the only route to the fix. The change is three steps and a cache `path`/`key`/`restore-keys` triple, all visible in the diff a reviewer reads.
Admits-answer-rejected-route: `config read first` does not apply: no `batten.toml` key selects a workflow cache path — the value is GitHub Actions' own and has no projection in this repository's config. `patch run first` does not apply either: there is no generator or task that emits this workflow, so there is no upstream artifact to patch and re-emit; the file is hand-maintained and is its own source of truth.

Admits: a9a9a523ecb285b98020fd8456634f1c2b427856132965d5588cce74bd1ad070
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: batten.toml
Admits-head: 116ebc1
Admits-epoch: 6c89a13a27eb648df22dd80a90d058b5dfbb32e883ced257202104a3f598b416
Admits-author: alec@wenzowski.com
Admits-prev: -
Admits-answer-lost: Without it the predicate cannot load, so the `perf` cache fix ships as a workflow edit with no gate behind it — exactly the half-change non-negotiable rule 2 refuses, leaving the next author free to reintroduce a merge-base SHA in the cached path and silently lose every hit again.
Admits-answer-precondition: `batten.toml` is this repository's one authority for `[[verdict]]` rows, and a policy module raising a token no row declares fails to LOAD. So the new `cache-path-is-rebase-stable` predicate cannot exist without writing this file; no other surface can declare its verdict class on its behalf. The addition is one `[[verdict]]` row and its route, read in the diff.
Admits-answer-rejected-route: `config read first` is the route being taken rather than rejected in spirit — but as a substitute for the write it does not apply: reading `batten.toml` cannot add a row to it, and the verdict registry has no override surface that admits a token from elsewhere. `patch run first` does not apply: `batten.toml` is hand-maintained policy, emitted by no generator, so there is no upstream artifact to patch.
…not a total

An across-turn poll was invisible to every gate. Measured over one real
transcript: `ReadNotifications` called 1079 times, 59% of every tool call
in the session, every one with identical arguments and the identical "No
queued notifications" result. All 1079 carry `role: assistant`, so this
judges agent conduct and cannot fire on harness behaviour.

No landed arm could reach it. `run-shape` keys on a backgrounded `sleep`
and CLOUD-489 on `until`/`while` inside one command string; repetition
spread across turns carries neither token, and a harness verb with no
argv gives a `shape` row nothing to match.

THE FACT IS A MAXIMUM OVER IDENTITIES, NEVER A SUM. `repeated_calls` is
how far the session's most-repeated (tool, arguments) identity ran. A
running total across every identity is monotonic in the length of the
session and never resets, so it fires on anything long enough to repeat
anything: over the same transcript the sum was 1294 where the max was
1079 against a healthy ceiling of 38. A threshold derived from one
identity's recurrence and applied to that total is not a threshold.

The row's own §2 is refuted and needs rewriting on the tracker. It
proposed "consecutive calls to one tool": the polling arrived in 249
bursts whose longest run was 8, a length any healthy session reaches, so
a consecutive reading ships as coverage while deciding nothing.

Identity is the tool name and a DIGEST of the arguments, so no argument
text is retained. The RESULT is deliberately excluded — it separates
better (60x against 28x) and would oblige the parser to read every
result body, the one payload `transcript.rs` exists never to touch. A
replayed `tool_use` id is deduped: counting replays measures what the
host chose to re-emit rather than what the session did.

IT SHIPS AT `warn`. A `mediated_call` row at `deny` refuses every later
tool call once it fires, and no admission can clear it — `batten
override request` and `spend` are themselves mediated calls, so
requesting one requires making the call the deny refuses. This predicate
was landed at `deny` once and locked its own authoring session out at
~1300, push included. Promotion needs it shown silent against a real
transcript first.

The compiled tier drives the engine rather than a fabricated input,
which is what pins the off-by-one: N calls are N-1 recurrences, so
clearing 100 takes 102 calls and 101 is clean.

Refs: CLOUD-1341, CLOUD-1172, CLOUD-418, CLOUD-489

Admits: 3d6283f7ae44fe7b8349bbb4ef3aa0b22e03bc84ad4059098ff15576b3729337
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: batten.toml
Admits-head: a907565
Admits-epoch: 93d8b1112718944856b6e81b9cfac7c903dbb2c28470be1fc6e59786fb42f404
Admits-author: alec@wenzowski.com
Admits-prev: a387bad2ea98589d0f75455b02cba853248836b7c51d4707cf39585329262b06
Admits-answer-lost: The module cannot load, so CLOUD-1341 ships nothing. An unregistered `policy/*.rego` file is inert — the engine reads its rule set from the committed authority, and a module with no registering row is never evaluated. The extractor row is what makes the count reach the predicate at all: without it the fact is undefined, Rego reads undefined as does-not-hold, and the gate loads clean while deciding nothing, which is the dead-gate class `.claude/rules/policy-modules.md` exists to warn about.
Admits-answer-precondition: Registering a policy module IS an edit to the committed authority: a `[[rule]]` row naming the module, a `[[rule.extract]]` row binding the new `repeated-calls` extractor to the key the module reads, and the `[[verdict]]` plus `[[verdict.route]]` rows declaring the `turn ask twice` token. A module raising an undeclared token fails to load, and a declared verdict row nothing raises fails the load too, so the rows and the module are one indivisible change. There is no other surface that expresses them, and the write lands in PR #833's diff where a reviewer sees it.
Admits-answer-rejected-route: `config read first` is what this IS — that route names reading configuration before writing it, and the rows being added do not exist to be read, because the module is new. `patch run first` does not apply: the file is hand-authored and has no generator behind it, so a patch route would write the same bytes to the same protected path; the generated `schema/*.json` are derived FROM the crate rather than producing it.

Admits: 247eaff38bfbc59e1afae9cae684be633d05c8eca2d901051cad86757508560c
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: policy/a-repeated-call-is-not-progress.rego
Admits-head: a907565
Admits-epoch: 93d8b1112718944856b6e81b9cfac7c903dbb2c28470be1fc6e59786fb42f404
Admits-author: alec@wenzowski.com
Admits-prev: e4dad46c2f5408cdb8bdb3487a98b9e7035341caad7cb39e7e1d555605d0d1f9
Admits-answer-lost: CLOUD-1341 stays prose. The always-loaded instructions already state that a backgrounded task's exit notification is the wake-up and that re-asking is not waiting, and non-negotiable rule 2 calls a rule without a runnable gate half a change. The two landed arms cannot reach this family: one keys on a backgrounded `sleep`, the other on `until`/`while` inside a single command string, while the measured defect was 1079 identical harness-verb calls spread across turns with no argv to match. Without the module there is no gate at all.
Admits-answer-precondition: CLOUD-1341's remedy is a policy predicate, and a predicate is a Rego module: the threshold constant, the null guard and the eight test rules are the module's own body, which no configuration surface can express. The committed authority carries only the rows that register it. Writing the protected path is the only route left, and the file lands in PR #833's diff where a reviewer sees it.
Admits-answer-rejected-route: `config read first` does not apply: the committed authority has no surface for a predicate body or a threshold constant, and `.claude/rules/policy-modules.md` refuses a threshold spelled as a pattern row outright, so the number has to live in the consumer module. `patch run first` does not apply either: the module is hand-authored with no generator behind it, so a patch route would write the same bytes to the same protected path.
…a pull request

CLOUD-1331's keyed base arm collected nothing across 4 runs and 2 pull
requests: 0 hits, ~190 MB written and discarded each time. The first
diagnosis blamed the key; the second blamed the path. Both were reading
the same design, which optimised for a hit and never asked what a miss
costs.

An entry written from a pull request is scoped to `refs/pull/N/merge`
and no other pull request can read it. An entry written from the default
branch is readable by every branch. So the PR side now RESTORES and
never saves — `actions/cache/restore`, a key carrying no merge base
because nothing saves and there is no miss to engineer — and the
scheduled `perf` job on `main`, which already pays a full `--release`
build, stages that closure and saves it under the same key. No new
trigger, no second build.

The seed is never measured: it is copied to `target/perf/base-$SHA` with
its binary removed, so `base_arm_is_built()` stays false and cargo runs
against the base tree actually materialised.

CLOUD-840 measured this repository at 225.2 GB of Actions cache, 203.7 GB
of it across 1,144 PR-ref entries nothing can restore. This job now
writes zero of them, and the hit arrives on a pull request's FIRST run
rather than its second.

`cache-path-is-rebase-stable` gains the two sub-action spellings.
`actions/cache/restore` and `actions/cache/save` derive an entry's
version from `path` exactly as the composite does, so matching only
`actions/cache@` would have left the predicate live and reaching nothing
the moment a job split restore from save — which is what this commit
does. Its own case is declared rather than assumed.

Refs: CLOUD-1342, CLOUD-1331, CLOUD-840

Admits: b0862b76d140be6746e8cdd0fcaf3c28e5bc685ef40c4e3164a1079bf12d075b
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: .github/workflows/ci.yml
Admits-head: fb2987d
Admits-epoch: 93d8b1112718944856b6e81b9cfac7c903dbb2c28470be1fc6e59786fb42f404
Admits-author: alec@wenzowski.com
Admits-prev: f0a932f52d256fb53185a39552efa1b6ac28551baa5b3c020412742000541260
Admits-answer-lost: The job keeps writing ~190 MB per run of cache scoped to the pull request's own merge ref, which no later run can ever restore — CLOUD-840 measured 203.7 GB of such entries across 1,144 of them in this repository — and every pull request keeps compiling the base arm's whole release closure cold, 239 crates in 4m42s of a 13.5 min step. The design being replaced optimised for a hit and never asked what a miss costs.
Admits-answer-precondition: The change is a GitHub Actions step in the `perf` job — swapping the composite cache action for its restore-only sub-action, dropping the merge-base SHA from the key, and deleting the pull-request-side save step. No batten surface expresses a workflow step, so writing the protected path directly is the only route left, and the write lands in PR #833's diff where a reviewer sees it.
Admits-answer-rejected-route: `config read first` does not apply: the committed authority declares rules, verdicts and patterns and carries no representation of a workflow job or step, so there is nothing to read or set there. `patch run first` does not apply either: the workflow is hand-authored YAML with no generator behind it, so a patch route would write the same bytes to the same protected path.

Admits: a76c9a87c95fa32cf60a4b398c75005fdba89cb88255963faf88c77be092502f
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: .github/workflows/perf.yml
Admits-head: fb2987d
Admits-epoch: 93d8b1112718944856b6e81b9cfac7c903dbb2c28470be1fc6e59786fb42f404
Admits-author: alec@wenzowski.com
Admits-prev: 22dc584248885224d59327013514c1c438affb90392824a98c742447a18e3b7e
Admits-answer-lost: The restore-only step this branch lands has nothing to restore. Only a run on the default branch can write a cache entry every branch can read, and this scheduled job is the one place that already pays a full release build. Absent the seed, every pull request keeps compiling the base arm's whole release closure cold, and writing unreadable per-pull-request entries stays the only alternative.
Admits-answer-precondition: The change adds two GitHub Actions steps to the scheduled job that already runs on the default branch — staging the release closure it has just built, and saving it under the key the pull-request side restores. No batten surface expresses a workflow step, so writing the protected path directly is the only route left, and the write lands in PR #833's diff where a reviewer sees it.
Admits-answer-rejected-route: `config read first` does not apply: the committed authority declares rules, verdicts and patterns and carries no representation of a workflow job or step, so there is nothing to read or set there. `patch run first` does not apply either: the workflow is hand-authored YAML with no generator behind it, so a patch route would write the same bytes to the same protected path.

Admits: 7d1ad5bf1d04296f330e798f42f2a74aba2e5d7f78ab4b9a8d2e2474c9dea5e3
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: policy/ci-parity.rego
Admits-head: fb2987d
Admits-epoch: 93d8b1112718944856b6e81b9cfac7c903dbb2c28470be1fc6e59786fb42f404
Admits-author: alec@wenzowski.com
Admits-prev: d2e45c5830908bdc08f6f59141f3a9b031f68ad27ffdd7d133ebf108d0b30921
Admits-answer-lost: The gate goes dead against its own subject. The sub-actions derive an entry's version from the cached path exactly as the composite does, so an interpolated path is the identical silent total miss — the one measured at 4 runs, 0 hits, ~190 MB discarded each time — and after this branch the repository's only cache steps of that kind are the sub-action spellings the predicate cannot see. A gate that loads clean and matches nothing reads exactly like a clean tree.
Admits-answer-precondition: The predicate matches only the composite spelling of the cache action, and this branch moves the job to the restore-only sub-action plus a save on the scheduled side — so the gate that exists to refuse an interpolated cached path would stop covering the very steps this branch lands. Widening it and adding the discriminating case is a change to the Rego module itself; no configuration surface expresses a predicate body, so writing the protected path directly is the only route, and it lands in PR #833's diff where a reviewer sees it.
Admits-answer-rejected-route: `config read first` does not apply: the committed authority carries the rows that register this module, not the predicate body, so nothing there can widen which action spellings the rule matches. `patch run first` does not apply either: the module is hand-authored and has no generator behind it, so a patch route would write the same bytes to the same protected path.
…viting the second claim

The claim receipt is minted on `claim check`'s pullable path and keyed by
the branch checked out at that moment. That one fact makes two orderings
fail in opposite directions, and both are reachable by following the
instructions correctly the first time.

Claim, then branch: the receipt is minted against the branch you were
standing on, and the first edit on the real branch is refused with
`receipt read missing claim branch claim-needs-receipt`.

Claim twice: having branched first, the pullable message read as
go-do-this-and-come-back, so the next move was to move the row and run
`claim-check` again. The second run arrives after the row has left Todo,
reads it as held, and refuses `not-todo` — where the holder is the caller
ninety seconds earlier. The only route past is `--takeover` against
oneself, which writes a takeover record for a row nobody else touched,
and CLOUD-1139 needs that signal to stay rare. Measured twice, in two
sessions, both by an agent following the documented order.

NO ENGINE CHANGE, and the first revision of this row assumed one was
needed. Creating a branch writes only under `.git`, which
`claim-needs-receipt` never judges, so there is no ordering deadlock —
the deadlock was imagined. `claim.rs`'s `not-todo` decision is correct as
it stands for a row genuinely held elsewhere and is untouched.

Recognising a self-claim in the engine stays CLOUD-1139's: it needs to
tell this session from a sibling, and the row's assignee cannot, since
every fleet session carries the same configured accountable identity. An
assignee-keyed re-mint would let a sibling silently re-mint over a working
holder — that row's measured harm, reintroduced through this row's fix.

The gate decides what a gate can: whether the always-loaded file still
states the order and whether the triggered file still carries both failure
directions. Whether a given session actually claimed before branching is
not a property of the tree, and a rule resolving to it would be the model
verdict non-negotiable rule 3 forbids.

Two files because a budget forced it. `policy-budget` caps the index at
3500 tokens and 199 lines, and an earlier draft of this change blew both
at 3516/202; the index carries only the ORDER and the reason lives in the
rules file that loads at the trigger. Both arms are therefore required.

The compiled tier is what proves the two markdown files reach the module
at all: a dead gate and a tree that still states the order are
byte-identical on the decision surface, so each drift case can only go red
if the lines actually arrived.

Refs: CLOUD-1343, CLOUD-1139, CLOUD-786, CLOUD-733

Admits: 3fe3ee98cc0ec6d5f1a9ea595f823c2daa76332568b0ca2332363539c5979d99
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: batten.toml
Admits-head: f9f48e6
Admits-epoch: 84237be51e176ca21fe43fbe5f50492b50c188b8881b9745dfb2b4004a75a220
Admits-author: alec@wenzowski.com
Admits-prev: 3d6283f7ae44fe7b8349bbb4ef3aa0b22e03bc84ad4059098ff15576b3729337
Admits-answer-lost: The gate is dead rather than absent, which is strictly worse. Without the `line_sources` list the engine builds no lines for those two paths, every clause reads undefined, Rego takes undefined as does-not-hold, and the module loads clean while deciding nothing — indistinguishable on the decision surface from a tree that still states the order. That is the exact failure class `.claude/rules/policy-modules.md` records, and without the registering row the module is never evaluated at all.
Admits-answer-precondition: Registering the claim-order module IS an edit to the committed authority: a tree-scoped `[[rule]]` row naming it, the `line_sources` list that makes the two instruction files reach the predicate at all, and the `[[verdict]]` plus `[[verdict.route]]` rows declaring the `claim declare dropped` token. A module raising an undeclared token fails to load, and a declared verdict row nothing raises fails the load too, so the rows and the module are one indivisible change. There is no other surface that expresses them, and the write lands in PR #833's diff where a reviewer sees it.
Admits-answer-rejected-route: `config read first` is what this IS — that route names reading configuration before writing it, and the rows being added do not exist to be read, because the module is new. `patch run first` does not apply: the file is hand-authored and has no generator behind it, so a patch route would write the same bytes to the same protected path.

Admits: f4a603586d8e7912a35d4358fe20ed1955c112b132ecbaf9ad46eebc2bfa1742
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: policy/claim-order-is-stated.rego
Admits-head: f9f48e6
Admits-epoch: 84237be51e176ca21fe43fbe5f50492b50c188b8881b9745dfb2b4004a75a220
Admits-author: alec@wenzowski.com
Admits-prev: -
Admits-answer-lost: The order evaporates. The failure mode here is DRIFT — the remedy is prose, and prose is what a later edit silently rewords or drops, after which the next session walks into the same refusal that cost this one a full stop mid-work and a takeover record against its own claim. A prose assertion with no gate is exactly the half-change rule 2 refuses, and the Ready gate already refused an earlier revision of this row for carrying an empty tests array.
Admits-answer-precondition: CLOUD-1343's remedy is words in two instruction files, and non-negotiable rule 2 calls a rule without a runnable gate half a change. The gate over words is a Rego module: which literal phrases each file must still carry, which file a finding points at, and the could-not-look arm are the module's own body, and no configuration surface expresses a predicate. Writing the protected path is the only route, and the file lands in PR #833's diff where a reviewer sees it.
Admits-answer-rejected-route: `config read first` does not apply: the committed authority carries the rows that register a module and the literal phrases the predicate matches are not expressible there. `patch run first` does not apply either: the module is hand-authored with no generator behind it, so a patch route would write the same bytes to the same protected path.
…iling-run adjacency

The predicate this branch first landed was a second authority. CLOUD-1347
and CLOUD-1350 already design one module for the whole loop vocabulary,
and `a-repeated-call-is-not-progress` was CLOUD-1350's window-recurrence
detector under another name, computing another fact. Two predicates over
one question can disagree, and the disagreement is discovered by a
session being refused, so the module is renamed onto the declared
vocabulary rather than left beside it.

WHAT REPLACES IT IS NOT A RENAME. `repeated-calls` was the MAXIMUM
recurrences of any identity over the whole stream. `repeat-depth` is the
TRAILING run of identical calls, and `distinct-calls` is the progress
term beside it. Adjacency is what does the false-positive work: any
intervening distinct call clears the run, so the edit-then-retest loop is
false by construction rather than by carve-out, and the threshold is the
3 that opencode, hermes-agent and OpenHands all converge on rather than
a constant derived from one pathological session.

AND ADJACENCY IS WHY THIS ONE CANNOT LOCK A SESSION OUT. A whole-stream
maximum is monotonic: once it crossed its threshold it stayed crossed for
the rest of the session, which is how the earlier predicate refused every
subsequent tool call including its own author's push, with no route to an
admission — `batten override request` and `spend` are themselves mediated
calls. A trailing run resets on the next distinct call, so the escape is
automatic and the override CLOUD-1352 names is actually reachable. Every
member of this family is non-monotonic for that reason.

The fingerprint stays inside `transcript.rs`, hashed and dropped in the
same expression as `Event::HookOutput`'s digest, so only a run length is
projected and `Extraction` stays integers-only. A replayed `tool_use` id
is deduped: counting replays measures what the host chose to re-emit
rather than what the session did.

Severity stays `warn`. CLOUD-1352 owns the promotion and makes a measured
firing rate over this repository's own history a hard precondition.

Refs: CLOUD-1347, CLOUD-1341, CLOUD-1350, CLOUD-1352, CLOUD-1337

Admits: 04117a36b963f590a81de944785fade06bd0eb7b27cd2b35f33d59fa3bae4b07
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: batten.toml
Admits-head: 6420fb9
Admits-epoch: 6f8c3ea4c54e6263a8ca91ed3f034ea58de55f4324b3d93f7737b467fcc2467d
Admits-author: alec@wenzowski.com
Admits-prev: 3fe3ee98cc0ec6d5f1a9ea595f823c2daa76332568b0ca2332363539c5979d99
Admits-answer-lost: The module cannot load, so CLOUD-1341 ships nothing. An unregistered `policy/*.rego` file is inert — the engine reads its rule set from the committed authority, and a module with no registering row is never evaluated. The extractor row is what makes the count reach the predicate at all: without it the fact is undefined, Rego reads undefined as does-not-hold, and the gate loads clean while deciding nothing, which is the dead-gate class `.claude/rules/policy-modules.md` exists to warn about.
Admits-answer-precondition: Registering a policy module IS an edit to the committed authority: a `[[rule]]` row naming the module, a `[[rule.extract]]` row binding the new `repeated-calls` extractor to the key the module reads, and the `[[verdict]]` plus `[[verdict.route]]` rows declaring the `turn ask twice` token. A module raising an undeclared token fails to load, and a declared verdict row nothing raises fails the load too, so the rows and the module are one indivisible change. There is no other surface that expresses them, and the write lands in PR #833's diff where a reviewer sees it.
Admits-answer-rejected-route: `config read first` is what this IS — that route names reading configuration before writing it, and the rows being added do not exist to be read, because the module is new. `patch run first` does not apply: the file is hand-authored and has no generator behind it, so a patch route would write the same bytes to the same protected path; the generated `schema/*.json` are derived FROM the crate rather than producing it.

Admits: 368813d8fbd4cfef084b59aaa4b102fdf06d6bdb3d21ceaf0df87cc61629d2a8
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: policy/repetition-without-progress.rego
Admits-head: 6420fb9
Admits-epoch: 6f8c3ea4c54e6263a8ca91ed3f034ea58de55f4324b3d93f7737b467fcc2467d
Admits-author: alec@wenzowski.com
Admits-prev: -
Admits-answer-lost: The branch ships a second authority over a question the tracker already owns. A family of rows designs one module holding the whole loop vocabulary; landing a differently-named module computing a differently-named fact means whoever builds that family writes the second one, and this repository refuses exactly that — two predicates over one question can disagree, and the disagreement is discovered by a session being refused. Without the restructure the duplicate is what lands.
Admits-answer-precondition: The predicate is a Rego module: the threshold constant, the three-valued posture guard, the null guard and the ten test rules are the module's own body, and no configuration surface expresses a predicate. This replaces a module landed earlier on this same branch under a name that duplicated an already-designed family, so the write is a restructure onto the declared vocabulary rather than a new gate. It lands in PR #833's diff where a reviewer sees it.
Admits-answer-rejected-route: `config read first` does not apply: the committed authority has no surface for a predicate body or a threshold constant, and a threshold spelled as a pattern row is refused outright, so the number has to live in the consumer module. `patch run first` does not apply either: the module is hand-authored with no generator behind it, so a patch route would write the same bytes to the same protected path.

Admits: 8a3ff5e4351e38ed26fdc4fd7eeaa30b80bd7d096ab7ebb2433024111f576156
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: policy/a-repeated-call-is-not-progress.rego
Admits-head: 6420fb9
Admits-epoch: 6f8c3ea4c54e6263a8ca91ed3f034ea58de55f4324b3d93f7737b467fcc2467d
Admits-author: alec@wenzowski.com
Admits-prev: 247eaff38bfbc59e1afae9cae684be633d05c8eca2d901051cad86757508560c
Admits-answer-lost: Two modules over one question ship together. The registering rows can be pointed at the replacement, but leaving the file behind leaves a second authority in the tree that a later reader may register again, and the two can disagree over exactly the cases neither author had in mind. Keeping it would also leave its rule id in the mutation census naming a gate nothing registers.
Admits-answer-precondition: The module being removed was landed earlier on this same branch and duplicates an already-designed family's detector under a different name. Its replacement lands in the same change on the declared path, so the removal is half of one restructure rather than a deletion of coverage — every predicate it carried survives, renamed onto the vocabulary the family declares. There is no surface that retires a module except removing the file and its registering rows.
Admits-answer-rejected-route: `config read first` does not apply: the committed authority registers modules and cannot delete one, and pointing the rows elsewhere is what leaves the orphan this removal exists to prevent. `patch run first` does not apply either: there is no generator behind the module, so no patch route reaches it.
…f answering zero

Four of five `Extraction` members reached no consumer, and repetition —
the one thing a doom loop is made of — had no member at all. This is the
spine the predicates stand on, landed alone because the risk is the
plumbing rather than the detection.

`Extraction::of` widens from `&Counts` to `&Stream`. `Counts` is
unchanged: it is the `-J` capability report's own shape and its comments
refuse members no reader of that document needs, so runs live in
`Stream::repeats()` beside `Stream::counts()` rather than as fields on it.

THE LOAD-BEARING HALF IS THE PER-EXTRACTION CAPABILITY. Zero is a real
answer and means the extractor ran. An extraction whose underlying event
kind never appears is not a session that did none of it — it is a host
that does not record it, and answering zero there is a false green over a
session nobody measured. `of` returns `Option<usize>` and the projection
omits the key, so a module reads undefined and Rego takes that as
does-not-hold. Decided in the engine: a per-module conjunct asking whether
this host records turns is a dead gate on every harness but the one its
author tested.

The bound is stated rather than absorbed: this cannot separate a host that
records no hook runs from a session that triggered none. It resolves that
toward could-not-look, which is the safe direction — a missing answer is
reported, where a false zero is indistinguishable from a clean session.

`agent-turn-run` is the first member and deliberately the simplest: a
trailing run of assistant turns carrying no tool call, over the
already-typed `Event::Turn`. No hashing, so no argument or result is read
even internally. It maps to OpenHands' monologue detector at 3+.

The honest claim is narrower than loop detection, and the module header
says so. Termination is undecidable and every quantity here is a
monotonically growing count, so there is no ranking function to be had.
What the literature buys is the shape of the declaration: a SET of
extractions with a set of thresholds, supplying an effective bound on a
SUSPECTED feedback path. It does not detect non-termination.

Adjacency also makes the member non-monotonic — one action clears the run
— which is the property any later promotion depends on. CLOUD-894 owns the
firing-rate ceiling and CLOUD-1352 the promotion; this ships at `warn`.

The compiled tier is what proves the capability rather than the author's
arithmetic. It caught a wrong fixture doing it: a `user` record parses as
a turn, so the first cross-compatibility case recorded turns after all and
answered a real zero. A host recording hook runs and no turn boundaries is
the shape the claim is about.

Also retires the two now-spent rows in `policy/harness-declared.json`.
CLOUD-1079 has landed, so the user-level hooks those rows excused are no
longer provisioned and the merged surfaces carry zero hook commands —
`harness-wiring` reported both, which is the gate working rather than
misfiring. Reproduced against this same tree with `--config-from a907565`,
so the finding predates this diff and is not caused by it. Leaving them is
not neutral once their owner has landed: they would silently excuse
anything that later matched either pattern by name.

Refs: CLOUD-1344, CLOUD-1079, CLOUD-1049, CLOUD-418, CLOUD-894

Admits: d0ac8a4640f2a747b01bedc8449254f45d80865a7de7f0d3f765c40975627f93
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: batten.toml
Admits-head: aed8b44
Admits-epoch: 3ce215f766f4c55fa0be890ef01ffdb5bfc93d7e597090d5f2098d074c5f389f
Admits-author: alec@wenzowski.com
Admits-prev: 04117a36b963f590a81de944785fade06bd0eb7b27cd2b35f33d59fa3bae4b07
Admits-answer-lost: The module cannot load, so CLOUD-1341 ships nothing. An unregistered `policy/*.rego` file is inert — the engine reads its rule set from the committed authority, and a module with no registering row is never evaluated. The extractor row is what makes the count reach the predicate at all: without it the fact is undefined, Rego reads undefined as does-not-hold, and the gate loads clean while deciding nothing, which is the dead-gate class `.claude/rules/policy-modules.md` exists to warn about.
Admits-answer-precondition: Registering a policy module IS an edit to the committed authority: a `[[rule]]` row naming the module, a `[[rule.extract]]` row binding the new `repeated-calls` extractor to the key the module reads, and the `[[verdict]]` plus `[[verdict.route]]` rows declaring the `turn ask twice` token. A module raising an undeclared token fails to load, and a declared verdict row nothing raises fails the load too, so the rows and the module are one indivisible change. There is no other surface that expresses them, and the write lands in PR #833's diff where a reviewer sees it.
Admits-answer-rejected-route: `config read first` is what this IS — that route names reading configuration before writing it, and the rows being added do not exist to be read, because the module is new. `patch run first` does not apply: the file is hand-authored and has no generator behind it, so a patch route would write the same bytes to the same protected path; the generated `schema/*.json` are derived FROM the crate rather than producing it.

Admits: 67106f144c06ca5c3aeaae3482a922741375d1168c3b3a6016979aadbe61b4a4
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: policy/repetition-without-progress.rego
Admits-head: aed8b44
Admits-epoch: 3ce215f766f4c55fa0be890ef01ffdb5bfc93d7e597090d5f2098d074c5f389f
Admits-author: alec@wenzowski.com
Admits-prev: 368813d8fbd4cfef084b59aaa4b102fdf06d6bdb3d21ceaf0df87cc61629d2a8
Admits-answer-lost: The spine row ships with no predicate reading it, which is the half-change non-negotiable rule 2 refuses: an extraction nothing consumes is exactly the defect CLOUD-1344 was filed about, since four of five existing members already reach no consumer. The branch would add a sixth unread member while claiming to fix that.
Admits-answer-precondition: CLOUD-1344's predicate is a Rego module: the adopted threshold, the null guard and the seven test rules are the module's own body, and no configuration surface expresses a predicate. This replaces the predicate landed earlier on this same branch, which was sequenced ahead of the spine row it is blocked on, so the write is a re-sequencing onto the row that must land first. It lands in PR #833's diff where a reviewer sees it.
Admits-answer-rejected-route: `config read first` does not apply: the committed authority has no surface for a predicate body or a threshold, and a threshold spelled as a pattern row is refused outright. `patch run first` does not apply either: the module is hand-authored with no generator behind it, so a patch route would write the same bytes to the same protected path.

Admits: 06de726f4ec787c03543081e96d8e7ae1d46932d5306fe7b68e4465d56f8db24
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: policy/harness-declared.json
Admits-head: aed8b44
Admits-epoch: 6785dbe2aaf01f887e2ec5902922351fd259cb877a4e4d8c830aeefa1d3079ba
Admits-author: alec@wenzowski.com
Admits-prev: -
Admits-answer-lost: The gate stays red on every commit for a reason nobody caused, which is how a correct refusal gets skipped per commit until it is skipped by habit. Worse, the rows keep excusing two commands by name: were anything to reintroduce a hook matching either pattern, the exemption would silently cover it, so leaving them is not neutral once their owner has landed.
Admits-answer-precondition: The exemption table is a policy data file the engine reads, and the two rows in it are now spent: the issue that owns them has landed, the user-level hooks they excused are no longer provisioned, and the merged surfaces carry zero hook commands. `harness-wiring` reports both, which is the gate working rather than misfiring. Retiring a row means editing that file; there is no other surface for it, and the write lands in PR #833's diff where a reviewer sees it.
Admits-answer-rejected-route: `config read first` does not apply: the committed authority declares which external surfaces are read and holds no exemption rows, which is deliberate — an exemption table is data a gate reads rather than part of the gate. `patch run first` does not apply either: the file is hand-authored with no generator behind it, so a patch route would write the same bytes to the same protected path.
`claim-order-is-stated` fired on any repository carrying an instructions
file, so it refused a fixture whose `AGENTS.md` is the single word
`instructions` and which has no `.claude/rules/` surface at all.
`cli::the_committed_repo_config_gates_a_repository` runs the whole
committed ruleset over exactly that tree and asserts its entire output,
so it went red — the foreign-tree failure every tree-scoped row here has
to answer, caught by the one case that was watching for it.

The guard is keyed on the TRIGGERED file rather than the index, and that
is the whole of the reasoning. This row is about a SPLIT: the order in
the always-loaded file, the reason in the file that loads at the trigger.
A tree with no `.claude/rules/toolchain.md` has not made that split and
is answering for nothing here. Guarding on the index instead would be
circular, since an absent index is precisely what one arm exists to
refuse.

Same shape as `hk-fix-selection`'s `governed`, which asks whether the
repository has the config it judges before judging it.

Both tiers gain the case, so the fixture shape is pinned rather than
merely un-refused: `test_an_index_without_the_rules_file_is_not_judged`
and `an_index_without_the_rules_file_is_not_judged`.

Refs: CLOUD-1343

Admits: 03d9f351561fe5a7cbd22d78ad942b3f541a3ac5a92b241d8c327d36becf58b9
Admits-rule: protected-mutation
Admits-verdict: path write refused
Admits-subject: policy/claim-order-is-stated.rego
Admits-head: 889f97b
Admits-epoch: fb78c0a9b65b102b88a160ecb71c69eb089570a736ea14b96c2ea053177664fd
Admits-author: alec@wenzowski.com
Admits-prev: f4a603586d8e7912a35d4358fe20ed1955c112b132ecbaf9ad46eebc2bfa1742
Admits-answer-lost: The compiled suite stays red on a case that asserts the committed configuration's whole output over a fixture, so the branch cannot land at all. Worse as a shipped property: a consumer adopting this row would have it fire on any repository with an instructions file and no triggered rules file, which is a refusal about a document split that repository never made — the foreign-tree failure every tree-scoped row here has to answer.
Admits-answer-precondition: The module judges any tree carrying an index file, so it fires on a fixture repository whose stub instructions file has never stated the claim order — a foreign tree answering for a split this repository's own instruction surface has. The fix is a not-applicable guard in the predicate body, which no configuration surface expresses; the committed authority registers the module and cannot narrow what it decides. It lands in PR #833's diff where a reviewer sees it.
Admits-answer-rejected-route: `config read first` does not apply: the committed authority declares the module's sources and severity and has no surface for a not-applicable conjunct, which is part of the predicate. `patch run first` does not apply either: the module is hand-authored with no generator behind it, so a patch route would write the same bytes to the same protected path.
CLOUD-1342 replaced the composite `actions/cache` with
`actions/cache/restore` in `ci.yml` and added `actions/cache/save` to
`perf.yml`. Both are SHA-pinned, so both are `pkg:github` components in
the inventory, and neither was in the table `sbom-inventory` reads —
`sbom-action-unenriched` reported 2.

Same repository, same commit as the composite entry already in the
table, so the licence and the copyright line are the same facts rather
than a guess.

The counts that rule decides over come from the RECORDED scan rather
than from this file, so the table alone does not clear it: the record is
regenerated by `mise run record-sbom`, and it lives out of tree, so a
fresh container has to re-run that before `batten-check` can pass.

Refs: CLOUD-1342, CLOUD-667
Two workspace lints, both against the spine landed two commits ago, and
both fixed rather than suppressed.

`struct_excessive_bools` on `Records`. A struct of four flags asks a
reader to remember which one is which, and clippy's own remedy — a
closed vocabulary — is the better shape here anyway: `Kind` is an enum
and `records()` returns a `BTreeSet<Kind>`, so a caller asks `contains`
and a kind added later is a variant every exhaustive match decides or
fails to compile over. That is the same argument `Extraction`'s own
closed set already makes one module across.

`match_same_arms` in `Stream::repeats()`. Here the arms really are one
decision: a tool call and a user turn both end the model's trailing
monologue — it acted, or the operator spoke — so they merge under one
comment rather than carrying an `expect` over an arm that decides the
same thing twice. `Stream::counts()`'s `expect` stays, because its two
silent arms are genuinely different reasons.

No behaviour changes. The compiled tier is unchanged and still green,
including the pair that carries CLOUD-1344: a host recording no turns
answers could-not-look, and a recorded session with no run answers a
real zero.

Refs: CLOUD-1344
`a_clean_final_message_says_nothing` asserts silence and got
`unlanded: 1 commit(s) not on the landing target` — a finding from the
checkout the suite was running in, not from its fixture.

The Stop tier reads the out-of-tree findings store, and `hook()` set no
state home, so it inherited the ambient one. The consequence is worse
than one flaky case: the test passes on a tree with nothing unlanded and
fails on any branch carrying unpushed work, which is every branch this
suite is ever run from mid-development. It went unnoticed because the
green path is the one a clean checkout takes.

The suite already knew. Its unlanded fixture's own runner contains the
state home and says why in as many words — "an ambient one would let a
real session's findings decide a fixture's verdict" — and the posture
half simply never did it. This applies the same containment there, so
both halves of the file now hold the same invariant.

Derived per fixture from the directory name rather than shared, because
nextest runs each case in its own process and a shared scratch name is a
wipe under another process's read.

All 16 cases pass, including `unlanded_work_at_a_declared_stopping_point_is_pointed_at`
— the containment isolates the store rather than emptying it, so the
cases that must READ a finding still do.

Refs: CLOUD-97
@wenzowski
wenzowski marked this pull request as ready for review September 3, 2026 04:45
@wenzowski
wenzowski force-pushed the claude/cloud-1331-perf-base-arm-6pgmy2 branch from 0be1024 to 55ca205 Compare September 3, 2026 04:45
@sonarqubecloud

sonarqubecloud Bot commented Sep 3, 2026

Copy link
Copy Markdown

❌ The last analysis has failed.

See analysis details on SonarQube Cloud

@wenzowski

Copy link
Copy Markdown
Contributor Author

/fast-forward

@wenzowski
wenzowski merged commit 55ca205 into main Sep 3, 2026
11 of 12 checks passed
@wenzowski
wenzowski deleted the claude/cloud-1331-perf-base-arm-6pgmy2 branch September 3, 2026 05:04

Copy link
Copy Markdown
Contributor Author

Punt audit — session close

This PR is merged (main at 55ca205a, all ten required checks green). Written here because the session is being archived; everything below now has a durable home, and this comment is the index rather than the record.

Filed on the rows that own them

finding home
Recovery gap: override request/spend are themselves mediated calls, so an admission can never clear a mediated_call deny. Survivability today rests on predicate non-monotonicity, not on any recovery path. CLOUD-1051
The composed disarm trap — disarm guard plus a mediated_call deny leave no exit, each defensible alone. CLOUD-1051
CLOUD-1350's §2 is not implementable as written: the progress conjunct names no ratio or threshold, and the declared window has no path from config to the reduction (Extraction::of(self, &Stream) -> Option<usize> carries no per-row parameter). CLOUD-1350
landed-check.sh cannot be extended — the write is allowed (tree-scoped rule) but batten check --rule shell-retirement answers exit 2 shell-rule-retired. CLOUD-1127 §1 re-scoped to a retirement; §7 still names tests/landed-check.bats, which the retirement deletes. CLOUD-1127
closing-key-check's refusal text — the designed fix ("decline just that one") needs scope, and is blocked behind the same retirement. CLOUD-1127
The &Stream widening took the extraction reduction from one walk to 2 × declared extractions. My own landed work, not fixed here. CLOUD-1345
A session-long live instance of the interpreter hole. CLOUD-1141
CLOUD-1341's claims object is stale (repeats-may-go-unpriced is declared nowhere), and ready lint passes over it — the clause checks the object's shape, not whether <suite>:<mutation> resolves against the tree. CLOUD-1341
The two declared mutations in policy/repetition-without-progress.rego have never been shown to redden. mise run mutant was started once, killed, never re-run — so CLOUD-418's "shown able to fail" is unverified for this landing. CLOUD-1344

Measured, gate-visible, no note needed

batten ready lint --issue over the family: CLOUD-1344 carries a correct claims object; CLOUD-1345, CLOUD-1347, CLOUD-1350 and CLOUD-1352 each report claims-object-absent; CLOUD-1345 additionally reports bump-disagrees-with-type (perf implies no bump); CLOUD-1346 has no stored payload in this repository's capture store. Any agent pulling those rows gets the refusal on demand, so recording it per-row would be a copy with no reader — this paragraph exists only so the sweep is not paid for twice.

Two bot signals on this PR, unactioned at merge

  • CodeRabbit hit its included-review limit. Its last actual review covered batten.toml only, up to a9075652. The final 18-file diff that merged as 55ca205a was never line-reviewed by it; the summary it posted is a walkthrough, not a review, and reports no findings.
  • SonarQube Cloud: "the last analysis has failed", 2026-09-03T04:45Z, ~18 minutes before the merge. It is not in the required check set, so the landing loop merged over it correctly — but nothing has looked at why the analysis failed, and no row owns that.

Corrections to my own conduct in this session, for the record

  • I moved CLOUD-1341's state three times on a misreading of what DO-NOT-CLOSE does. The token is real batten vocabulary (closing-key-check.sh:64) but has no effect on the tracker; I had inferred an automation behaviour that does not exist. The row is in Backlog, which is where the tree puts it.
  • I built the extraction family out of order once (aed8b44), violating CLOUD-1344's own acceptance, and reverted the engine files to a907565 to build the spine first.
  • #MUTANT rows initially named .rego test_ rules rather than cases in the declared suite; caught by reading feat(land): four of the lap's six steps as verbs, and three decisions as predicates #829.

Dies with the container, deliberately not preserved

/root/.claude/plans/breezy-whistling-puffin.md — the landing plan. Every live item in it is either landed or in the table above.


Generated by Claude Code

Copy link
Copy Markdown
Contributor Author

Index correction — the audit table above was itself incomplete. Two more findings are now on CLOUD-1051, which is where the exit question lives:

  • The stop hook demands a push the policy engine refuses. The Stop posture asks for committed-and-pushed before the turn ends; a mediated_call deny over that push is decided on the same boundary with no arbitration between them. Same mechanism as the composed trap, reached by a second route.
  • BATTEN_HOOK_BYPASS is live, not retired. Measured today: 21 files, hook.rs / facts.rs / lib.rs / verdict.rs / doctor.rs / batten.toml / schema/policy-call.schema.json, pinned by five test modules. A handoff asserted it was retired; it is not. The prose must not be purged — that would make the docs wrong about a working mechanism, and it is currently the only unmediated route out of a mediated_call deny.

Neither became a new row: filing-needs-a-search wants a duplicate search first, and that is out of budget at session close.


Generated by Claude Code

Copy link
Copy Markdown
Contributor Author

Second index correction — the audit above recorded the shell findings without citing the campaign that owns them. All bash here is scheduled for retirement under CLOUD-843, so landed-check.sh and closing-key-check.sh are wave items, not standalone obstacles. Links now on the rows:

the item belongs against
Both programs are still in the tree, so both are in the 131 unretired. CLOUD-1151 (waves 2 and 3 have no owner) — now carries the measured blocker for each
landed-check's successor must be a policy module over input.tree["commit-meta"], not a CLI verb. CLOUD-1182 (a retirement does not declare its successor kind)
The refusal arriving at verify rather than at the write — what I measured this session. CLOUD-1131, which states exactly that and is marked Done; recorded as a reproduction, not a verdict on it
tests/landed-check.bats is what makes §7 stale — a removed case is not admitted, and a frozen suite pins a remedy string. CLOUD-1294, CLOUD-1299
closing-key-check's "decline just that one" needs a granularity the ratchet does not have. CLOUD-1108 (the ratchet is FILE-granular) — not a blocker of its own
~15s per deleted governed path against a Surface::Check with no ceiling. CLOUD-1321
The two-shapes rule itself. CLOUD-1132

And the mutation findings have a campaign frame too: CLOUD-1355 (a declared mutation naming a case that does not exist is only reachable from the nightly sweep) is CLOUD-843's child and the precedent for both — linked from CLOUD-1341 and CLOUD-1344, with the distinction stated so they are not merged: 1355 is the module's declaration, 1341 is the issue's obligation, and nothing reports the second at all.


Generated by Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant