Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
20 commits
Select commit Hold shift + click to select a range
6aa2ecc
fix(verdict): measure the single-token dictionary the verdict grammar…
wenzowski Sep 1, 2026
a3ba695
fix(verdict): convert every refusal class to the declared three-word …
wenzowski Sep 1, 2026
6470ef2
fix(policy): make the stays-bash route clear the verdict that offers it
wenzowski Sep 1, 2026
49d6fcd
fix(hook): every mediated deny carries a declared class
wenzowski Sep 1, 2026
3111d37
fix(hook): the hot path emits a class and its pointers, and stops
wenzowski Sep 1, 2026
f7539e0
fix(hook): verb and operand attribution stops at a newline
wenzowski Sep 1, 2026
80a1f81
fix(hook): resolve a relative operand against the caller's cwd, and s…
wenzowski Sep 1, 2026
7051181
fix(hook): a bare directory destination is inside the protected set
wenzowski Sep 1, 2026
6103071
feat(hook): a generic read of a memory names the tool that answers it
wenzowski Sep 1, 2026
0d786e1
fix(rules): derive the mediated column refusal from the fact model
wenzowski Sep 1, 2026
942ca55
fix(receipt): a branch behind its own receipt is not a restart
wenzowski Sep 1, 2026
8ce896c
feat(hook): one budget for the advisory channel, not per producer
wenzowski Sep 1, 2026
046659f
fix(hook): bound what a session's hooks cost it, and make a repeat de…
wenzowski Sep 1, 2026
52be523
fix(verdict): convert the classes main added after the grammar landed
wenzowski Sep 1, 2026
30a07ed
refactor(tests): retire the fact-record-keying suite onto the compile…
wenzowski Sep 1, 2026
872f5bf
fix(verdict): convert the lock-complete classes main landed mid-rebase
wenzowski Sep 1, 2026
0e3d9dc
fix(verdict): convert fixture registry ids to the three-word grammar
wenzowski Sep 1, 2026
4571135
fix(config): restore the pattern id the rename sweep hit, and declare…
wenzowski Sep 1, 2026
9737a46
feat(policy): a remedy resolves to a declared command or rule
wenzowski Sep 1, 2026
707ab41
fix(trust): a verdict rename is not a hatch newly added
wenzowski Sep 1, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
24 changes: 22 additions & 2 deletions .claude/rules/policy-modules.md
Original file line number Diff line number Diff line change
Expand Up @@ -47,7 +47,7 @@ above rather than with the authoring judgement below.
```rego
violation contains {
"rule": "shell-rule-retired",
"verdict": "V-SHELL-RULE-EDITED",
"verdict": "shell edit refused",
"subjects": [{"path": path}],
} if { ... }
```
Expand Down Expand Up @@ -241,7 +241,27 @@ A **newline is whitespace, not a separator** — bash disagrees, and the bound i
deliberate rather than an oversight: promoting it would change every landed
`pipeline` verdict. So the shell following a heredoc's terminator joins the
segment its opener was written in, and a two-command call written across lines is
judged as one. It under-denies, which is the sanctioned direction.
judged as one segment.

**AND "it under-denies, which is the sanctioned direction" IS MEASURED
BACKWARDS** (CLOUD-1287). That sentence stood here and was false of the arm that
matters most: one segment means `effective_program` resolves the FIRST line's
program for every operand on every line, so a declared `protected_readers` entry
was unreachable from any script. Measured over the shipped binary, one protected
path, the same read twice: `stat -c %s batten.toml` allowed, and the identical
`stat` written on line two after `cd /tmp` REFUSED, naming `cd`. That is an
OVER-deny, on a read, which is the direction that gets a guard switched off
rather than the sanctioned one.

The bound above still holds for segment identity — `terminator` is unmoved and no
landed `pipeline` verdict changed. What changed is narrower and lives in the
engine rather than in a module: `hook::line_bounded_words` splits a segment's own
`raw` at newlines and re-enters `segments` per line, and only the mutation walk
and the unknown-program walk read it. Both ask "which program was handed this
operand", a question a line answers and a segment does not. So a module reading
`input.call.segments` sees exactly what it saw before, and must not grow its own
line splitting to compensate — that would be the second authority two sections
up already refuses.

There is **one parser**, and a module must not grow a second: no `split` of the
command line, in Rego or in Rust. **The reason is not effort, and giving it as
Expand Down
6 changes: 3 additions & 3 deletions .claude/rules/toolchain.md
Original file line number Diff line number Diff line change
Expand Up @@ -51,7 +51,7 @@ So a change touching one has exactly two shapes:
`conserves` arm per deleted path, and drop the gate from `$MUTANT_GATES`.
2. **Leave the file alone.**

`V-SHELL-RULE-EDITED` declares one route, `R-PORT-AND-RETIRE`, with no override
`shell edit refused` declares one route, `rule read first`, with no override
and no `bypass_env`. That is not an oversight to be worked around; it is the
whole design.

Expand Down Expand Up @@ -116,7 +116,7 @@ on the arm, beside the successor it qualifies:
```

A `policy/*.rego` or preset successor needs no field — its path already decides
it — and `V-SUCCESSOR-KIND-UNDECLARED` refuses only the engine-source arm that
it — and `shell port unnamed` refuses only the engine-source arm that
omits one. **It does not refuse a verb**, and that is the point rather than a
softening: a gate needing stdin, spawning with its own arguments, or performing a
write cannot be a tree-scoped module, so the choice has to stay available and
Expand Down Expand Up @@ -459,7 +459,7 @@ call` with no `CLOUD-*` key **in that same paragraph** stops the lap. Two open
**Both environment variables are gone rather than ported** (CLOUD-1051):
`BATTEN_FILED_HERE_BYPASS` and `BATTEN_FILED_HERE_OVERLAP` were knowable
strings anyone could spend without articulating anything, and the override is
`V-FILED-OVER-OWN-DIFF`'s declared route with its precondition, issued and
`issue file same`'s declared route with its precondition, issued and
spent through `batten override request`/`spend`. `land` still calls
`mise run filed-here-check` by name — an inline `batten check` on that row now
— so that call site is byte-identical and `land.sh` never
Expand Down
2 changes: 1 addition & 1 deletion .github/workflows/auto-release-land.yml
Original file line number Diff line number Diff line change
Expand Up @@ -313,7 +313,7 @@ jobs:
# AND IT IS SET HERE RATHER THAN IN THE TASK, which is forced rather than
# chosen. `mise-tasks/release-due.sh` is `governed_at_head` under
# `policy/shell-retirement.rego` — it carries a shebang and a `#MISE
# description=` line — so editing its default raises V-SHELL-RULE-EDITED,
# description=` line — so editing its default raises shell edit refused,
# which declares no override route and no `bypass_env`; `tests/release-due.bats`
# is governed the same way. `mise.toml [env]` is ungoverned but WRONG: it
# leaks into `mise run test:bats` and breaks the three cases pinned to the
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,7 @@ typed := `package batten.example

violation contains {
"rule": "a-gate",
"verdict": "V-A-CLASS",
"verdict": "a class probe",
"subjects": [{"path": "a.rs"}],
} if {
input.call.operation == "write"
Expand Down Expand Up @@ -52,7 +52,7 @@ msg_as_a_value := `package batten.example

violation contains {
"rule": "a-gate",
"verdict": "V-A-CLASS",
"verdict": "a class probe",
"subjects": [{"artifact": "msg"}],
} if {
input.call.operation == "write"
Expand Down
57 changes: 57 additions & 0 deletions .serena/memories/core.md
Original file line number Diff line number Diff line change
Expand Up @@ -134,6 +134,31 @@ err)` takes **both** channels and the resolved `Mode`, so a verb can write a
no locking at all — and the questions come off a class's declared
`override.precondition` in `verdict.rs`. The gate never grades an answer;
presence and non-emptiness are the whole predicate (rule 3).
- `advisory.rs` — the advisory CHANNEL and what it may cost (CLOUD-896). A LEAF
beside `refusal.rs`, and the pairing is the whole placement: `refusal` bounds
ONE emitted deny line, this bounds ONE emission of the whole channel, so the two
answer the same question over the two documents a boundary can produce.
CLOUD-461 coalesced the FRAMING — one `additionalContext` object per call — and
bounded nothing about volume: three producers (`drain::render`,
`contract::render`, the dispatched handler) shared no rate budget, and
`[drain] token_budget` bounds only one of them, so the channel's real ceiling
was whatever the set summed to. `[advisory] max_tokens` supersedes it — one
fact, one authority. `admit` sorts by `AdvisoryTier` (CLOUD-80's severity as
required response latency, `Reverse` because the derive is weakest-first) and
fills until the ceiling is spent; the tier is carried from the PUSH SITE in
`lib.rs` rather than inferred here, because "how soon must this be answered" is
a property of what is said and the boundary has only a string. **The remainder
is dropped AND COUNTED** — a truncated report that reads as complete is the
false green in advisory form — and the count line is a count and a ceiling,
never the dropped text. **The first entry is always admitted**, even alone over
budget, so the count line can never be the only thing said. An UNDECLARED
ceiling emits exactly what it emitted before, in the boundary's own order: that
is the anti-vacuity half, and it is what keeps this consumer's number out of
every other consumer's engine (rule 1). `validate` refuses `max_tokens = 0` at
load — a channel switched off wearing a budget's clothes. `trust.rs` carries
`AdvisoryCeilingRaised`: smaller is stricter, absent is unenforced rather than
zero. It reaches `budget` for the estimator `refusal` already reaches, because a
second one would be a second authority over what a token costs.
- `action.rs` — the `[[hook.action]]` plugin surface (CLOUD-91), house-style §9's
"repo-specific cleanup or keepalive is reconstructed here, not hardcoded". A row
names an event and argv already on the operator's PATH. **`fire` returns
Expand Down Expand Up @@ -1064,6 +1089,38 @@ transcript CONTENT needs 1029 first, and nothing landed authorises one.
housing it there would close a module cycle. Bound (CLOUD-211): a mediated deny
comes only from a computable predicate, never a judge verdict, so the shape
models no advisory output — no confidence, no severity, no "maybe".
- `hookcost.rs` — what this repository's own hooks cost the session that runs
them (CLOUD-417). Rule 4 is stated per-CHECK and enforced per-CHECK, so nobody
had measured the hooks IN AGGREGATE, where one compliant line is emitted
hundreds of times and every copy stays in context forever. Measured on one
captured transcript (758 turns, 5.83 MB): `hook_success` 1181 KB +
`hook_additional_context` 42 KB — **hook output alone is 20% of the
transcript**, the largest contributor saying one identical true thing every
turn. Same shape CLOUD-896 found one layer down, and the same answer: put the
ceiling on the aggregate, because the aggregate is what is spent. Two
predicates on `[hook_output]`, and the second is most of the win —
`max_tokens` bounds the session, `max_repeats` makes **"silence on success is
the default"** and **"a repeat is a pointer to the first, not a copy"**
DECIDABLE rather than prose, because a hook saying the same thing every turn is
a digest repeated. Floor on `max_repeats` is 1, never 0: saying it once is the
report, and refusing the first emission would silence the finding rather than
its restatement — which the row puts explicitly out of scope. Pointer-only
structurally: `transcript.rs` hashes the emitted text
(`identity::context_fingerprint`) and DROPS it at the parse, so a measurement
of an over-wide channel cannot itself carry what the channel said; the report
keeps eight hex characters, a count and the first copy's line. Grouping key is
(producer, digest) — two hooks emitting one string are two producers, not a
repeat. `Event::HookOutput` is APPENDED to the transcript vocabulary and gated
on the host's `hook_*` tag PREFIX rather than an enumerated set, because an
unrecognized tag counted as zero is the silent under-report this row ends.
Empty output is not an emission, which is what keeps three silent records from
hashing alike and manufacturing a violation out of the behaviour being asked
for. `Reading::line` is ONE line and `batten policy hooks` prints nothing
around it — the self-applying property, asserted in both tiers. Not in
`verify` and not in the hk gate: a transcript is a property of the WORLD, not
of the commit (`lock-complete` vs `lock-currency`), so `mise run hook-cost` is
a hand run — which is also the row's acceptance clause, the 20% figure shipped
as a re-runnable command rather than a number in an issue body.
- `markers.rs` — counted suppression markers (CLOUD-36): how many times policy
was waved through, and where. Tokens are config, never crate constants (rule
1); hits are pointer-only (`path:line` + marker id, rule 4) and `counts`
Expand Down
41 changes: 40 additions & 1 deletion Cargo.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

16 changes: 16 additions & 0 deletions Cargo.toml
Original file line number Diff line number Diff line change
Expand Up @@ -202,6 +202,22 @@ insta = { version = "1", default-features = false }
# Inline fixtures for the snapshot cases, so a multi-line expectation reads as
# the block it is rather than as escaped `\n`s.
indoc = "2"
# The BPE tokenizer the verdict vocabulary is measured against (CLOUD-1284's arm
# 4). DEV-ONLY, and that placement is the whole of why it is affordable.
#
# `budget.rs` estimates bytes/4 on purpose -- "an exact count needs a tokenizer,
# a vocabulary and a network fetch, and a budget gate that fails because a
# download failed is worse than one 10% out" -- and `batten hook` runs on every
# mediated tool call under CLOUD-689's budget. A tokenizer in the SHIPPED binary
# would put an embedded merge table on the config-load path and contradict both.
# A tokenizer in the TEST binary contradicts neither: the vocabulary is a fixed
# table in `batten.toml`, so its token counts are a property of the commit, and
# the one place that property has to hold is a gate that reads the commit.
#
# So arm 4 is enforced where it is decidable and costs nothing at runtime. The
# pin it measures under is declared as data beside the vocabulary, never baked in
# here -- `bench/tokens/method.toml`'s discipline for its byte divisor.
tiktoken-rs = "0.12"
# `min_batten_version` comparison. Cargo's own flavour of SemVer, so the config
# key means what a Rust consumer already expects it to.
flate2 = { version = "1", default-features = false, features = ["rust_backend"] }
Expand Down
22 changes: 19 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -112,9 +112,25 @@ Batten's output contract answers all three at once. A finding is a **pointer, no
a payload** — a count and a `path:line`, never the matched content — so a wrapped
tool's two thousand lines become one. Output is **byte-stable**, so an unchanged
repository renders identical bytes and the agent's prefix cache stays warm instead
of being invalidated by a reordered map or a timestamp. And a refusal **points at
the fix**: a deny names the rule, the reason, and the command to run instead,
which is one hop to right rather than a round of guessing.
of being invalidated by a reordered map or a timestamp.

And a refusal **points at the fix — the class name IS the pointer.** A mediated
deny emits one line: a declared three-word class and the pointers it applies to,
as in `shell edit refused mise-tasks/land.sh:845`. The reason, the routes out
(the escape hatch and the override alike) and the class's full definition are
one hop away, at `batten policy explain "shell edit refused"`.

That is a decision rather than an omission, and it is the same argument as the
three pains above turned on the tool's own output. The reason and the remedy do
not change between firings, and a mediated refusal fires hundreds of times in a
long session, so inlining them means paying per firing for text that was
declared once. The name carries the class because the names are a declared
vocabulary rather than free text — three positional words, each glossed — which
is what makes one hop cheap and the elision honest rather than merely shorter.
The pointer, which DOES change per firing, stays inline: this shortens the
prose, never the operand a reader acts on. The ceiling on that line is declared
in `batten.toml` and gated, so "one line" is a property of the data and not of
an author's restraint.

Magnitude belongs to the benchmark, not to this page. A benchmark is the proof,
measured per capability against a named workload with a stated baseline and run
Expand Down
Loading
Loading