diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index 9d31fa8..7bfe237 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -46,6 +46,13 @@ jobs: # read-only guarantee would be asserted only by unit tests, never by a whole run. - run: npm run demo -- --agents - run: npm run demo -- --agents --provider anthropic + # The board agent as writer, replayed offline on both providers. Same reasoning as the two + # steps above, with more riding on it: this is the only mode in which a model reaches the + # tracker, so it is the only mode where a whole run demonstrates the governed write path + # refusing what the gates refuse. Off by default in normal use, so nothing else would exercise + # these recordings. + - run: npm run demo -- --agents --board-writes + - run: npm run demo -- --agents --board-writes --provider anthropic # Never --labels here: that path calls a judge model, and this job has no secrets. - run: npm run eval @@ -57,13 +64,14 @@ jobs: # warning. A recording made against a prompt the code no longer sends is a reply that may not # be representative — the demo's entire claim is that it replays the real prompts. # - # If this fails: either re-record (`npm run record -- --all --agents --provider

`) or work + # If this fails: either re-record (`npm run record -- --all --agents [--board-writes] --provider

`) or work # out why the prompt moved. Do not silence it. - name: no cassette drift run: | set -uo pipefail drift=0 - for mode in "" "--twice" "--provider anthropic" "--agents" "--agents --provider anthropic"; do + for mode in "" "--twice" "--provider anthropic" "--agents" "--agents --provider anthropic" \ + "--agents --board-writes" "--agents --board-writes --provider anthropic"; do # shellcheck disable=SC2086 n=$(npm run demo --silent -- $mode 2>&1 | grep -c "recorded against a different prompt" || true) echo "demo ${mode:-(default)}: $n drifted cassette(s)" diff --git a/AGENTS.md b/AGENTS.md index b81ecd4..e726172 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -136,6 +136,21 @@ happened to say on one day; if the feature needed a model to disagree in order t the demonstration would be the weather. But it does mean the honest claim is **"the path is proven by test, not by recording"** — and a reader who wants to see it fire should run the tests, not the demo. +**The same is true of the board-write recordings, and for a reason worth recording.** Across all eight +scenarios on both providers, `--board-writes` produces **zero refused writes**. Every write the board +agent originates passes the gates. + +The first recorded attempt was not like that: it produced five refusals in a single scenario, all of +them `unresolvable field(s): FINAL_DESC`. That was not the model failing the gate — `create_task` +listed `description` as optional while the gate requires it, so the tool was lying about what a valid +write looks like and the model believed it. Fixing the schema took the refusals to zero. + +Which leaves the same honest position as above: the governed write path is wired and exercised end to +end, and **no shipped recording shows a gate refusing a write.** What proves it does is +`governedTracker.test.ts` with scripted replies — an off-roster assignee, an unknown list key, a +credential-touching title, a subtask whose parent history was never read, and an operation with no +manifest form at all. Each is refused, and the inner adapter is asserted never to have seen it. + ### This is what "authority to write" means The internal spec this repo was built from describes the Board agent as *"the orchestrator above the