Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 10 additions & 2 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -46,6 +46,13 @@ jobs:
# read-only guarantee would be asserted only by unit tests, never by a whole run.
- run: npm run demo -- --agents
- run: npm run demo -- --agents --provider anthropic
# The board agent as writer, replayed offline on both providers. Same reasoning as the two
# steps above, with more riding on it: this is the only mode in which a model reaches the
# tracker, so it is the only mode where a whole run demonstrates the governed write path
# refusing what the gates refuse. Off by default in normal use, so nothing else would exercise
# these recordings.
- run: npm run demo -- --agents --board-writes
- run: npm run demo -- --agents --board-writes --provider anthropic
# Never --labels here: that path calls a judge model, and this job has no secrets.
- run: npm run eval

Expand All @@ -57,13 +64,14 @@ jobs:
# warning. A recording made against a prompt the code no longer sends is a reply that may not
# be representative — the demo's entire claim is that it replays the real prompts.
#
# If this fails: either re-record (`npm run record -- --all --agents --provider <p>`) or work
# If this fails: either re-record (`npm run record -- --all --agents [--board-writes] --provider <p>`) or work
# out why the prompt moved. Do not silence it.
- name: no cassette drift
run: |
set -uo pipefail
drift=0
for mode in "" "--twice" "--provider anthropic" "--agents" "--agents --provider anthropic"; do
for mode in "" "--twice" "--provider anthropic" "--agents" "--agents --provider anthropic" \
"--agents --board-writes" "--agents --board-writes --provider anthropic"; do
# shellcheck disable=SC2086
n=$(npm run demo --silent -- $mode 2>&1 | grep -c "recorded against a different prompt" || true)
echo "demo ${mode:-(default)}: $n drifted cassette(s)"
Expand Down
15 changes: 15 additions & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -136,6 +136,21 @@ happened to say on one day; if the feature needed a model to disagree in order t
the demonstration would be the weather. But it does mean the honest claim is **"the path is proven by
test, not by recording"** — and a reader who wants to see it fire should run the tests, not the demo.

**The same is true of the board-write recordings, and for a reason worth recording.** Across all eight
scenarios on both providers, `--board-writes` produces **zero refused writes**. Every write the board
agent originates passes the gates.

The first recorded attempt was not like that: it produced five refusals in a single scenario, all of
them `unresolvable field(s): FINAL_DESC`. That was not the model failing the gate — `create_task`
listed `description` as optional while the gate requires it, so the tool was lying about what a valid
write looks like and the model believed it. Fixing the schema took the refusals to zero.

Which leaves the same honest position as above: the governed write path is wired and exercised end to
end, and **no shipped recording shows a gate refusing a write.** What proves it does is
`governedTracker.test.ts` with scripted replies — an off-roster assignee, an unknown list key, a
credential-touching title, a subtask whose parent history was never read, and an operation with no
manifest form at all. Each is refused, and the inner adapter is asserted never to have seen it.

### This is what "authority to write" means

The internal spec this repo was built from describes the Board agent as *"the orchestrator above the
Expand Down
Loading