Record the board-write cassettes, and fix the two minor bugs - #21
Merged
Merged
Conversation
BOARD_AGENT_WRITES shipped with no recordings, so the one mode that shows an agent being refused by a gate was the one mode nobody could watch. A flag that is documented and never exercised is the same shape as the four unreachable modules this repo has already shipped. Now it replays offline on both providers, no key. Recording one scenario before all sixteen paid for itself twice: - create_task listed `description` as optional while the gate requires it. The tool was lying to the model about what a valid write looks like, and the model believed it: five refusals in one run, every one the schema's fault. Required now, and the parent_id description says to read the parent's comments first because the evidence gate will ask. - Operations were attributed to planned actions by exact signature only. The agent is invited to reword, so it commented on the right card in its own words, the signature missed, and Pass 2d reported "no comment landed on t200" for a comment plainly visible in the trace landing on t200. The write was correct and the attribution was wrong, which is worse than the reverse: it makes a working run look broken and teaches a reader to discount the audit. Matching now falls back to the target card, leaving Pass 2d free to ask its own question. What the recordings actually show, which is why they were worth buying: the agent reads cards and their history before writing, then departs from the plan where the board tells it something the pipeline could not — commenting on the duplicate's card rather than silently skipping it, and filing the email copy as a subtask of the redesign it had just read. 55 new cassettes, both providers. Zero drift across all seven demo modes: the tool list is part of the prompt fingerprint, so that zero is what proves the write tools did not leak into the read-only path.
3 tasks done
digitalmasterykit-rgb
added a commit
that referenced
this pull request
Aug 26, 2026
* Run the board-write modes in CI, so their recordings cannot rot PR #21 added 55 cassettes for the mode where the board agent does the writing. Nothing exercised them: CI ran four demo modes and the board-write pair was not among them, and the mode is off by default so no ordinary run reaches it either. That is the exact failure this job's own comment warns about two lines above — the Anthropic set "nearly rotted" the same way, behind docs that cited it. A recording nothing replays is a recording nothing can tell you has stopped matching the prompt. More rides on these two than on the others. This is the only mode in which a model reaches the tracker, so it is the only mode where a whole run demonstrates the governed write path refusing what the gates refuse — asserted by unit tests otherwise, never end to end. Both modes added to the demo steps and to the drift loop. Verified by running the drift check exactly as CI runs it: seven modes, zero drifted cassettes, and both new steps exit 0. * Say plainly that no board-write recording shows a gate refusing I described this mode as the one that demonstrates an agent being refused by a gate. Checked it: across all eight scenarios on both providers, --board-writes produces zero refused writes. The claim was wrong. The first recording did show five refusals, which is where the impression came from — but they were all `unresolvable field(s): FINAL_DESC`, and that was the tool spec's fault, not the gate catching a bad model. create_task listed `description` as optional while the gate requires it. Fixing the schema took refusals to zero, which is the correct outcome and also removes the only footage of a refusal. So the same disclosure the role-agent recordings already carry now applies here: the path is wired and exercised end to end, no recording shows it firing, and what proves it is governedTracker.test.ts with scripted replies — off-roster assignee, unknown list key, credential-touching title, uncited subtask, and an op with no manifest form, each refused with the inner adapter asserted never to have seen it. A green --board-writes run is evidence the mechanism runs, not evidence it bites. --------- Co-authored-by: harishghasolia07 <100846446+harishghasolia07@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
BOARD_AGENT_WRITESshipped in #19 with no recordings, so the one mode that demonstrates an agentbeing refused by a gate was the one mode nobody could watch. A flag that is documented and never
exercised is the same shape as the four unreachable modules this repo has already shipped.
It now replays offline on both providers, no API key:
Recording one scenario before all sixteen paid for itself twice
1. The tool spec was lying to the model.
create_tasklisteddescriptionas optional; the gaterequires it. Five refusals in one recorded run — every one the schema's fault, not the model's. After
the fix: 0 refusals, 3 created, 2 commented.
2. Correct writes were reported as failures. Ops were attributed to planned actions by exact
signature only. The agent is invited to reword, so it commented on the right card in its own words,
the signature missed, and Pass 2d reported
no comment landed on t200— for a comment plainly visiblein the trace landing on t200. The write was right and the attribution was wrong, which is the worse
failure to ship: it makes a working run look broken and trains a reader to discount the audit.
Matching now falls back to the target card, so Pass 2d still gets to ask its own question
independently.
What the recordings show
This is why they were worth buying — the agent investigates before it writes, then departs from the
plan where the live board tells it something the pipeline could not:
That last one is the interesting departure: the plan said skip as duplicate; the agent left a note
on the card instead.
Verification
npm test— 1018 passing ·npx tsc --noEmit·npm run lintall cleanfingerprint, so that zero is the proof the write tools did not leak into the read-only path
DEEPSEEK_API_KEYandANTHROPIC_API_KEYunset