diff --git a/ADAPTERS.md b/ADAPTERS.md index 41255dd..d8d1be7 100644 --- a/ADAPTERS.md +++ b/ADAPTERS.md @@ -112,9 +112,11 @@ add/rem against the current assignees. **Status vocabulary is per list, and casing is load-bearing.** ClickUp rejects a status whose spelling does not match that list's own vocabulary. The adapter reads the vocabulary and sends back -the tracker's exact casing, with the fallback chain `not started` → `to do` → `todo` → first -`type=="open"`. A status that does not exist **fails loudly**, rather than leaving the card where it -was and reporting success. +the tracker's exact casing. When the model asks for `not started` (what it emits by default; every +board spells it differently), the adapter tries exact-name matches against the list's own vocabulary +in order: `to do` → `todo` → `open`. It matches on the status's name, not its ClickUp `type` field — +a list whose "open" bucket is spelled anything else won't match. A status that does not exist +**fails loudly**, rather than leaving the card where it was and reporting success. **Auth headers differ and neither is guessable.** ClickUp takes the raw token with **no `Bearer` prefix**. The ClickUp fake rejects a `Bearer ` prefix specifically, so an adapter that adds one fails @@ -229,8 +231,8 @@ timestamp ISO and none of them epoch-zero — which is the check that matters, b which exercises the multipart walk and the base64url decode against mail a person actually sent rather than a fixture built to be walkable. -**Drive: live-verified, 2026-08-13**, against a sheet carrying two comments — one open, one resolved, -with a reply on the resolved one. File metadata, revisions and comments all confirmed: the name +**Drive: live-verified, 2026-08-13**, against a sheet carrying two comments, each with one reply — +one thread open, one resolved. File metadata, revisions and comments all confirmed: the name resolves, revisions carry an author and an ISO `modifiedTime`, and the comment came back with author, content and timestamp intact. diff --git a/AGENTS.md b/AGENTS.md index 1fa6279..6578638 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -128,11 +128,13 @@ happened to say on one day; if the feature needed a model to disagree in order t the demonstration would be the weather. But it does mean the honest claim is **"the path is proven by test, not by recording"** — and a reader who wants to see it fire should run the tests, not the demo. -### This is what §5's "authority to write" means +### This is what "authority to write" means -PRD §5 describes the Board agent as *"the orchestrator above the role agents, holding board state and -authority to write."* Read literally that sounds like a write handle, and building it that way would -put a model in the write path and cost the guarantee the README leads with. +The internal spec this repo was built from describes the Board agent as *"the orchestrator above the +role agents, holding board state and authority to write."* Read literally that sounds like a write +handle, and building it that way would put a model in the write path and cost the guarantee the +README leads with. (That spec is private and not shipped in this repo — the quote is given here in +full so the argument stands on its own without it.) **Production does not work that way either.** Its board agent proposes, and one script enforces the protected-status guard, the duplicate check and read-only mode. "Authority to write" there means *its diff --git a/ARCHITECTURE.md b/ARCHITECTURE.md index 1b89087..4bff3e4 100644 --- a/ARCHITECTURE.md +++ b/ARCHITECTURE.md @@ -209,14 +209,17 @@ only one of them involves a model at the moment of blocking. | | Gates | Decided by | |---|---|---| -| **Pure code over structured data** | unknown list key · assignee not in team roster · assignee not valid for list · referenced/parent/RELATE task id not on the board · subtask list ≠ parent list · RELATE self-link · evidence not cited · uncertain field(s) · vague update — card not confirmed · update — card match not confident · possible missed duplicate · registry degraded · **critical — credentials / client PII / production deploy / client-facing send** | The board, the registry, and a literal read of the manifest. No model is consulted. | +| **Pure code over structured data** | unknown list key · assignee not in team roster · assignee not valid for list · referenced task id not on the board · parent task id not on the board · RELATE link id not on the board · subtask list ≠ parent list · RELATE self-link · evidence not cited · uncertain field(s) · unresolvable field(s) · vague update — card not confirmed · update — card match not confident · possible missed duplicate · possible intra-run duplicate · over-subtasking · conflicting updates to the same card · uncategorized · registry degraded · independent verification unavailable · **critical — credentials / client PII / production deploy / client-facing send** | The board, the registry, and a literal read of the manifest. No model is consulted. | | **A code rule over a model's stated verdict** | legitimacy — may not be a trackable task | `legitimacyHolds()` combines Pass 2b's legitimacy verdict with 2a's confidence and the source's ASR provenance — but its `not_a_task` branch fires on the verdict alone, no other input required. Genuinely a model decision, not a softened one. | | **Two independent model reads disagreeing** | category dispute | A different *shape* of model decision — a disagreement between two reads, not a threshold on one — but not the only gate a live model verdict can decide. | **This is the opposite of what the design anticipated.** The system this was extracted from expected -deterministic blocking to be the rare case and model judgement the norm; here thirteen of fifteen -gates never ask a model anything. That is not an accident of porting — it is what happens when the -model's job is narrowed to producing a *manifest* and every structural claim in that manifest is +deterministic blocking to be the rare case and model judgement the norm; here twenty-two of +twenty-four named gates never ask a model anything (`grep -rohE "gate: [\`'\"][^\`'\"]*[\`'\"]" +src/pipeline/gates/ src/pipeline/passes/contractCheck.ts | sort -u | wc -l`, plus the two +constant-named gates it misses — re-derive it yourself rather than trust this number, since it has +already drifted once as gates were added). That is not an accident of porting — it is what happens +when the model's job is narrowed to producing a *manifest* and every structural claim in that manifest is checked against data the pipeline already holds. The practical consequence: **most holds are reproducible.** Feed the same manifest and board twice diff --git a/LIMITATIONS.md b/LIMITATIONS.md index 6b7919d..84f3ff7 100644 --- a/LIMITATIONS.md +++ b/LIMITATIONS.md @@ -328,8 +328,10 @@ them sits directly under the claim that human-in-the-loop is this system's stron ### The approval surface is a CLI, not a chat app -**The gates that decide to hold are here in full.** All fifteen of them, with the ordering, the -questions, and the persistence. What is not here is production's asking-and-answering surface: ten +**The gates that decide to hold are here in full.** All of them — twenty-two pure-code, two decided +by a model's verdict (see [ARCHITECTURE.md](ARCHITECTURE.md#human-holds) for the count and how to +re-derive it yourself) — with the ordering, the questions, and the persistence. What is not here is +production's asking-and-answering surface: ten Block Kit modules — clarify, duplicate-resolution, critical approval, task rating — with modals, per-actor authorisation so the wrong person cannot resolve someone else's question, and TTLs on pending slots. Roughly 4,600 lines. diff --git a/PROVIDERS.md b/PROVIDERS.md index 5443c02..1eec17d 100644 --- a/PROVIDERS.md +++ b/PROVIDERS.md @@ -21,12 +21,12 @@ written. Re-derive it with `npm run demo` and `npm run demo -- --provider anthro | Scenario | DeepSeek | Claude | | |---|---|---|---| -| `01-meeting-mixed` | 6 items · 4 created | 6 items · **3 created** (UPDATE 2 / SUBTASK 1 vs 1 / 2) | differs | +| `01-meeting-mixed` | 6 items · 3 created | 6 items · **3 created** (UPDATE 2 / SUBTASK 1 vs 1 / 2) | differs | | `02-meeting-duplicates` | 2 items · 0 created | 2 items · 0 created | identical | | `03-meeting-noise` | 0 items | 0 items | identical | | `04-channel-messages` | 4 items · 3 created | **3 items · 2 created** | differs | | `05-corrections` | 1 item · 1 created | 1 item · 1 created | identical | -| `06-github-activity` | 4 items · 0 created · 2 held | 4 items · **1 created · 1 held** | differs | +| `06-github-activity` | 4 items · 0 created · 2 held | 4 items · **1 created · 2 held** | differs | | `07-email-thread` | 2 items · 1 created | **1 item** · 1 created | differs | | `08-drive-activity` | 4 items · 2 created | **1 item · 0 created** | differs | diff --git a/README.md b/README.md index b590392..af3e318 100644 --- a/README.md +++ b/README.md @@ -1,16 +1,16 @@ # Triage -A production ops-agent pipeline: meeting transcripts, channel logs, GitHub activity, email threads -and document activity in — governed tracker writes out, with human-in-the-loop gates on everything -it is not sure about. +Built by [Agent Loopr](https://github.com/agentloopr). A production ops-agent pipeline: meeting +transcripts, channel logs, GitHub activity, email threads and document activity in — governed tracker +writes out, with human-in-the-loop gates on everything it is not sure about. ![Eight scenarios running offline through the real prompts, parsers and gates, then a redelivery that costs zero tokens](assets/demo.gif) -*Real captured output, replayed at reading speed — the actual run takes ~40ms. Nothing above is staged.* +*Real captured output, replayed at reading speed — the actual run takes under a second. Nothing above is staged.* ```bash npm ci -npm run demo # 8 scenarios, offline, ~40ms, no API key +npm run demo # 8 scenarios, offline, under a second, no API key npm run demo -- --twice # a redelivery costs zero tokens ``` @@ -40,8 +40,9 @@ does not exist, and the only alternative — a model grading a model — is a sy itself. Volume and hold rate are honest; accuracy is not reported. See [LIMITATIONS.md](LIMITATIONS.md). -It is extracted from a system that has been running in production. **The architecture is identical to -what we run; the tuned few-shot examples are replaced with generic ones.** +It is extracted from a system that has been running in production. **The core structure is what we +run** — the passes, the gates, the blind second read — **with the tuned few-shot examples replaced by +generic ones and one real generalization on top** (see below). [EXTRACTION.md](EXTRACTION.md) records exactly what changed on the way out and why. ### What this is one half of @@ -52,7 +53,7 @@ shapes of input: | | Agent path | This repo | |---|---|---| | Input | one conversational request, ambiguous, a human present | 6–14 items, uniform policy, nobody watching | -| Decides by | a model, over a long tool-using loop | deterministic code, in passes 2a/2b | +| Decides by | a model, over a long tool-using loop | deterministic code, in two passes — categorize, then a blind re-check (the shape is below) | | Reaches the tracker via | the same single writer | the same single writer | **This repo is the second path**, and it is the one worth publishing: an agent is good at one @@ -96,10 +97,10 @@ the real parsers and the real gates: ```bash npm ci -npm run demo # all eight scenarios, offline, ~40ms +npm run demo # all eight scenarios, offline, under a second npm run demo -- --twice # proves a redelivery costs zero tokens npm run demo -- --provider anthropic # the same scenarios, replayed from a Claude recording -npm run demo -- --agents # with the agent layer on (PRD §5), also offline +npm run demo -- --agents # with the agent layer on (see AGENTS.md), also offline ``` ``` @@ -209,7 +210,7 @@ These are the **only** paths that need credentials. `pull` plans without writing `--write`; `poll` and `serve` write by default (`poll --dry-run` to plan only). Every fixture, test and demo stays offline because they start from a recorded payload rather than a live read. -**An optional agent layer** (PRD §5) sits between the gates and the writer: a board agent that +**An optional agent layer** sits between the gates and the writer: a board agent that delegates to eight role agents with **read-only** tools. It is off by default. It may **propose** a different category, list, assignee or description — and every proposal is re-run through the same gates, so one the gates refuse becomes a hold rather than a write. **The agent never writes, and