Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 7 additions & 5 deletions ADAPTERS.md
Original file line number Diff line number Diff line change
Expand Up @@ -112,9 +112,11 @@ add/rem against the current assignees.

**Status vocabulary is per list, and casing is load-bearing.** ClickUp rejects a status whose
spelling does not match that list's own vocabulary. The adapter reads the vocabulary and sends back
the tracker's exact casing, with the fallback chain `not started` → `to do` → `todo` → first
`type=="open"`. A status that does not exist **fails loudly**, rather than leaving the card where it
was and reporting success.
the tracker's exact casing. When the model asks for `not started` (what it emits by default; every
board spells it differently), the adapter tries exact-name matches against the list's own vocabulary
in order: `to do` → `todo` → `open`. It matches on the status's name, not its ClickUp `type` field —
a list whose "open" bucket is spelled anything else won't match. A status that does not exist
**fails loudly**, rather than leaving the card where it was and reporting success.

**Auth headers differ and neither is guessable.** ClickUp takes the raw token with **no `Bearer`
prefix**. The ClickUp fake rejects a `Bearer ` prefix specifically, so an adapter that adds one fails
Expand Down Expand Up @@ -229,8 +231,8 @@ timestamp ISO and none of them epoch-zero — which is the check that matters, b
which exercises the multipart walk and the base64url decode against mail a person actually sent
rather than a fixture built to be walkable.

**Drive: live-verified, 2026-08-13**, against a sheet carrying two comments — one open, one resolved,
with a reply on the resolved one. File metadata, revisions and comments all confirmed: the name
**Drive: live-verified, 2026-08-13**, against a sheet carrying two comments, each with one reply —
one thread open, one resolved. File metadata, revisions and comments all confirmed: the name
resolves, revisions carry an author and an ISO `modifiedTime`, and the comment came back with author,
content and timestamp intact.

Expand Down
10 changes: 6 additions & 4 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -128,11 +128,13 @@ happened to say on one day; if the feature needed a model to disagree in order t
the demonstration would be the weather. But it does mean the honest claim is **"the path is proven by
test, not by recording"** — and a reader who wants to see it fire should run the tests, not the demo.

### This is what §5's "authority to write" means
### This is what "authority to write" means

PRD §5 describes the Board agent as *"the orchestrator above the role agents, holding board state and
authority to write."* Read literally that sounds like a write handle, and building it that way would
put a model in the write path and cost the guarantee the README leads with.
The internal spec this repo was built from describes the Board agent as *"the orchestrator above the
role agents, holding board state and authority to write."* Read literally that sounds like a write
handle, and building it that way would put a model in the write path and cost the guarantee the
README leads with. (That spec is private and not shipped in this repo — the quote is given here in
full so the argument stands on its own without it.)

**Production does not work that way either.** Its board agent proposes, and one script enforces the
protected-status guard, the duplicate check and read-only mode. "Authority to write" there means *its
Expand Down
11 changes: 7 additions & 4 deletions ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -209,14 +209,17 @@ only one of them involves a model at the moment of blocking.

| | Gates | Decided by |
|---|---|---|
| **Pure code over structured data** | unknown list key · assignee not in team roster · assignee not valid for list · referenced/parent/RELATE task id not on the board · subtask list ≠ parent list · RELATE self-link · evidence not cited · uncertain field(s) · vague update — card not confirmed · update — card match not confident · possible missed duplicate · registry degraded · **critical — credentials / client PII / production deploy / client-facing send** | The board, the registry, and a literal read of the manifest. No model is consulted. |
| **Pure code over structured data** | unknown list key · assignee not in team roster · assignee not valid for list · referenced task id not on the board · parent task id not on the board · RELATE link id not on the board · subtask list ≠ parent list · RELATE self-link · evidence not cited · uncertain field(s) · unresolvable field(s) · vague update — card not confirmed · update — card match not confident · possible missed duplicate · possible intra-run duplicate · over-subtasking · conflicting updates to the same card · uncategorized · registry degraded · independent verification unavailable · **critical — credentials / client PII / production deploy / client-facing send** | The board, the registry, and a literal read of the manifest. No model is consulted. |
| **A code rule over a model's stated verdict** | legitimacy — may not be a trackable task | `legitimacyHolds()` combines Pass 2b's legitimacy verdict with 2a's confidence and the source's ASR provenance — but its `not_a_task` branch fires on the verdict alone, no other input required. Genuinely a model decision, not a softened one. |
| **Two independent model reads disagreeing** | category dispute | A different *shape* of model decision — a disagreement between two reads, not a threshold on one — but not the only gate a live model verdict can decide. |

**This is the opposite of what the design anticipated.** The system this was extracted from expected
deterministic blocking to be the rare case and model judgement the norm; here thirteen of fifteen
gates never ask a model anything. That is not an accident of porting — it is what happens when the
model's job is narrowed to producing a *manifest* and every structural claim in that manifest is
deterministic blocking to be the rare case and model judgement the norm; here twenty-two of
twenty-four named gates never ask a model anything (`grep -rohE "gate: [\`'\"][^\`'\"]*[\`'\"]"
src/pipeline/gates/ src/pipeline/passes/contractCheck.ts | sort -u | wc -l`, plus the two
constant-named gates it misses — re-derive it yourself rather than trust this number, since it has
already drifted once as gates were added). That is not an accident of porting — it is what happens
when the model's job is narrowed to producing a *manifest* and every structural claim in that manifest is
checked against data the pipeline already holds.

The practical consequence: **most holds are reproducible.** Feed the same manifest and board twice
Expand Down
6 changes: 4 additions & 2 deletions LIMITATIONS.md
Original file line number Diff line number Diff line change
Expand Up @@ -328,8 +328,10 @@ them sits directly under the claim that human-in-the-loop is this system's stron

### The approval surface is a CLI, not a chat app

**The gates that decide to hold are here in full.** All fifteen of them, with the ordering, the
questions, and the persistence. What is not here is production's asking-and-answering surface: ten
**The gates that decide to hold are here in full.** All of them — twenty-two pure-code, two decided
by a model's verdict (see [ARCHITECTURE.md](ARCHITECTURE.md#human-holds) for the count and how to
re-derive it yourself) — with the ordering, the questions, and the persistence. What is not here is
production's asking-and-answering surface: ten
Block Kit modules — clarify, duplicate-resolution, critical approval, task rating — with modals,
per-actor authorisation so the wrong person cannot resolve someone else's question, and TTLs on
pending slots. Roughly 4,600 lines.
Expand Down
4 changes: 2 additions & 2 deletions PROVIDERS.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,12 +21,12 @@ written. Re-derive it with `npm run demo` and `npm run demo -- --provider anthro

| Scenario | DeepSeek | Claude | |
|---|---|---|---|
| `01-meeting-mixed` | 6 items · 4 created | 6 items · **3 created** (UPDATE 2 / SUBTASK 1 vs 1 / 2) | differs |
| `01-meeting-mixed` | 6 items · 3 created | 6 items · **3 created** (UPDATE 2 / SUBTASK 1 vs 1 / 2) | differs |
| `02-meeting-duplicates` | 2 items · 0 created | 2 items · 0 created | identical |
| `03-meeting-noise` | 0 items | 0 items | identical |
| `04-channel-messages` | 4 items · 3 created | **3 items · 2 created** | differs |
| `05-corrections` | 1 item · 1 created | 1 item · 1 created | identical |
| `06-github-activity` | 4 items · 0 created · 2 held | 4 items · **1 created · 1 held** | differs |
| `06-github-activity` | 4 items · 0 created · 2 held | 4 items · **1 created · 2 held** | differs |
| `07-email-thread` | 2 items · 1 created | **1 item** · 1 created | differs |
| `08-drive-activity` | 4 items · 2 created | **1 item · 0 created** | differs |

Expand Down
23 changes: 12 additions & 11 deletions README.md
Original file line number Diff line number Diff line change
@@ -1,16 +1,16 @@
# Triage

A production ops-agent pipeline: meeting transcripts, channel logs, GitHub activity, email threads
and document activity in — governed tracker writes out, with human-in-the-loop gates on everything
it is not sure about.
Built by [Agent Loopr](https://github.com/agentloopr). A production ops-agent pipeline: meeting
transcripts, channel logs, GitHub activity, email threads and document activity in — governed tracker
writes out, with human-in-the-loop gates on everything it is not sure about.

![Eight scenarios running offline through the real prompts, parsers and gates, then a redelivery that costs zero tokens](assets/demo.gif)

*Real captured output, replayed at reading speed — the actual run takes ~40ms. Nothing above is staged.*
*Real captured output, replayed at reading speed — the actual run takes under a second. Nothing above is staged.*

```bash
npm ci
npm run demo # 8 scenarios, offline, ~40ms, no API key
npm run demo # 8 scenarios, offline, under a second, no API key
npm run demo -- --twice # a redelivery costs zero tokens
```

Expand Down Expand Up @@ -40,8 +40,9 @@ does not exist, and the only alternative — a model grading a model — is a sy
itself. Volume and hold rate are honest; accuracy is not reported. See
[LIMITATIONS.md](LIMITATIONS.md).

It is extracted from a system that has been running in production. **The architecture is identical to
what we run; the tuned few-shot examples are replaced with generic ones.**
It is extracted from a system that has been running in production. **The core structure is what we
run** — the passes, the gates, the blind second read — **with the tuned few-shot examples replaced by
generic ones and one real generalization on top** (see below).
[EXTRACTION.md](EXTRACTION.md) records exactly what changed on the way out and why.

### What this is one half of
Expand All @@ -52,7 +53,7 @@ shapes of input:
| | Agent path | This repo |
|---|---|---|
| Input | one conversational request, ambiguous, a human present | 6–14 items, uniform policy, nobody watching |
| Decides by | a model, over a long tool-using loop | deterministic code, in passes 2a/2b |
| Decides by | a model, over a long tool-using loop | deterministic code, in two passes — categorize, then a blind re-check (the shape is below) |
| Reaches the tracker via | the same single writer | the same single writer |

**This repo is the second path**, and it is the one worth publishing: an agent is good at one
Expand Down Expand Up @@ -96,10 +97,10 @@ the real parsers and the real gates:

```bash
npm ci
npm run demo # all eight scenarios, offline, ~40ms
npm run demo # all eight scenarios, offline, under a second
npm run demo -- --twice # proves a redelivery costs zero tokens
npm run demo -- --provider anthropic # the same scenarios, replayed from a Claude recording
npm run demo -- --agents # with the agent layer on (PRD §5), also offline
npm run demo -- --agents # with the agent layer on (see AGENTS.md), also offline
```

```
Expand Down Expand Up @@ -209,7 +210,7 @@ These are the **only** paths that need credentials. `pull` plans without writing
`--write`; `poll` and `serve` write by default (`poll --dry-run` to plan only). Every fixture, test
and demo stays offline because they start from a recorded payload rather than a live read.

**An optional agent layer** (PRD §5) sits between the gates and the writer: a board agent that
**An optional agent layer** sits between the gates and the writer: a board agent that
delegates to eight role agents with **read-only** tools. It is off by default. It may **propose** a
different category, list, assignee or description — and every proposal is re-run through the same
gates, so one the gates refuse becomes a hold rather than a write. **The agent never writes, and
Expand Down
Loading