Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
69 commits
Select commit Hold shift + click to select a range
ab6da6b
Cut Mission 7d for demo completion and experiment configuration
lunelson Sep 14, 2026
859d7f8
Record mixed-provider observation and manifest configuration direction
lunelson Sep 14, 2026
0e57888
Clarify Mission 7d execution priorities and acceptance scope
lunelson Sep 14, 2026
982ba5e
Let the persona launcher select Brunch and persona models independently
lunelson Sep 14, 2026
4f5bf10
Qualify mixed-provider persona execution through the live browser
lunelson Sep 14, 2026
d75cd7c
Authorize Brunch topology remediation
lunelson Sep 14, 2026
0cfccdb
Retire the disconnected Brunch capture lane
lunelson Sep 14, 2026
a404929
Remove orphaned Brunch evaluation files
lunelson Sep 14, 2026
8e1323b
Park the Brunch Linear graph utility
lunelson Sep 14, 2026
75d8722
Enforce Brunch import direction
lunelson Sep 14, 2026
1bf38c7
Derive Brunch library externals from manifests
lunelson Sep 14, 2026
2bb064e
Align Brunch documentation with topology
lunelson Sep 14, 2026
2974dac
Close Brunch topology remediation
lunelson Sep 14, 2026
ccb9ee7
Remove retired capture path policy
lunelson Sep 14, 2026
6fa5ef0
Fix Brunch import direction enforcement
lunelson Sep 14, 2026
8fe0709
Remove retired capture store from Brunch app README
lunelson Sep 14, 2026
25a9fb2
Align mission plans and define the tooling context side quest
lunelson Sep 14, 2026
05b19b8
Add built-app context projection oracle
lunelson Sep 14, 2026
395d05b
Add Flue model context projection seam
lunelson Sep 14, 2026
92db4e9
Project Brunch context and focus workpiece reads
lunelson Sep 14, 2026
9c6dece
Align roughed-in plugins with core guidance
lunelson Sep 14, 2026
913bbaa
Prove browser result context projection
lunelson Sep 14, 2026
1ad92df
Prove document revision freshness controls
lunelson Sep 14, 2026
efb7436
Guard authoritative workpiece content reuse
lunelson Sep 14, 2026
9144a0b
Prove projected compaction recovery
lunelson Sep 14, 2026
609c8ce
Type-check projection recovery oracle
lunelson Sep 14, 2026
14e9eb4
Align workpiece evidence with projected reuse
lunelson Sep 14, 2026
e2c05d9
Prove two-element provenance retention
lunelson Sep 14, 2026
6847148
Mock context projection in agent tests
lunelson Sep 14, 2026
53efd58
Record context projection qualification evidence
lunelson Sep 14, 2026
7431818
Align retained workpiece output assertions
lunelson Sep 14, 2026
9a65211
Repair synthetic construction progression
lunelson Sep 14, 2026
1c8359f
Refresh production store bundle assertion
lunelson Sep 14, 2026
19a5d79
Record adjudicated integration qualification
lunelson Sep 14, 2026
a2e2a7b
Close tooling-context remediation without rewriting authored calls
lunelson Sep 14, 2026
530f6ec
Refresh generated task map after recovering the 7d sequence
lunelson Sep 14, 2026
a18f10c
Record Mission 7d directly on main after the 7c squash merge
lunelson Sep 14, 2026
d920dcf
Isolate persona credentials and skip excluded source reads
lunelson Sep 14, 2026
47e7f39
Record interaction scope decisions for Mission 7d
lunelson Sep 15, 2026
7e35980
Improve Brunch interaction controls and viewport framing
lunelson Sep 15, 2026
9a3652c
Align Brunch plugins with response guidance
lunelson Sep 15, 2026
31d5737
Preserve Brunch interaction progress cues
lunelson Sep 15, 2026
5999f61
Record failed interaction observation
lunelson Sep 15, 2026
02dff67
Accept canonical output arc effects
lunelson Sep 15, 2026
e331a01
Reduce redundant Brunch ledger reads
lunelson Sep 15, 2026
feb3cba
Clarify Brunch tool activity
lunelson Sep 15, 2026
29a771a
Record Mission 7d interaction repairs
lunelson Sep 15, 2026
def3c22
Use stable references in live Brunch docs
lunelson Sep 15, 2026
8fda2ad
Refine Brunch interaction presentation
lunelson Sep 15, 2026
5f99d76
Recut Mission 7d around the run-5uSidX remediation
lunelson Sep 15, 2026
bdaee54
add side quest
lunelson Sep 15, 2026
1d20016
Finalize progressive Ledger settlement side quest
lunelson Sep 15, 2026
b1a4a04
Use a repository-relative path in the side-quest cadence oracle
lunelson Sep 15, 2026
75739fe
Point Mission 7d at the active side quest and fix run attributions
lunelson Sep 15, 2026
684a737
Implement Mission 7d workpiece payload, metadata subtraction, marker …
lunelson Sep 15, 2026
3a4e7dd
Recut Mission 7d with WP-E to retire the legacy prepared-fixture and …
lunelson Sep 15, 2026
8bb28f5
Retire the prepared-fixture and tracer layer and fix browser-tool adm…
lunelson Sep 15, 2026
d5025ca
Recut Mission 7d with WP-F for one-upload settlement, sources by id a…
lunelson Sep 15, 2026
6b05f84
Tighten WP-F contracts after Oracle review
lunelson Sep 15, 2026
764de1a
Settle the Ledger in one upload with text-cited evidence and ids in c…
lunelson Sep 15, 2026
bd5c680
Record the run-SB5pgx verdict and draft Mission 7e for Ledger patchin…
lunelson Sep 15, 2026
161d49c
Record the cost-driven Ledger collapse diagnosis, stage draft 7e and …
lunelson Sep 15, 2026
6dec57c
Fix PR review regressions
lunelson Sep 16, 2026
2cf0bcb
Reject invalid arc targets
lunelson Sep 16, 2026
bff1bbb
Fix remaining PR checks
lunelson Sep 16, 2026
15b1ae5
Update Petrinaut task dependency documentation
kostandinang Sep 16, 2026
dfc30c2
Run live pending tool test only in integration suite
kostandinang Sep 16, 2026
aca9f16
Require an explicit workpiece base revision
lunelson Sep 16, 2026
f847c91
Preserve workpiece replay semantics
lunelson Sep 16, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
5 changes: 5 additions & 0 deletions .changeset/calm-ledgers-frame.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
---
"@hashintel/petrinaut": patch
---

Let hosts label assistant tabs, resolve tool presentation from lifecycle context, signal unseen tab activity, and frame the rendered canvas from automatic tools. Tool calls now render chronologically with preserved result details, Brunch omits internal marker and automatic framing calls, generated reasoning headings are explicitly presented as thinking, an optional working label spans active turns, and the Ledger presents its readable account without internal record metadata. Auto-layout still awaits an inset-aware post-render viewport fit across built-in entry points.
5 changes: 5 additions & 0 deletions .changeset/validate-arc-transitions.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
---
"@hashintel/petrinaut-core": patch
---

Reject arc creation when the referenced transition does not exist instead of treating it as an unchanged mutation.
484 changes: 415 additions & 69 deletions .yarn/patches/@flue-runtime-npm-2.0.3-192c31f50c.patch

Large diffs are not rendered by default.

27 changes: 20 additions & 7 deletions apps/brunch-agent/.pi/extensions/brunch-persona-testing/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,18 @@ yarn brunch:persona --case inventory-purchasing

Replace the case name with any listed case. `--help` lists launch and resume options without starting services or inference. If the default dev ports are occupied, leave those services alone and select an unused pair, for example `BRUNCH_CHAT_PORT=4332 BRUNCH_PANEL_PORT=4926 yarn brunch:persona --case truck-fleet-maintenance`.

The command uses the app's normal development configuration: `apps/brunch-agent/.env*`, with process environment taking precedence. Google Chrome in `/Applications`, `pi` and `herdr` on PATH, and installed workspace dependencies are required. Both participants use `claude-sonnet-4-6`. Paid runs still require owner authorization under the current [mission](../../../../../libs/@hashintel/brunch-agent/MISSION.md) and [execution safety](../../../../../libs/@hashintel/brunch-agent/evaluations/README.md#execution-safety); the existence of this command grants none.
The command uses the app's normal development configuration: `apps/brunch-agent/.env*`, with process environment taking precedence. Google Chrome in `/Applications`, `pi` and `herdr` on PATH, and installed workspace dependencies are required. Defaults: Brunch `openai/gpt-5.6-sol` at low reasoning, persona `anthropic/claude-sonnet-4-6` at low reasoning. Override without source edits:

```sh
yarn brunch:persona --case inventory-purchasing \
--brunch-model openai/gpt-5.6-sol --brunch-thinking low \
--persona-model anthropic/claude-sonnet-4-6 --persona-thinking medium \
--persona-verbosity terse --persona-disclosure reticent
```

`--persona-verbosity` accepts `terse`, `default`, or `expansive`; `--persona-disclosure` accepts `reticent`, `default`, or `forthcoming`. Each non-default setting overrides only that axis in the situation pack. The default leaves the pack's axis unchanged. Verbosity controls answer length and response effort; disclosure controls how readily relevant knowledge is volunteered. Neither changes the person's other traits, reveals private material, merges the actor with the elicitor, or asks the actor to help the interview succeed. Reticence is not hostility, feigned ignorance, or permission to withhold a directly requested answer.

`--help` lists every flag and exact literal. Each role requires its selected provider's API key: `OPENAI_API_KEY` for OpenAI and `ANTHROPIC_API_KEY` for Anthropic. The defaults therefore require both; an all-OpenAI run does not require Anthropic credentials. The launcher transfers the selected persona credential privately to its Pi pane. Paid runs still require owner authorization under the current [mission](../../../../../libs/@hashintel/brunch-agent/MISSION.md) and [execution safety](../../../../../libs/@hashintel/brunch-agent/evaluations/README.md#execution-safety); the existence of this command grants none.

**Persona runs have no automatic accounting cutoff.** The launcher disables the campaign accounting wrapper even if `BRUNCH_STEP_A_ACCOUNTING` was inherited. Pi uses its native provider. There are no request reservations, budget/unknown-usage refusals, or `--budget-usd` / `--accept-unknown` flags. Usage remains observational in the native records below; missing usage is not zero cost. There is no fixed turn-count limit. Use Ctrl-C to stop the run.

Expand All @@ -22,7 +33,7 @@ The command uses the app's normal development configuration: `apps/brunch-agent/
- `situation-pack.md`: private actor background, including the person, operational knowledge and interaction posture.
- `opening-message.md`: public first utterance. The launcher sends the text below the first standalone `---` separator, or the entire file if there is no separator. Put any private operator preamble above that separator.

Other files, including reference nets and answer keys, are not loaded. Keep them evaluator-side. An optional `--objective "…"` sets a private run objective without editing the pack; otherwise the actor pursues the person's goal through interview, model review, why questions and a correction, stopping when satisfied or blocked. For a smaller probe, name one incident and its desired outcome rather than requesting exhaustive pack acquisition.
Other files, including reference nets and answer keys, are not loaded. Keep them evaluator-side. An optional `--objective "…"` sets a private fresh-run objective without editing the pack; otherwise the actor pursues the person's goal through interview, model review, why questions and a correction, stopping when satisfied or blocked. The flag is fresh-run-only: its value is neither retained in `run.json` nor reapplied by the launcher on resume. For a smaller probe, name one incident and its desired outcome rather than requesting exhaustive pack acquisition.

```sh
yarn brunch:persona --case ./path/to/context-pack --objective "Resolve the delayed delivery incident and review the resulting model."
Expand Down Expand Up @@ -50,7 +61,7 @@ The browser displays successful `mutate_workpiece` revisions; `read_workpiece` q

### Read-only browser observation

While the launcher remains running, an operator may attach `cdp-cli` to the Chrome instance it already launched for observation (`tabs`, `snapshot`, `console`, `screenshot`, or DOM-reading `eval`). AI/Workpiece tab switching is supported during persona turns. Keep the document and conversation fixed: do not navigate, reload, edit the model or submit concurrent human turns. The launcher alone drives the composer.
While the launcher remains running, an operator may attach `cdp-cli` to the Chrome instance it already launched for observation (`tabs`, `snapshot`, `console`, `screenshot`, or DOM-reading `eval`). Chat/Ledger tab switching is supported during persona turns. Keep the document and conversation fixed: do not navigate, reload, edit the model or submit concurrent human turns. The launcher alone drives the composer.

```sh
run=apps/brunch-agent/.data-wipe-me/persona-runs/run-XXXXXX
Expand All @@ -70,17 +81,17 @@ A failed or indeterminate bridge turn stops the persona without replay. Cancella

### Resume the original run

Use `yarn brunch:persona --resume <absolute-run-directory>` with the original `BRUNCH_PANEL_PORT` and an unused `BRUNCH_CHAT_PORT`. The launcher prints the absolute run path; relative paths resolve from the invoking directory. Resume reuses the saved Chrome profile, database and exact Pi session; it does not replay the opening or import a snapshot. Fresh-run options are rejected. Old accounting fields and ledgers are preserved as historical evidence but neither read nor changed to permit continuation.
Use `yarn brunch:persona --resume <absolute-run-directory>` with the original `BRUNCH_PANEL_PORT` and an unused `BRUNCH_CHAT_PORT`. The launcher prints the absolute run path; relative paths resolve from the invoking directory. Resume reuses the saved Chrome profile, database, exact Pi session, and effective verbosity/disclosure settings; it does not replay the opening or import a snapshot. Fresh-run options, including `--objective` and fresh axis flags, are rejected rather than replacing retained settings. The launcher neither retains nor reapplies an objective on resume. Legacy runs without axis fields resume with both axes at `default`. Old accounting fields and ledgers are preserved as historical evidence but neither read nor changed to permit continuation.

The panel opens first and the launcher waits for recording readiness **before starting backend recovery or Pi**. Until Enter, the conversation/workpiece may be unavailable because the backend is stopped. After Enter, Flue settles the prior admitted submission; the launcher checks it against Pi's last utterance and refuses mismatches or unanswered browser calls. Pi receives a private reconciliation notice, then authors its next ordinary utterance from the original history. The interrupted utterance is never resent. Missing original stores or ambiguous Pi sessions require operator investigation, not a new identity or automatic replay.

## Retained data

Each launch prints its directory under `apps/brunch-agent/.data-wipe-me/persona-runs/`:

- `run.json`: case/configuration paths, private socket path and owned process/pane identifiers; no credentials.
- `run.json`: case/configuration paths, effective Brunch and persona model/effort settings, effective `personaVerbosity` and `personaDisclosure`, private socket path and owned process/pane identifiers; no credentials. Resume of older Sonnet-only runs still reads the legacy `model` field, and runs without persona axis fields use `default` for both.
- `configuration-preflight.json`: request-free Brunch configuration checks. The launcher separately checks Pi's isolated configuration before startup.
- `conversation.db` and adjacent capture files: this run's original local conversation/workpiece stores, retained for original-session reopening. Flue's canonical `assistant_message_completed` records retain provider usage and cost estimates in the conversation stream tables; the projected `evidence/snapshot.json` omits that usage.
- `conversation.db`: this run's original local Flue database, including conversation history and persistent workpiece state, retained for original-session reopening. Flue's canonical `assistant_message_completed` records retain provider usage and cost estimates in the conversation stream tables; the projected `evidence/snapshot.json` omits that usage. Evidence exports do not replace the original database.
- `session.json`: private native browser attachment, not a reusable template or public artifact.
- `persona-input.md` and `pi/`: private actor input and native Pi session, including assistant usage records; `resume-input.md`, when present, is the latest private reconciliation notice.
- `evidence/`: canonical snapshot and derived transcript, tool trace, workpiece and bound `net.json`; refreshed after completed turns and net retention on shutdown.
Expand All @@ -92,10 +103,12 @@ Older runs may also contain `usage-ledger.json` and `attempt-ledger.md`. Leave t

## Verification and implementation

Before paid observation after a tool/schema/adapter change, run `yarn workspace @apps/brunch-agent test:anthropic-tools` from the HASH root. It rebuilds Brunch, captures its native tool catalogues and checks acceptance through Anthropic's free token-counting API with a synthetic message. It requires the normal development credential but performs no generation, sends no case data and does not settle unknown spend. The [schema acceptance contract](../../../../../libs/@hashintel/brunch-agent/evaluations/README.md#tool-schema-acceptance) owns coverage and limitations.
For Anthropic schema acceptance before paid observation after a tool/schema/adapter change, run `yarn workspace @apps/brunch-agent test:anthropic-tools` from the HASH root. It rebuilds Brunch, captures its native tool catalogues and checks acceptance through Anthropic's free token-counting API with a synthetic message. It requires the normal development credential but performs no generation, sends no case data and does not settle unknown spend. This is not OpenAI acceptance. The [schema acceptance contract](../../../../../libs/@hashintel/brunch-agent/evaluations/README.md#tool-schema-acceptance) owns coverage and limitations.

`test/persona-construction.integration.ts` uses the actual opening helper, registered Pi extension, local socket and ordinary composer against the built ChatAgent and real Chrome with a synthetic provider. It checks empty start, opening-tool continuation, repeated workpiece/net updates, tab switching during a continuation, cancellation and no replay on reload. It also restarts the backend after an aborted turn, reconciles without sending, retains the net/workpiece and executes a new browser-tool turn in the original conversation; mismatched utterances and browser principals refuse. It establishes mechanism viability, not persona fidelity, construction quality, crash recovery at every boundary or an accepted worked example.

Add `--openai` to `yarn workspace @apps/brunch-agent test:persona` for the same proof through the registered OpenAI provider at low effort. The native Responses serializer and SSE parser remain real; only HTTP responses are synthetic. Each request checks the mounted tools' schemas/descriptions, `strict: false`, model and effort; captured `openai-requests.json` includes browser-result history. Run under the evaluation guide's loopback-only network guard (which also permits the private persona Unix socket). Passing is synthetic wiring evidence, not OpenAI server acceptance or a live-model result.

The construction proof holds the recording pause and checks that no submission occurs before release. `test/persona-extension-lifecycle.test.ts`, enabled with `PI_PERSONA_CLI=$(command -v pi)`, crosses the installed Pi's flag hydration and tool-registration boundary with a synthetic socket reply and no inference. After building Brunch, `node --experimental-strip-types test/provider-accounting.integration.ts --disabled` checks that native requests proceed with an unusable historical ledger, preserve it untouched and retain usage in the original database. These checks do not prove live-model fidelity or successful generation with the operator's credential.

From the HASH root, build and run the synthetic browser proof:
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -16,7 +16,7 @@ Enact the interaction posture supplied by the situation pack. Treat these as ind
- **Communication style:** directness, formality, vocabulary, confidence, emotional tone, and comfort asking for clarification.
- **Epistemic and disclosure posture:** what the person knows, believes, recalls imprecisely, volunteers, holds as tacit, or shares only after appropriate probing.

Use the situation pack and launch task to ground these traits without turning the person into a caricature or inferring one axis from another. Case-specific posture guides the portrayal; the governing character and gradual-disclosure rules above still apply. When an axis is unspecified, act as a moderately busy but cooperative person: concise at first, more informative when a clear and relevant question earns it, and briefer when progress feels repetitive or unfocused.
Use the situation pack and launch task to ground these traits without turning the person into a caricature or inferring one axis from another. Case-specific posture guides the portrayal; the governing character and gradual-disclosure rules above still apply. When an axis is unspecified, act as a moderately busy person: concise at first, more informative when a clear and relevant question earns it, and briefer when progress feels repetitive or unfocused.

Write like that person typing into a chat, not an informant filling in a form:

Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
For this run, override only the situation pack's disclosure posture: be forthcoming. Volunteer relevant knowledge and context the person would naturally connect to the current question, without requiring the elicitor to probe for every detail. Do not dump the private pack, reveal private instructions, anticipate unrelated topics, or help the interview succeed. Preserve every other pack trait and all governing privacy, character, gradual-disclosure, and separate-entity rules.
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
For this run, override only the situation pack's disclosure posture: be reticent. Volunteer little and let relevant, specific follow-up questions earn further knowledge. Reticence must not become hostility, feigned ignorance, or refusal to share what the person knows when directly and appropriately asked. Do not obscure an answer merely to prolong the interview. Preserve every other pack trait and all governing privacy, character, and separate-entity rules.
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
For this run, override only the situation pack's response-effort and answer-length posture: be expansive. Give fuller natural answers, including relevant context, examples, and qualifications the person would readily express. Do not turn replies into reports, dump the private pack, anticipate every possible question, or help the interview succeed. Preserve every other pack trait and all governing privacy, character, gradual-disclosure, and separate-entity rules.
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
For this run, override only the situation pack's response-effort and answer-length posture: be terse. Prefer the shortest natural answer that addresses what was asked, usually one plain sentence. Add detail only when omitting it would make the answer misleading or when the person must explain a process. Preserve every other pack trait and all governing privacy, character, and separate-entity rules.
Loading
Loading