Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
38 commits
Select commit Hold shift + click to select a range
62a4dd3
Cut Mission 8 for FE-1745
kostandinang Sep 16, 2026
b72ec57
Expose canonical experiment preparation
kostandinang Sep 16, 2026
2b92f62
Document the experiment preparation boundary
kostandinang Sep 16, 2026
1f5c280
Add scenario and metric mutations
kostandinang Sep 16, 2026
e1cde1b
Add experiment readiness guidance
kostandinang Sep 16, 2026
5dd7463
Add the Brunch experiment draft tool
kostandinang Sep 16, 2026
3e14899
Render session-only experiment drafts
kostandinang Sep 16, 2026
ed7c7fd
Test experiment drafts without execution
kostandinang Sep 16, 2026
18b63c1
Add the support-desk staffing persona
kostandinang Sep 16, 2026
2211616
Record the demo run and tune staffing load
kostandinang Sep 16, 2026
5dc4f7b
Fix experiment draft approval and session isolation
kostandinang Sep 16, 2026
e1d15bf
Guard experiment runs and preserve active progress
kostandinang Sep 16, 2026
d6820ca
Reserve stale-net proposals for model reads
kostandinang Sep 16, 2026
559245d
Update Mission 8 after parent merge
kostandinang Sep 17, 2026
2ae360b
Adapt experiment fixture to current optimizer context
kostandinang Sep 17, 2026
546d811
Reuse the crash recovery revision fixture
kostandinang Sep 17, 2026
f6e616c
Wait for live tool stream before asserting failure
kostandinang Sep 17, 2026
25efcc1
Prevent duplicate interactive tool continuations
kostandinang Sep 17, 2026
64374c9
Prevent stream-end duplicate tool continuations
kostandinang Sep 17, 2026
1ae5c59
Harden Petrinaut experiment drafting and mutation tracking
kostandinang Sep 18, 2026
5dc79df
Expose Petrinaut experiment host foundations
kostandinang Sep 18, 2026
8112b7d
Show drafted experiments on Simulate mode
kostandinang Sep 18, 2026
a1d1f3f
Merge Petrinaut experiment foundations
kostandinang Sep 18, 2026
7163b52
Stub ResizeObserver in draft experiment tests
kostandinang Sep 18, 2026
12354f1
Use observed model hash for experiment drafts
kostandinang Sep 18, 2026
80d7877
Retry Petrinaut docs deployment
kostandinang Sep 18, 2026
508e74b
Document experiment draft hash source
kostandinang Sep 18, 2026
8abbf4d
Show draft preparation while submission is pending
kostandinang Sep 20, 2026
ad9f843
Remove the host Simulate draft indicator slot
kostandinang Sep 21, 2026
5853245
Merge the local foundation without the draft indicator slot
kostandinang Sep 21, 2026
6712044
Keep Brunch experiment proposals in chat without a Simulate badge
kostandinang Sep 21, 2026
1a5607b
Address experiment drafting review feedback
kostandinang Sep 22, 2026
c27d034
Prevent implicit draft submission retries
kostandinang Sep 22, 2026
fab5718
Merge signed Petrinaut foundation into Brunch experiments
kostandinang Sep 22, 2026
876c101
Harden Brunch experiment draft recovery
kostandinang Sep 22, 2026
e4db0c3
Verify normalized experiment drafts run directly
kostandinang Sep 22, 2026
68a0c40
Merge remote-tracking branch 'origin/main' into kostandin/fe-1745-pet…
kostandinang Sep 22, 2026
f73bc53
Merge branch 'kostandin/fe-1745-petrinaut-foundation' into kostandin/…
kostandinang Sep 22, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .changeset/brunch-petrinaut-host-integration.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,4 +2,4 @@
"@hashintel/petrinaut": patch
---

Expose revision identity, editor commands, current-definition diagnostics, and user-guide content to host-owned assistant tools, while showing unfinished tool progress and refusing unavailable title edits. Keep host transport and store resources current across route changes, with stable React subscriptions and capability-based read-only titles.
Expose revision identity, editor commands, current-definition diagnostics, optimizer availability, and user-guide content to host-owned assistant tools, while showing unfinished tool and experiment progress, refusing unavailable title edits, and rejecting unavailable optimization before an experiment record is created. Keep host transport and store resources current across route changes, with stable React subscriptions and capability-based read-only titles.
5 changes: 5 additions & 0 deletions .changeset/expose-prepare-experiment.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
---
"@hashintel/petrinaut": patch
---

`prepareExperiment` is exported from `@hashintel/petrinaut/react`, so a host can resolve an experiment request against the live model and build its optimization input without starting a run.
5 changes: 5 additions & 0 deletions .changeset/selected-batch-scenarios-metrics.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
---
"@hashintel/petrinaut-core": patch
---

`selectedMutationBatchSchema` admits `addScenario`, `updateScenario`, `removeScenario`, `addMetric`, `updateMetric` and `removeMetric`, so a selected batch can save the scenario and metric an experiment names.
2 changes: 1 addition & 1 deletion apps/brunch-agent/src/agents/chat-agent/agent.ts
Original file line number Diff line number Diff line change
Expand Up @@ -227,7 +227,7 @@ export function ChatAgent({ id }: AgentProps) {
Call ping when you need to confirm the server tool path.
Submit at most one browser tool call per proposal, separately from server tools, and wait for its correlated client result before further browser work. Invalid proposals fail as a whole; do not rely on sibling execution order.
A client-tool-result signal is JSON [{ toolCallId, toolName, output, metadata? }]. Treat output as the browser's canonical result for that call and continue helping the user once; never reapply a completed mutation. For a joined root arc, metadata.mutationRecord contains verified observations and effects, not assistant prose or user testimony. Failed, stale, no-op and unknown attempts are not causes.
A ${NET_STALE_SIGNAL} signal at the start of a user turn means this conversation holds no verified read of the net now open in Petrinaut, or the net changed after your last verified read. When it is present, call ${readPetrinautNetToolName} in its own proposal and wait for its browser result before explaining, reviewing, interviewing about, or changing the model, and do not say the net is unavailable or ask for an upload or description. When it is absent, the most recent ${readPetrinautNetToolName} result in this conversation is the current net.
A ${NET_STALE_SIGNAL} signal at the start of a user turn means this conversation holds no verified read of the net now open in Petrinaut, or the net changed after your last verified read. When it is present, make ${readPetrinautNetToolName} the entire proposal: do not call activate_skill, read_skill_resource, or any other tool in the same proposal. Wait for its browser result before activating required skills, explaining, reviewing, interviewing about, or changing the model, and do not say the net is unavailable or ask for an upload or description. When it is absent, the most recent ${readPetrinautNetToolName} result in this conversation is the current net.
`.replace(/^\s+|\s+$/gu, ""),
);
if (browserContext)
Expand Down
7 changes: 7 additions & 0 deletions apps/brunch-agent/src/agents/chat-agent/tool-catalogue.ts
Original file line number Diff line number Diff line change
Expand Up @@ -16,6 +16,7 @@ export interface OrdinaryBrunchToolCatalogueEntry {
| "workpiece"
| "petrinaut-read"
| "petrinaut-mutation"
| "petrinaut-experiment-draft"
| "explanation"
| "diagnostic";
}
Expand Down Expand Up @@ -83,6 +84,12 @@ export const ordinaryBrunchToolCatalogue: readonly OrdinaryBrunchToolCatalogueEn
executionOwner: "petrinaut-website",
role: "petrinaut-mutation",
},
{
name: "draft_petrinaut_experiment",
definitionOwner: "sdcpn-plugin",
executionOwner: "petrinaut-website",
role: "petrinaut-experiment-draft",
},
{
name: "read_workpiece",
definitionOwner: "brunch-core",
Expand Down
4 changes: 3 additions & 1 deletion apps/brunch-agent/src/conversation/net-ledger.ts
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,7 @@
*/
import {
canonicalContent,
isDraftPetrinautExperimentToolName,
isLayoutPetrinautNetToolName,
isMutatePetrinautNetToolName,
isReadPetrinautDocsToolName,
Expand Down Expand Up @@ -92,7 +93,8 @@ const resultMessages = (snapshot: FlueConversationSnapshot) =>
const isNonMutatingBrowserTool = (name: string): boolean =>
isReadPetrinautNetToolName(name) ||
isReadPetrinautDiagnosticsToolName(name) ||
isReadPetrinautDocsToolName(name);
isReadPetrinautDocsToolName(name) ||
isDraftPetrinautExperimentToolName(name);

/** A model-selected ID selects a recorded browser observation, never a model-supplied hash. */
export const recordedBrowserObservation = async (
Expand Down
1 change: 1 addition & 0 deletions apps/brunch-agent/src/conversation/why.ts
Original file line number Diff line number Diff line change
Expand Up @@ -415,6 +415,7 @@ export const queryWorkpiece = async (input: {
case "differential-equation":
case "type":
case "scenario":
case "metric":
return locateRootState(definition, {
kind: target.kind,
name: target.id,
Expand Down
14 changes: 13 additions & 1 deletion apps/brunch-agent/test/chat-agent-compaction.test.ts
Original file line number Diff line number Diff line change
@@ -1,3 +1,4 @@
import { useInstruction } from "@flue/runtime";
import { afterEach, beforeEach, expect, test, vi } from "vitest";

import { useBrunchAgent } from "@hashintel/brunch-agent/flue";
Expand All @@ -19,7 +20,7 @@ vi.mock(
);
vi.mock("@flue/runtime", async (importOriginal) => ({
...(await importOriginal<typeof import("@flue/runtime")>()),
useInstruction: () => undefined,
useInstruction: vi.fn<typeof useInstruction>(),
useContextProjection: () => undefined,
useInitialData: () => undefined,
useDelivery: () => ({ kind: "user", body: "test" }),
Expand Down Expand Up @@ -60,6 +61,17 @@ test("the production ChatAgent supplies no compaction override when unset", asyn
);
});

test("a stale-net turn reserves the proposal for the browser read", async () => {
const { ChatAgent: renderChatAgent } =
await import("../src/agents/chat-agent/agent.ts");
renderChatAgent({ id: "test-instance" });
expect(useInstruction).toHaveBeenCalledWith(
expect.stringContaining(
"do not call activate_skill, read_skill_resource, or any other tool in the same proposal",
),
);
});

test("the production ChatAgent forwards an independent OpenAI specifier and thinking level", async () => {
vi.stubEnv("BRUNCH_CHAT_MODEL", "openai/gpt-5.6-sol");
vi.stubEnv("BRUNCH_CHAT_THINKING", "low");
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -20,6 +20,8 @@ import {
import { createFlueClient } from "@flue/sdk";

import {
draftPetrinautExperimentInputSchema,
draftPetrinautExperimentToolName,
queryWorkpieceInputSchema,
mutatePetrinetInputSchema,
mutatePetrinautNetToolName,
Expand Down Expand Up @@ -330,6 +332,77 @@ try {
);
assert(tools.length > 0);
for (const tool of tools) assert.deepEqual(tool.input_schema, expected);

// The session experiment draft carries core's request schema natively:
// the sent schema must be byte-identical to the Zod source and must still
// refuse a constraint (there is no constraint carriage) once serialized.
const draftTool = ordinaryRequest.serialized.tools.find(
(tool) => tool.name === draftPetrinautExperimentToolName,
);
assert(draftTool, `${method} must carry the experiment draft tool`);
assert.deepEqual(
draftTool.input_schema,
draftPetrinautExperimentInputSchema["~standard"].jsonSchema.input({
target: "draft-2020-12",
}),
);
const draftExperiment = {
name: "Synthetic staffing",
scenarioId: "scenario-peak",
scenarioParameterValues: { agents: { mode: "range", min: 2, max: 8 } },
runCount: 20,
seed: 1,
dt: 0.5,
maxTime: 120,
metricIds: ["metric-wait"],
execution: {
mode: "optimize",
objectiveMetricId: "metric-wait",
direction: "minimize",
steps: 5,
runsPerStep: 4,
},
};
const draftEnvelope = {
observation: { toolCallId: "read-1", baseHash: "a".repeat(64) },
declarations: [
{ subject: "maxTime", statement: "120 model minutes: the peak." },
],
basis: { kind: "absent", reason: "Synthetic control." },
unsupported: [],
};
const validateDraft = (arguments_: Record<string, unknown>): void => {
validateToolArguments(
{
name: draftTool.name,
description: "Captured draft tool",
parameters: draftTool.input_schema as Tool["parameters"],
},
{
type: "toolCall",
id: "draft-schema-control",
name: draftTool.name,
arguments: arguments_,
},
);
};
assert.doesNotThrow(() =>
validateDraft({ ...draftEnvelope, experiment: draftExperiment }),
);
assert.throws(() =>
validateDraft({
...draftEnvelope,
experiment: { ...draftExperiment, constraints: [] },
}),
);
assert.throws(() =>
validateDraft({
...draftEnvelope,
declarations: [],
experiment: draftExperiment,
}),
);
assert.throws(() => validateDraft(draftEnvelope));
}
assert.equal(networkAttempts, 0);
process.stdout.write(
Expand Down
20 changes: 20 additions & 0 deletions apps/brunch-agent/test/net-ledger.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,7 @@ import { createHash } from "node:crypto";
import { describe, expect, test } from "vitest";

import {
draftPetrinautExperimentToolName,
layoutPetrinautNetToolName,
deriveMutationEffects,
mutatePetrinetInputSchema,
Expand Down Expand Up @@ -338,6 +339,25 @@ describe("the net ledger is a projection over Flue history", () => {
expect(JSON.stringify(snapshot)).toBe(before);
});

test("ignores a completed experiment draft because it cannot change the net", async () => {
const events = await deriveNetLedger(
snapshotOf([
...readTurn("read-1", emptyNet),
assistantCall("draft-1", draftPetrinautExperimentToolName),
resultDelivery("draft-1", draftPetrinautExperimentToolName, {
status: "drafted",
summary: "Prepared only",
diagnostics: [],
}),
]),
browser,
);

expect(events.map((event) => [event.kind, event.toolCallId])).toEqual([
["read", "read-1"],
]);
});

test("introduces no identities: every event names a tool call the assistant made, at its own position", async () => {
const snapshot = fullHistory();
const calls = assistantCallsOf(snapshot);
Expand Down
Original file line number Diff line number Diff line change
@@ -1,4 +1,5 @@
import {
draftPetrinautExperimentToolName,
layoutPetrinautNetToolName,
mutatePetrinautNetToolName,
READ_PETRINAUT_DOCS_TOOL_NAME,
Expand All @@ -16,24 +17,45 @@ export const brunchClientToolNames: ReadonlySet<string> = new Set([
READ_PETRINAUT_DOCS_TOOL_NAME,
]);

/** Batched construction: reads, one batch mutation and layout. */
/** Batched construction: reads, one batch mutation, layout and one session draft. */
export const batchedConstructionClientToolNames: ReadonlySet<string> = new Set([
...brunchClientToolNames,
readPetrinautNetToolName,
readPetrinautDiagnosticsToolName,
mutatePetrinautNetToolName,
layoutPetrinautNetToolName,
draftPetrinautExperimentToolName,
]);

/**
* The Brunch-named tools are host dynamic tools in every mode: the transport
* projects them as `dynamic-tool` parts and the panel routes them to the
* host's automatic tools rather than to Petrinaut's static registry.
* The Brunch-named tools the browser executes without asking: the panel
* routes them to the host's automatic tools.
*/
export const brunchPetrinautDynamicToolNames: ReadonlySet<string> = new Set([
export const brunchPetrinautAutomaticToolNames: ReadonlySet<string> = new Set([
READ_PETRINAUT_DOCS_TOOL_NAME,
readPetrinautNetToolName,
readPetrinautDiagnosticsToolName,
layoutPetrinautNetToolName,
mutatePetrinautNetToolName,
]);

/**
* The Brunch-named tools the browser renders as a widget the person acts on:
* the panel routes them to the host's interactive tools. The experiment draft
* auto-submits its prepared state so Brunch's turn continues, and keeps its
* Run and Dismiss actions for the person.
*/
export const brunchPetrinautInteractiveToolNames: ReadonlySet<string> = new Set(
[draftPetrinautExperimentToolName],
);

/**
* The Brunch-named tools are host dynamic tools in every mode: the transport
* projects them as `dynamic-tool` parts and the panel routes them to the
* host's automatic or interactive tools rather than to Petrinaut's static
* registry.
*/
export const brunchPetrinautDynamicToolNames: ReadonlySet<string> = new Set([
...brunchPetrinautAutomaticToolNames,
...brunchPetrinautInteractiveToolNames,
]);
Loading
Loading