Skip to content

Feat/benchmark orchestrator - #265

Merged
robdmac merged 31 commits into
mainfrom
feat/benchmark-orchestrator
Jul 13, 2026
Merged

robdmac merged 31 commits into
mainfrom
feat/benchmark-orchestrator

Conversation

@robdmac

@robdmac robdmac commented Jul 13, 2026

Copy link
Copy Markdown
Contributor

No description provided.

robdmac and others added 30 commits July 12, 2026 22:56
…g work

Orcabot chat could write to terminals (terminal_send_input) and manipulate the
canvas, but had no way to READ anything back — so it was a blind launcher and
couldn't answer "how's it going?" with real data. Add a read_file tool
(workspace-scoped, read-only) that reuses the existing internal-token file API
(sandboxFetch → GET /sessions/:id/file), with a max_bytes tail for large/live
logs and a no-traversal path guard. This is the linchpin that turns chat from
blind launcher into a real orchestrator — it reads progress logs / results
(.scb-run.log, runs.jsonl, result.json) instead of guessing.

(Chose read_file over terminal_read: the scrollback endpoint requires per-PTY
X-MCP-Secret the control plane can't supply, and benchmark progress lives in
files anyway.)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
… template

- scb-visualize: watches the host-tmux executor's .scb_tmux/runs.jsonl and, for
  each agent-under-test that starts, creates a read-only Orcabot terminal tailing
  its per-run logfile with a note above naming the problem — via the local
  sandbox MCP server (create_note/create_terminal), authed with ORCABOT_MCP_SECRET
  from its own PTY (dashboard auto-resolved server-side). Read-only by
  construction: a viewer only tails a file, so it can't steer/corrupt the run.
  `run` mode wraps the benchmark; `watch` mode observes an existing run.
- SlopCodeBench template: drop the auto-opened Claude Code terminal — Orcabot
  chat is now the orchestrator. Keep the note + results blocks; rewrite the setup
  guide to drive the run via chat's tools (create_terminal / read_file /
  secrets_create) and launch through scb-visualize for the live per-agent view.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
…gler-dev)

The HTTP seeder targets the desktop d1-shim (:9001), which doesn't reach a local
wrangler-dev miniflare D1. Add `--sql` to emit the DELETE+INSERT (status
'approved', with setup_guide) for `wrangler d1 execute --local --file`. Author is
AUTHOR_ID (local D1 enforces the author_id FK, unlike D1 remote, so it must be a
real user id).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
Two failures hit during a live localhost run:
- git clone wedges on a leftover /workspace/slop-code-bench from a prior attempt
  ("destination path already exists"). Make the clone idempotent (reuse a valid
  checkout, else rm -rf + clone) and guard the venv build too.
- the runner PTY died instantly on a bad boot command (dead PTY → "failed to
  connect"), and bare `scb-visualize` isn't on PATH. Use bin/scb-visualize and
  append `; echo "[runner exited $?]"; exec bash` so failures stay visible.
Also have the orchestrator verify setup via read_file before launching instead of
declaring "done" blind.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
…), not a self-terminating boot_command

A boot_command terminal dies the moment its command finishes; when setup is
already done the idempotent setup is a <1s no-op, so the setup terminal vanished
as 'pty not found'/'failed to connect'. Have the orchestrator create a plain
shell and send setup via terminal_send_input so it stays alive.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
The sandbox image had python3/pip/venv but not uv, so slop-code-bench's
`uv sync` failed with "uv: command not found" on a freshly-built docker sandbox
(desktop worked only via a host-mounted/manually-set-up workspace). Copy the
official uv binary into /usr/local/bin so `uv sync` works out of the box after a
rebuild (localhost docker + desktop VM, whose rootfs derives from this image).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
…equence setup

terminal_send_input is fire-and-forget, so chat had no way to know when a command
finished — setup (clone, uv sync) would run while chat sat idle, with no callback
to continue. run_command runs a shell command in the sandbox workspace and BLOCKS
until it exits, returning {exit_code, output}. Chat's agentic loop then sequences
automatically: run_command(setup) -> see exit 0 -> read_file verify -> launch.

Implemented control-plane-only (no VM rebuild): wrap the command to write
stdout/stderr to a .out file and the exit code to a .exit marker (base64'd to
sidestep quoting), launch it as a headless PTY (createPty), then poll the existing
internal-token file API for the .exit marker until it appears or timeout (default
120s, max 300s). Cleans up marker files + the PTY. Long-running processes should
still use create_terminal + read_file, not this.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
…if missing

Now that run_command exists, the setup guide uses it (timeout 300s) instead of
fire-and-forget terminal_send_input — so chat blocks until clone+uv sync finish
and continues automatically instead of stalling. Also installs uv if the image
lacks it (until the Dockerfile-baked uv ships), and ends with SETUP_OK as the
success signal.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
… managers don't pop on secrets

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
The `-webkit-text-security: disc` CSS mask makes Safari classify a text
input as a password just like `type="password"` does — popping the
"save password" prompt and offering to autofill saved site credentials
(the Touch-ID thumbprint) right on the field. Since a key is pasted once
and saved secrets are already masked in the list, render the entry field
unmasked by default; add an opt-in `masked` prop for callers that accept
the manager popups.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
Replace the hardcoded splash dark-blue glass (--chat-input-bg, #e8edf5
text, #8ba3c4 icons, cyan caret, cyan hairline borders) on the chat
input/title bar and in-window compose row with theme tokens
(--background-elevated bar, --background-surface compose pill,
--foreground / --muted-foreground text, --accent-primary caret,
--border). The bar now blends with the dashboard and follows
light/dark/midnight themes. Drops the chat-input-splash autofill class
from the collapsed input (autofill already suppressed via nofill name +
one-time-code + ignore attrs). Send button stays primary blue.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
Expanded chat: make the compose input a white pill with black text in
every theme (dark/midnight would otherwise render dark-on-dark), and drop
the surrounding bar/border backgrounds so the input reads as part of the
chat window rather than sitting on its own panel. Collapsed floating pill
keeps its themed surface.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
…card

The dashboard Secrets panel value input and the chat AI-provider key card
still used a raw Input with WebkitTextSecurity:"disc", which Safari flags
as a password (save prompt + saved-credential autofill thumbprint). Swap
both to the shared unmasked SecretInput so no masked fields remain.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
run_command failed silently ("no output, no visibility"). Add revision
bump (chat-v20) plus console logs at each step: executeTool entry (tool +
arg keys + dashboard_id), run_command entry (cmd + timeout), access/sandbox
lookup misses, pty creation, poll completion (poll count + exit marker or
timeout), output byte count, and a stack trace on throw. Per CLAUDE.md:
prove where it stalls instead of speculating.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
…pup)

Masking a field via type=password or -webkit-text-security is exactly what
makes Safari/Chrome classify it as a password (save prompt + autofill
thumbprint). Instead, keep an ordinary type=text field and mask visually
with the self-hosted text-security-disc font (MIT, github.com/noppa/
text-security), embedded as a ~5KB base64 woff2 in globals.css. The browser
sees a normal text input — no popup, no thumbprint — but every glyph renders
as a bullet; input.value is still the real secret. A .secret-masked class
resets the placeholder to the normal font so it stays legible.

SecretInput now defaults masked=true again (masked={false} to reveal).
Dropped the monospace fontFamily override in AiProviderSetupCard that would
have beaten the disc font. Verify on real Safari.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
A slow/network-heavy run_command (uv install, git clone) blocks headlessly
with no chat feedback until it returns. Log a heartbeat every ~10s while
polling for the exit marker so the logs show it is still alive and when the
marker finally appears (vs a true hang).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
…un_command

The setupGuide contradicted itself — step 1 said "use run_command (headless,
blocking)" while the tool list advertised create_terminal for setup, so the
model picked whichever and the user saw an unexplained uv-sync terminal. Since
uv is now baked into the image, setup is just clone + uv sync (worth watching).
Step 1 now opens a visible "setup" terminal via create_terminal, tees output to
/workspace/.scb-setup.log, and the orchestrator confirms completion by polling
read_file(".scb-setup.log") for SETUP_OK before proceeding. No run_command.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
The chat auto-minimizes on the first pointer-down outside its panel. That's
wrong when the chat IS the primary UI: arriving from the splash chat bar, or a
template auto-kicking a setup walkthrough. Track a chatIsPrimary flag (set from
isTransitionTarget and from consuming an initial prompt — both the splash-bar
and template-walkthrough paths) and skip the auto-minimize when it's set. Manual
collapse via the chevron still works.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
A setup terminal opened but ran no visible command. Log the exact name/agentic/
boot_command the model passed to create_terminal so we can see whether the model
sent the full setup script or mangled/truncated it.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
The multi-line setup script left the model to assemble a boot_command, and it
came out as a bare cd (terminal showed a prompt, no uv sync). Replace step 1 with
a VERBATIM create_terminal call whose boot_command is a single line with no inner
double-quotes (nothing to mangle): it prints to the terminal (visible uv sync) and
writes /workspace/.scb-setup.done=SETUP_OK only on success. The orchestrator checks
the marker once via read_file; if absent it tells the user to watch the terminal
and say "go" when SETUP_OK appears — and is told NOT to claim background monitoring
(it can't), which is what made setup look stuck.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
The sandbox replays its 64KB scrollback ring buffer to a client on WS attach
(hub.go:782), but TerminalBlock wrote \x1b[2J\x1b[H (clear screen) on every
connect — wiping the replayed output of a fast boot_command that already ran
before we attached. Result: a create_terminal whose boot_command finished
quickly (e.g. an already-set-up benchmark: prints SETUP_OK in <1s) showed a
blank terminal. Now clear only on RE-connect (to avoid a duplicated replay);
the first attach keeps the replayed scrollback so the boot output is visible.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
…re-attach

The sandbox only replays its scrollback ring to RE-connecting clients (hub.go:782),
so on a terminal's very first attach a fast boot_command's output (an already-set-up
run prints SETUP_OK in <1s) is produced before the client attaches and never shown —
the terminal looked blank until a manual reload (which reconnects → replays). Prepend
`sleep 2` (plus a "== Orcabot setup ==" banner) so the client attaches during the
sleep and all setup output streams live and visible. Verified via Playwright on a
fresh dashboard: terminal now shows the banner, the ls path, and SETUP_OK.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
…paths

The sandbox file API resolves ?path= relative to the workspace root (/workspace)
via Join(root, TrimPrefix(path,"/")) — so an absolute "/workspace/x" double-nests
to /workspace/workspace/x and 404s. read_file prepended "/workspace/" and
run_command read/deleted its markers by absolute path, so BOTH could never find
any file: read_file always errored ("It may not exist yet") and the orchestrator
reported setup "still running" forever even after SETUP_OK; run_command always
timed out reading its exit marker. Pass paths relative to /workspace instead
(shell redirects keep the absolute path, correct inside the VM).

Proven against the live sandbox: path=.scb-setup.done -> "SETUP_OK" [200];
path=/workspace/.scb-setup.done -> E79710 not found [404].

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
… (P1) + honest timeout report (P2)

P1 (privilege escalation): run_command, terminal_send_input and terminal_start_agent
only checked that a dashboard_members row existed, so a read-only VIEWER could
execute arbitrary shell commands / drive terminals and mutate or delete the shared
workspace. Require role IN (owner, editor) — matching session creation. read_file /
terminal_get_sessions stay viewer-accessible (read-only).

P2 (misleading status): on run_command timeout the code deletes the PTY (which kills
the process group) but then reported "Command still running". Report it was
TERMINATED (terminated:true) and return isError:true, since it did not complete.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
…history (P2)

runs.jsonl is append-only across all runs, and `run` mode launches a fresh watcher
per benchmark run that started with an empty `seen` set — so every launch recreated
a note + a permanent `tail -F` terminal for every historical run. In `run` mode,
snapshot the pre-existing records as already-seen so only runs THIS invocation
appends get surfaced (standalone `watch` still shows everything). NOTE: also lives
in the fork (robdmac/slop-code-bench bin/scb-visualize) — push there too.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
…ion markers (P3)

P2: the disc masking font used font-display:swap, so during load (or on font
failure) the fallback readable font would flash the real secret in clear text.
Use font-display:block — text renders INVISIBLE during the (near-zero, inline)
load window instead of falling back to readable.

P3: bump revision markers on files changed this session that kept stale ones —
TerminalBlock (v9->v10 first-connect-no-clear), dashboards page, AiProviderSetupCard,
and the ASR/Matrix/WhatsApp/Telegram/GoogleChat/Teams blocks (secret-input sweep).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
…P2) + complete ASR revision marker (P3)

P2: font-display:block still falls back to a READABLE font if the disc masking
font fails to decode, exposing the type="text" secret permanently. Verify the font
is actually available (document.fonts.check after document.fonts.ready) and, if not,
fall back to the -webkit-text-security CSS mask — which hides the value independently
of any font (at the cost of a possible password-manager prompt, only in that
near-impossible failure case). Masking no longer depends on font availability.
Verified live: normal path keeps font mask + webkitTextSecurity:none (no popup),
document.fonts.check('16px text-security-disc') === true.

P3: ASRSettingsDialog had only the REVISION comment — add the runtime revision
constant + timestamped module-load log required by the coding standard.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
Previous version initialized fontOk=true, so before the verification effect ran
(after paint) a mounted field rendered via the readable fallback font — exposing
the secret for a frame if the font was already broken (repro: cached corrupt
font). And when document.fonts was unavailable it stayed fontOk=true forever
(fail open).

Now a 3-state machine that starts fail-closed:
  unknown (initial) -> color:transparent (invisible, font-independent, NO
    -webkit-text-security so no password-manager prompt for normal users)
  ok      -> disc font mask (dots, no prompt) — the steady state
  failed  -> -webkit-text-security mask (dots, hides value regardless of font;
    possible manager prompt, only in genuine failure) — also used when the Font
    Loading API is absent
Mask props are applied AFTER {...style} so a caller can't override them.

Also: inline @font-face fonts load LAZILY (only when used), so document.fonts.check
reported the font absent until it was used — a deadlock that made every field fall
to the CSS mask. Explicitly document.fonts.load() the font, then check; a corrupt
font makes load() reject -> failed. Verified live: steady state resolves to the
font mask (secret-masked, webkitTextSecurity:none).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
…eal (P2)

The fail-closed pre-verification state used an inline `color: transparent`, but the
global `::selection { color: white }` (globals.css) overrides it — selecting the
field revealed the secret (repro: programmatic select showed VISIBLE_SECRET).
`::selection` can't be styled inline, so move the pending state to a `secret-pending`
class that forces the fill transparent AND overrides ::selection / ::-moz-selection
(color + -webkit-text-fill-color transparent, !important). caret-color is
unaffected so the caret still shows.

Verified (injecting the rule against the live global selection style): pending
input has normal-fill, selection-color and selection-fill all rgba(0,0,0,0)
(hidden), while a plain input's selection is rgb(255,255,255) (the bug).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
The pending `.secret-pending` state sets the text (and -webkit-text-fill-color)
transparent, so `caret-color: auto` resolved to transparent too — the insertion
point was invisible until font verification completed (Chromium reported
caret-color rgba(0,0,0,0)). Set an explicit `caret-color: var(--foreground)` and
correct the misleading comment.

Verified: secret-pending input has caret-color rgb(232,237,245) (visible) while
-webkit-text-fill-color stays rgba(0,0,0,0) (value hidden).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
@robdmac
robdmac merged commit 972e443 into main Jul 13, 2026
1 check passed
@robdmac
robdmac deleted the feat/benchmark-orchestrator branch July 13, 2026 10:31
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant