Repository navigation
Feat/benchmark orchestrator - #265
Merged
Merged
Conversation
…g work Orcabot chat could write to terminals (terminal_send_input) and manipulate the canvas, but had no way to READ anything back — so it was a blind launcher and couldn't answer "how's it going?" with real data. Add a read_file tool (workspace-scoped, read-only) that reuses the existing internal-token file API (sandboxFetch → GET /sessions/:id/file), with a max_bytes tail for large/live logs and a no-traversal path guard. This is the linchpin that turns chat from blind launcher into a real orchestrator — it reads progress logs / results (.scb-run.log, runs.jsonl, result.json) instead of guessing. (Chose read_file over terminal_read: the scrollback endpoint requires per-PTY X-MCP-Secret the control plane can't supply, and benchmark progress lives in files anyway.) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
… template - scb-visualize: watches the host-tmux executor's .scb_tmux/runs.jsonl and, for each agent-under-test that starts, creates a read-only Orcabot terminal tailing its per-run logfile with a note above naming the problem — via the local sandbox MCP server (create_note/create_terminal), authed with ORCABOT_MCP_SECRET from its own PTY (dashboard auto-resolved server-side). Read-only by construction: a viewer only tails a file, so it can't steer/corrupt the run. `run` mode wraps the benchmark; `watch` mode observes an existing run. - SlopCodeBench template: drop the auto-opened Claude Code terminal — Orcabot chat is now the orchestrator. Keep the note + results blocks; rewrite the setup guide to drive the run via chat's tools (create_terminal / read_file / secrets_create) and launch through scb-visualize for the live per-agent view. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
…gler-dev) The HTTP seeder targets the desktop d1-shim (:9001), which doesn't reach a local wrangler-dev miniflare D1. Add `--sql` to emit the DELETE+INSERT (status 'approved', with setup_guide) for `wrangler d1 execute --local --file`. Author is AUTHOR_ID (local D1 enforces the author_id FK, unlike D1 remote, so it must be a real user id). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
Two failures hit during a live localhost run:
- git clone wedges on a leftover /workspace/slop-code-bench from a prior attempt
("destination path already exists"). Make the clone idempotent (reuse a valid
checkout, else rm -rf + clone) and guard the venv build too.
- the runner PTY died instantly on a bad boot command (dead PTY → "failed to
connect"), and bare `scb-visualize` isn't on PATH. Use bin/scb-visualize and
append `; echo "[runner exited $?]"; exec bash` so failures stay visible.
Also have the orchestrator verify setup via read_file before launching instead of
declaring "done" blind.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
…), not a self-terminating boot_command A boot_command terminal dies the moment its command finishes; when setup is already done the idempotent setup is a <1s no-op, so the setup terminal vanished as 'pty not found'/'failed to connect'. Have the orchestrator create a plain shell and send setup via terminal_send_input so it stays alive. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
The sandbox image had python3/pip/venv but not uv, so slop-code-bench's `uv sync` failed with "uv: command not found" on a freshly-built docker sandbox (desktop worked only via a host-mounted/manually-set-up workspace). Copy the official uv binary into /usr/local/bin so `uv sync` works out of the box after a rebuild (localhost docker + desktop VM, whose rootfs derives from this image). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
…equence setup
terminal_send_input is fire-and-forget, so chat had no way to know when a command
finished — setup (clone, uv sync) would run while chat sat idle, with no callback
to continue. run_command runs a shell command in the sandbox workspace and BLOCKS
until it exits, returning {exit_code, output}. Chat's agentic loop then sequences
automatically: run_command(setup) -> see exit 0 -> read_file verify -> launch.
Implemented control-plane-only (no VM rebuild): wrap the command to write
stdout/stderr to a .out file and the exit code to a .exit marker (base64'd to
sidestep quoting), launch it as a headless PTY (createPty), then poll the existing
internal-token file API for the .exit marker until it appears or timeout (default
120s, max 300s). Cleans up marker files + the PTY. Long-running processes should
still use create_terminal + read_file, not this.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
…if missing Now that run_command exists, the setup guide uses it (timeout 300s) instead of fire-and-forget terminal_send_input — so chat blocks until clone+uv sync finish and continues automatically instead of stalling. Also installs uv if the image lacks it (until the Dockerfile-baked uv ships), and ends with SETUP_OK as the success signal. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
… managers don't pop on secrets Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
…s it Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
The `-webkit-text-security: disc` CSS mask makes Safari classify a text input as a password just like `type="password"` does — popping the "save password" prompt and offering to autofill saved site credentials (the Touch-ID thumbprint) right on the field. Since a key is pasted once and saved secrets are already masked in the list, render the entry field unmasked by default; add an opt-in `masked` prop for callers that accept the manager popups. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
Replace the hardcoded splash dark-blue glass (--chat-input-bg, #e8edf5 text, #8ba3c4 icons, cyan caret, cyan hairline borders) on the chat input/title bar and in-window compose row with theme tokens (--background-elevated bar, --background-surface compose pill, --foreground / --muted-foreground text, --accent-primary caret, --border). The bar now blends with the dashboard and follows light/dark/midnight themes. Drops the chat-input-splash autofill class from the collapsed input (autofill already suppressed via nofill name + one-time-code + ignore attrs). Send button stays primary blue. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
Expanded chat: make the compose input a white pill with black text in every theme (dark/midnight would otherwise render dark-on-dark), and drop the surrounding bar/border backgrounds so the input reads as part of the chat window rather than sitting on its own panel. Collapsed floating pill keeps its themed surface. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
…card The dashboard Secrets panel value input and the chat AI-provider key card still used a raw Input with WebkitTextSecurity:"disc", which Safari flags as a password (save prompt + saved-credential autofill thumbprint). Swap both to the shared unmasked SecretInput so no masked fields remain. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
run_command failed silently ("no output, no visibility"). Add revision
bump (chat-v20) plus console logs at each step: executeTool entry (tool +
arg keys + dashboard_id), run_command entry (cmd + timeout), access/sandbox
lookup misses, pty creation, poll completion (poll count + exit marker or
timeout), output byte count, and a stack trace on throw. Per CLAUDE.md:
prove where it stalls instead of speculating.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
…pup)
Masking a field via type=password or -webkit-text-security is exactly what
makes Safari/Chrome classify it as a password (save prompt + autofill
thumbprint). Instead, keep an ordinary type=text field and mask visually
with the self-hosted text-security-disc font (MIT, github.com/noppa/
text-security), embedded as a ~5KB base64 woff2 in globals.css. The browser
sees a normal text input — no popup, no thumbprint — but every glyph renders
as a bullet; input.value is still the real secret. A .secret-masked class
resets the placeholder to the normal font so it stays legible.
SecretInput now defaults masked=true again (masked={false} to reveal).
Dropped the monospace fontFamily override in AiProviderSetupCard that would
have beaten the disc font. Verify on real Safari.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
A slow/network-heavy run_command (uv install, git clone) blocks headlessly with no chat feedback until it returns. Log a heartbeat every ~10s while polling for the exit marker so the logs show it is still alive and when the marker finally appears (vs a true hang). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
…un_command
The setupGuide contradicted itself — step 1 said "use run_command (headless,
blocking)" while the tool list advertised create_terminal for setup, so the
model picked whichever and the user saw an unexplained uv-sync terminal. Since
uv is now baked into the image, setup is just clone + uv sync (worth watching).
Step 1 now opens a visible "setup" terminal via create_terminal, tees output to
/workspace/.scb-setup.log, and the orchestrator confirms completion by polling
read_file(".scb-setup.log") for SETUP_OK before proceeding. No run_command.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
The chat auto-minimizes on the first pointer-down outside its panel. That's wrong when the chat IS the primary UI: arriving from the splash chat bar, or a template auto-kicking a setup walkthrough. Track a chatIsPrimary flag (set from isTransitionTarget and from consuming an initial prompt — both the splash-bar and template-walkthrough paths) and skip the auto-minimize when it's set. Manual collapse via the chevron still works. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
A setup terminal opened but ran no visible command. Log the exact name/agentic/ boot_command the model passed to create_terminal so we can see whether the model sent the full setup script or mangled/truncated it. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
The multi-line setup script left the model to assemble a boot_command, and it came out as a bare cd (terminal showed a prompt, no uv sync). Replace step 1 with a VERBATIM create_terminal call whose boot_command is a single line with no inner double-quotes (nothing to mangle): it prints to the terminal (visible uv sync) and writes /workspace/.scb-setup.done=SETUP_OK only on success. The orchestrator checks the marker once via read_file; if absent it tells the user to watch the terminal and say "go" when SETUP_OK appears — and is told NOT to claim background monitoring (it can't), which is what made setup look stuck. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
The sandbox replays its 64KB scrollback ring buffer to a client on WS attach (hub.go:782), but TerminalBlock wrote \x1b[2J\x1b[H (clear screen) on every connect — wiping the replayed output of a fast boot_command that already ran before we attached. Result: a create_terminal whose boot_command finished quickly (e.g. an already-set-up benchmark: prints SETUP_OK in <1s) showed a blank terminal. Now clear only on RE-connect (to avoid a duplicated replay); the first attach keeps the replayed scrollback so the boot output is visible. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
…re-attach The sandbox only replays its scrollback ring to RE-connecting clients (hub.go:782), so on a terminal's very first attach a fast boot_command's output (an already-set-up run prints SETUP_OK in <1s) is produced before the client attaches and never shown — the terminal looked blank until a manual reload (which reconnects → replays). Prepend `sleep 2` (plus a "== Orcabot setup ==" banner) so the client attaches during the sleep and all setup output streams live and visible. Verified via Playwright on a fresh dashboard: terminal now shows the banner, the ls path, and SETUP_OK. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
…paths
The sandbox file API resolves ?path= relative to the workspace root (/workspace)
via Join(root, TrimPrefix(path,"/")) — so an absolute "/workspace/x" double-nests
to /workspace/workspace/x and 404s. read_file prepended "/workspace/" and
run_command read/deleted its markers by absolute path, so BOTH could never find
any file: read_file always errored ("It may not exist yet") and the orchestrator
reported setup "still running" forever even after SETUP_OK; run_command always
timed out reading its exit marker. Pass paths relative to /workspace instead
(shell redirects keep the absolute path, correct inside the VM).
Proven against the live sandbox: path=.scb-setup.done -> "SETUP_OK" [200];
path=/workspace/.scb-setup.done -> E79710 not found [404].
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
… (P1) + honest timeout report (P2) P1 (privilege escalation): run_command, terminal_send_input and terminal_start_agent only checked that a dashboard_members row existed, so a read-only VIEWER could execute arbitrary shell commands / drive terminals and mutate or delete the shared workspace. Require role IN (owner, editor) — matching session creation. read_file / terminal_get_sessions stay viewer-accessible (read-only). P2 (misleading status): on run_command timeout the code deletes the PTY (which kills the process group) but then reported "Command still running". Report it was TERMINATED (terminated:true) and return isError:true, since it did not complete. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
…history (P2) runs.jsonl is append-only across all runs, and `run` mode launches a fresh watcher per benchmark run that started with an empty `seen` set — so every launch recreated a note + a permanent `tail -F` terminal for every historical run. In `run` mode, snapshot the pre-existing records as already-seen so only runs THIS invocation appends get surfaced (standalone `watch` still shows everything). NOTE: also lives in the fork (robdmac/slop-code-bench bin/scb-visualize) — push there too. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
…ion markers (P3) P2: the disc masking font used font-display:swap, so during load (or on font failure) the fallback readable font would flash the real secret in clear text. Use font-display:block — text renders INVISIBLE during the (near-zero, inline) load window instead of falling back to readable. P3: bump revision markers on files changed this session that kept stale ones — TerminalBlock (v9->v10 first-connect-no-clear), dashboards page, AiProviderSetupCard, and the ASR/Matrix/WhatsApp/Telegram/GoogleChat/Teams blocks (secret-input sweep). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
…P2) + complete ASR revision marker (P3)
P2: font-display:block still falls back to a READABLE font if the disc masking
font fails to decode, exposing the type="text" secret permanently. Verify the font
is actually available (document.fonts.check after document.fonts.ready) and, if not,
fall back to the -webkit-text-security CSS mask — which hides the value independently
of any font (at the cost of a possible password-manager prompt, only in that
near-impossible failure case). Masking no longer depends on font availability.
Verified live: normal path keeps font mask + webkitTextSecurity:none (no popup),
document.fonts.check('16px text-security-disc') === true.
P3: ASRSettingsDialog had only the REVISION comment — add the runtime revision
constant + timestamped module-load log required by the coding standard.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
Previous version initialized fontOk=true, so before the verification effect ran
(after paint) a mounted field rendered via the readable fallback font — exposing
the secret for a frame if the font was already broken (repro: cached corrupt
font). And when document.fonts was unavailable it stayed fontOk=true forever
(fail open).
Now a 3-state machine that starts fail-closed:
unknown (initial) -> color:transparent (invisible, font-independent, NO
-webkit-text-security so no password-manager prompt for normal users)
ok -> disc font mask (dots, no prompt) — the steady state
failed -> -webkit-text-security mask (dots, hides value regardless of font;
possible manager prompt, only in genuine failure) — also used when the Font
Loading API is absent
Mask props are applied AFTER {...style} so a caller can't override them.
Also: inline @font-face fonts load LAZILY (only when used), so document.fonts.check
reported the font absent until it was used — a deadlock that made every field fall
to the CSS mask. Explicitly document.fonts.load() the font, then check; a corrupt
font makes load() reject -> failed. Verified live: steady state resolves to the
font mask (secret-masked, webkitTextSecurity:none).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
…eal (P2)
The fail-closed pre-verification state used an inline `color: transparent`, but the
global `::selection { color: white }` (globals.css) overrides it — selecting the
field revealed the secret (repro: programmatic select showed VISIBLE_SECRET).
`::selection` can't be styled inline, so move the pending state to a `secret-pending`
class that forces the fill transparent AND overrides ::selection / ::-moz-selection
(color + -webkit-text-fill-color transparent, !important). caret-color is
unaffected so the caret still shows.
Verified (injecting the rule against the live global selection style): pending
input has normal-fill, selection-color and selection-fill all rgba(0,0,0,0)
(hidden), while a plain input's selection is rgb(255,255,255) (the bug).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
The pending `.secret-pending` state sets the text (and -webkit-text-fill-color) transparent, so `caret-color: auto` resolved to transparent too — the insertion point was invisible until font verification completed (Chromium reported caret-color rgba(0,0,0,0)). Set an explicit `caret-color: var(--foreground)` and correct the misleading comment. Verified: secret-pending input has caret-color rgb(232,237,245) (visible) while -webkit-text-fill-color stays rgba(0,0,0,0) (value hidden). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XBxVYcEidYfRm5Gf4j7nPL
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.