Specstride. From specs to tested code. An autonomous coding orchestrator that steers your agent through implementation, critic review, and verification, phase by phase: a spec-driven agent loop with a critic gate (an LLM-as-a-judge that approves a phase only on cited evidence), wrapped in an outer loop that tunes its own budgets from its telemetry, checks whether each change helped, and rolls it back when it did not.
Every step, signed off. The mark is the track folded into an S; the rail under it shows the feature's real phases, and each gate opens only on approved evidence (animated version). The live timeline re-stamps that rail each time a gate opens or holds. The inner loop is the Ralph technique (fresh context every pass); Specstride adds the gate that holds each phase until its evidence is approved.
Try it in a minute (needs Python 3.13, Bash, the Claude Code CLI for
the proposer and ANTHROPIC_API_KEY for the critic, or swap in codex and OPENAI_API_KEY):
git clone https://github.com/mairp/specstride && cd specstride
mkdir -p /tmp/demo/specs/001-greeting && cp examples/speckit-tasks.example.md /tmp/demo/specs/001-greeting/tasks.md
SPECSTRIDE_PROPOSER=claude SPECSTRIDE_CRITIC=claude ./specstride run -w /tmp/demo --verification off--verification off skips the pre-loop test plan: the default (required) needs a test command it can discover, and the empty demo directory has none yet, so it would stop with exit 3. Leave verification on in a real project.
Then run ./specstride watch -w /tmp/demo in another terminal to follow it. The full setup is in Quick start.
You hand it a GitHub Spec Kit feature — its
tasks.md, an ordered set of phases, each a list of tasks — and it drives a coding
agent phase by phase, but nothing advances until a critic approves it. The
feature's spec.md, plan.md and contracts ride along as read-only context for both
the agent and the critic. The human who used to eyeball each phase and click
"approved" is replaced by an LLM judge, the critic. You stay out of the inner loop;
you only arbitrate the phases the machines genuinely can't settle.
Specstride runs a coding agent in a repeating, self-checking loop: each pass is a fresh, stateless session that works until the phase's evidence file exists. On top of that plain loop it adds three things:
- An automated critic gate. An LLM critic checks each phase's evidence against the spec's acceptance criteria and the real code. Nothing advances until the critic approves it.
- A diagnose-and-accelerate micro-loop for stuck phases. A plain retry loop retries a rejected phase from scratch, so a stuck phase can burn pass after pass and learn nothing new. When a phase stalls on a new set of unmet criteria, Specstride runs the diagnostician. It reads the full rejection history and the untruncated files, then says whether the critic simply couldn't see the code or the gap is real, and what to change. The accelerator acts on that hint: a narrowed retry that fixes only the failing criteria and leaves approved work untouched. On one real 12-task phase, a full retry had spent 16 minutes re-deriving what turned out to be a two-file fix.
- An opt-in learning loop over its own runs. Every pass is recorded in
events.jsonl.specstride learnreads that history and suggests per-phase settings sized to what each phase actually measured, instead of one global setting sized for the worst phase. You apply a suggestion explicitly; it is bounded, keyed to the phase as written (edit the phase and the old decision stops applying), and reversible. After it is applied, every closed phase evaluates it against the baseline it was learned from —helped,neutral,regressedorinsufficient, always printed beside the smallest effect the data could show — and a guardrail breach (more MALFORMED verdicts, a new verification failure, a shift in first-attempt approval, …) reverts it automatically and quarantines the value. A pass that rewrites the telemetry it is judged by is excluded. It can never touch anything the critic reads. See Learning.
| Loop | Scope | What repeats | What carries over |
|---|---|---|---|
| Pass loop (inner) | one phase attempt | a fresh, stateless agent session per pass, until the phase's evidence file exists | only what is on disk |
| Gated phase loop (middle) | one run | proposer → critic → approve or retry; a stuck phase gets the diagnostician and a narrowed accelerator retry | the feedback and hint files, within the run |
| Learning loop (outer) | across runs | measure every pass, suggest (and, opt-in, apply) per-phase settings, then evaluate each applied one against the baseline it was learned from, reverting it automatically if a guardrail breaks | a per-phase observation and an append-only log of decisions, their baselines and their evaluations |
The agents never improve; they stay stateless workers. What improves is how the loop drives them. The learning loop is deliberately narrow: it can tune two allowlisted settings, it is off unless you turn it on, and the critic's independence is locked by a test. A loop that could tune its own judge would drift toward approving its own work.
You write the spec; the loop takes it to verified code. You step in only for the phases the machines genuinely can't settle.
📖 Full documentation lives in the
wiki/folder — start atwiki/Home.md. It covers the architecture, getting started, the CLI, spec formats, the on-disk contract, hardening, telemetry, and configuration.
mixture-of-loops (MoL) is an Agent Skills package for Claude Code, Codex, dsh, pi and prime. You point it at a Spec Kit feature, it reads all of the feature's artifacts, and it writes a launch contract plus an unattended Bash launcher that runs Specstride for you:
/mixture-of-loops derive a pipeline for specs/007-example # generate only
/mixture-of-loops derive a pipeline for specs/007-example and run it # generate, gate, run
Specstride answers "is this phase really done?". MoL answers "is this the right run to start, and did it finish?". Together:
| On its own, Specstride… | With MoL on top |
|---|---|
| runs whatever verification commands you pass it | the commands come only from what plan.md and tasks.md literally declare, each traced from source phase to Specstride gate, so a check the plan named can't be silently skipped. Ambiguous or conflicting declarations become blockers, not guesses |
| starts from the flags you type | starts from a validated contract that hashes every source file; a changed spec makes the launcher refuse instead of running against stale inputs |
stops on a transient failure and waits for you to specstride resume |
a supervisor relaunches only stops the run's own records classify as transient, within a budget the contract declares, and announces each relaunch |
leaves you to read run.log or specstride watch |
the harness reports progress read from the run's telemetry ([MOL-STATE]), never from the model's narration, and ends with one digest: stage, evidence paths, next action |
applies learned settings from learning/applied.json |
the contract pins which learned decisions the run may use; a newer specstride learn --apply waits for a new contract, and supervise.py retro reports how they evaluated without ever changing them |
The split keeps judgment where it can be checked: MoL records facts from the files and
refuses to render a launcher from an unvalidated contract, and Specstride's critic, its LLM
judge, still decides every phase. Install the skill with ./bin/onboard-skill --harness all --scope user
from a MoL checkout. How learning is shared between the two is described in
Under mixture-of-loops: who decides what.
Specstride gets better at driving its agents by learning from its own runs, and it
checks every change it learns before keeping it. The agents themselves never change;
what improves is how the loop sizes and paces their work. The whole layer is off
unless you set SPECSTRIDE_LEARNING.
flowchart LR
O[Observe<br/>every pass → events.jsonl] --> S[Suggest<br/>per-phase value + evidence]
S --> A[Apply<br/>you opt in; bounded, reversible]
A --> E[Evaluate<br/>at every closed phase,<br/>against the recorded baseline]
E -->|guardrail breach| R[Auto-revert<br/>+ quarantine]
E -->|helped / neutral / regressed| K[Keep, and report]
R --> O
K --> O
- Observe. Every pass is measured: wall time, time spent waiting, cost, how it ended. At each approved phase, a per-phase observation is written next to the feature's state.
- Suggest.
specstride learn --showproposes a value for each phase from what that phase actually measured, with the sample count behind it. It needs at least three samples, and ignores passes killed for going nowhere. - Apply. You apply one suggestion at a time (
specstride learn --apply). Each decision is clamped to a hard range, moves at most ±50 % per step, is filed under the phase as written (edit the phase and the decision stops applying), and records the baseline it was learned from.--revertand--offundo it. - Evaluate. After every approved phase, each applied decision is compared with its
baseline and labelled
helped,neutral,regressedorinsufficient. The effect is always printed beside the smallest effect the data could have shown, because runs are few and a small change is invisible at that scale. - Keep or roll back. Guardrails can veto a decision even when it made things cheaper: more malformed critic verdicts, more ungrounded citations, a new verification failure, bigger critic inputs, more phases handed back to a human, or a shift in first-attempt approval in either direction. A breach reverts the decision automatically, the only change the loop ever makes on its own, and blocks that value until there is fresh evidence.
Two knobs can move: the per-phase pass ceiling (proposer_timeout) and the yield poll
interval (yield_poll_interval). Nothing the critic reads, no breaker, and nothing in the
verification commands can ever be tuned; tests lock both the list of knobs and every
place the learning code is called from. A pass that rewrites the telemetry it is judged
by is excluded from every evaluation. Run through a mixture-of-loops contract, the
decisions a run may use are bound into the contract, so a later change waits for a new
contract instead of slipping into a relaunch.
Details, formulas and file formats: Learning below and
wiki/Learning.md.
When a mixture-of-loops contract drives the run, the harness running the skill does not decide how the loop improves. It records two choices when it derives the contract, and after that it only reads:
| Step | Who decides | Where it shows |
|---|---|---|
Learning mode (off, suggest, apply) |
the operator, in the request; the deriving harness writes it literally into each specstride stage's env (default off; a value left in the shell is stripped) |
the contract, so its digest records it |
| Which applied decisions the run may use | bound at derivation: configuration.learning pins the decision log up to one run id, with a hash (SPECSTRIDE_LEARNING_THROUGH) |
the contract; a later --apply waits for a new contract |
| Measure, suggest, evaluate, auto-revert on a guardrail breach | Specstride itself, inside the run | learning/, knob_evaluated, knob_auto_reverted |
| Applying a new decision | the operator (specstride learn --apply); never the harness, never automatic |
learning/applied.json |
| Relaunching after a transient stop | the supervisor, within the contract's budget; refused with learning-decisions-changed if the bound decisions changed |
the supervisor's report |
| Retrospective | supervise.py retro reads what Specstride measured and may suggest specstride learn --revert; it never applies or reverts |
runs/<id>/retrospectives/<digest>.json |
Install Specstride once (clone it wherever you keep tools); it is not the working directory. Each run points at your project:
-w/--workdir DIR— where the proposer works. All generated state lives under.specstride/features/<slug>/(gates, evidence, PROGRESS.md, verdicts), so the workdir root holds only your real artifacts. Default:$PWD.-s/--specs FILE— the spec, normally a Spec Kit feature'sspecs/<feature>/tasks.md. Inside a Spec Kit project you can leave it out: the feature is auto-discovered. A relative path resolves against the directory you launched from, not the workdir.--feature SLUG— the feature namespace for durable state (.specstride/features/<slug>/), for repos with more than one Spec Kit feature. AlsoSPECSTRIDE_FEATURE. Default: the feature dir's basename, ordefault.
So the same installed Specstride drives any project:
specstride run -w ~/projects/foo # one Spec Kit feature: discovered
specstride run -w ~/projects/foo -s ~/projects/foo/specs/002-billing/tasks.md(specstride is the single front-door command — see Install it permanently
just below.)
An existing codebase can be turned into the Spec Kit feature Specstride takes as input, describing the system as it is:
specstride reverse ~/projects/foo --dry-run # inventory, output dir, phases, argv; no LLM
specstride reverse ~/projects/foo # writes ~/projects/foo/specs/NNN-as-is-foo/It reads the source and never changes it. A deterministic inventory settles what
exists; a generated five-phase driver writes spec.md, plan.md, research.md,
data-model.md, contracts/, quickstart.md and tasks.md one gate at a time; and
every gate runs a deterministic Spec Kit linter plus a guard that fails if anything
outside the output and .specstride/ changed. Every requirement cites the lines it
came from in an invisible <!-- evidence: path:40-88 | kind=observed | … -->
comment the linter resolves. Discovery (the target's own tests) and git checkpoints
are off for these runs. With --tasks done (the default) every task is ticked, so
the result is a baseline a later run extends; --tasks open gives a rebuild plan.
specstride reverse --lint DIR lints any Spec Kit directory, hand-written ones
included. See the wiki's Reverse Engineering page.
Specstride can derive a Lisa-compatible VerificationPlan v1 before the first
proposer pass. The canonical JSON is hash-bound to the authoritative
specification, while TEST_PLAN.md is its human-readable projection.
--verification requiredis the default. It creates the plan, injects its obligations into proposer and critic, runs fixed-argv tests before each approval, and runs a cumulative release gate (including on an already-approved resume).--verification plancreates and injects the plan without executing its gates.--verification offexplicitly disables test-plan creation and execution.
Left alone, every executable command in the plan comes from discover_project(),
a per-framework prober that reads the filesystem: a package.json yields its
test/build/lint scripts, a pyproject.toml mentioning pytest yields exactly
one python3 -m pytest <workdir>. A spec whose phases name specific verification
commands is therefore gated on something it never asked for, and the gap is silent —
the gate passes, the evidence looks clean, and the commands the spec named were
never run.
--verification-commands FILE closes that gap. The document lists the commands the
gates MUST execute; each joins its phase's suite (after the discovered ones) and the
release suite:
{
"schema_version": "1.0.0",
"commands": [
{
"id": "p0-canonical-embedded",
"phase": 3,
"executable": "uv",
"args": ["run", "python", "scripts/verify_canonical_run_e2e.py", "--mode", "embedded"],
"cwd": "/abs/path/to/project",
"timeoutSec": 900,
"env": {"ADLC_SPECIALIST_SOURCE": "fixture"}
}
]
}phaseis the spec's phase number. A phase the spec does not define aborts the preflight — commands no gate could reach would otherwise sit unexecuted while the run reported success.executablemay be a bare name; it is resolved onPATHat plan time and stored absolute, so the plan records what will actually run. Unresolvable aborts.envis an overlay on the run environment, not a replacement, and is rendered into the command line the proposer andTEST_PLAN.mdsee.- Commands still execute with
shell=Falseby argv. There is no shell string form. - The document's hash is bound into the plan hash, so editing it after planning invalidates the plan rather than quietly changing what a gate checks.
- Declared commands satisfy
--verification requiredon their own: a project with nothing discoverable is no longer refused when it has declared commands to run. - A top-level
"discovery": "none"turns the discovered commands off: every gate runs the declared list and nothing else (the fingerprint and frameworks are still recorded, and the plan says discovery was disabled). Every phase then needs a declared command, which the preflight enforces.specstride reversesets it, so a run over a repo it only reads never executes that repo's tests. Beside it, an optional top-level"phaseTimeouts": {"N": SECONDS}sets per-phase pass ceilings.
Gate evidence records the revision it ran against — sourceRevision.revision from
git rev-parse HEAD plus workingTreeDirty — because an exit code proves nothing
if you cannot say which tree produced it. When the workdir is not a repository the
field carries available: false and a reason rather than being omitted.
A 90-minute live suite does not fit inside a proposer pass that must cite it, and a
cumulative gate re-runs it module by module every later phase (across feature
002-extproc-data-path, one live suite was launched 159 times). Five optional
fields on a declared entry move that measurement out of the pass and out of the
duplicated gate:
| Field | Default | Meaning |
|---|---|---|
stage |
"gate" |
"prestage" runs the command ONCE per attempt, before the proposer pass, and lets that phase's gate reuse the passing result. "both" pre-stages it and still executes it at the gate. ("pre" is accepted as a spelling of "prestage".) |
reportPath |
— | workdir-relative artifact the command produces; named in the block the proposer reads, so the pass writes evidence from a file instead of re-running the measurement |
reusePolicy |
"per-attempt" |
how far a passing pre-stage may travel: per-attempt (this attempt's gate), per-phase (any attempt of that phase), per-run (any later gate too) |
detached |
false |
launch it as a job Specstride owns (its own session, its own log) instead of blocking; its timeoutSec becomes a polled deadline — an expired deadline is reported with the job left running, never signalled |
cumulative |
true |
false gates the command at its own phase and at release only, instead of at every later phase gate |
The pre-stage report reaches the pass on the verification slice the proposer prompt
already carries; verification_plan.py prestage-report --plan … --phase N prints
the same block. An entry with no stage behaves exactly as it does today — same
command id, same gate, same evidence.
Reuse fails closed, and says where the result came from. A gate accepting a
result it did not observe is the one change here that can weaken a verdict, so it
is allowed only when every one of these holds: the plan hash matches, the command
id matches, git rev-parse HEAD is the same revision the pre-stage ran against,
and the working tree is clean at both ends. Anything else — a moved revision, a
dirty tree, an unavailable revision, an unknown attempt, a pre-stage that failed or
is still running — re-runs the command. Every adopted record carries reusedFrom
(the pre-stage evidence path, its phase/attempt, the plan hash and the revision),
so the gate document never claims an execution it did not perform.
Because a dirty tree refuses reuse, keep the run's own artifacts out of git status
— .specstride/ and testautomation/ in .gitignore — or the tree is dirty from the
first pass and the gate (correctly) re-runs everything. cumulative: false is a real
weakening of the cumulative-regression property, so it is opt-in per command and is
reported in the plan's assumptions block, never inferred.
By default, each feature gets isolated artifacts at
<workdir>/testautomation/<feature>/TEST_PLAN.md and
<workdir>/testautomation/<feature>/generated/. Operator overrides must be absolute,
resolve inside the workdir, and not target a final-path symlink. A maximum-observability
run against Lisa is:
SPECSTRIDE_AGENT_STREAM=true SPECSTRIDE_LIVE_DETAIL=full \
$SPECSTRIDE_HOME/specstride run \
--workdir ~/projects/app \
--specs ~/projects/app/specs/specification-bundle-v2/tasks.md \
--spec-format speckit-tasks \
--feature specification-bundle-v2 \
--verification required \
--test-plan ~/projects/app/testautomation/specification-bundle-v2/TEST_PLAN.md \
--generate-tests ~/projects/app/testautomation/specification-bundle-v2/generated \
--live \
--debug \
--telemetry \
--loki-url http://127.0.0.1:13011 \
--otel \
--otel-url http://127.0.0.1:13018Planning can also be run independently, before any loop:
/usr/bin/python3 \
$SPECSTRIDE_HOME/lib/verification_plan.py create \
--workdir ~/projects/app \
--specs ~/projects/app/specs/specification-bundle-v2/tasks.md \
--format speckit-tasks \
--output ~/projects/app/testautomation/specification-bundle-v2/TEST_PLAN.md \
--json-output ~/projects/app/.specstride/verification/verification-plan.json \
--generate-tests ~/projects/app/testautomation/specification-bundle-v2/generated \
--requiredThe Bash entry points (orchestrator.sh, proposer.sh, specstride) sit at the top
level; all Python components live under lib/ (lib/critic.py, lib/learn.py,
lib/present.py, and the Loki and OTLP shippers).
Typing ~/specstride/orchestrator.sh … every run gets old fast. The specstride
script is already the single front door for everything — it owns the routing
itself: specstride run … (or a leading -w/-s/--flag) starts the loop by
execing orchestrator.sh, while specstride status, specstride watch, specstride stop,
… run the inspection CLI. So all your shell rc needs is a thin pointer at the
script — no dispatch logic to copy, nothing to keep in sync.
Add this to ~/.bashrc (or ~/.zshrc):
# ── Specstride ─────────────────────────────────────────────────────────────
export SPECSTRIDE_HOME="$HOME/specstride" # wherever you cloned it — set once
export SPECSTRIDE_LIVE_DETAIL=full # richest live view — narrates assistant text + every tool call
# `specstride` owns its own run-vs-inspect routing, so this is just a pointer.
specstride() { "$SPECSTRIDE_HOME/specstride" "$@"; }
# ───────────────────────────────────────────────────────────────────────Reload once (source ~/.bashrc) and the one command drives every example below,
from any directory:
specstride run -w ~/projects/foo -s ~/projects/foo/ROADMAP.md # START a loop
specstride -w ~/projects/foo -s ~/projects/foo/ROADMAP.md # …same thing, leading flag
specstride status -w ~/projects/foo # inspect it
specstride watch -w ~/projects/foo # live status card
specstride stop -w ~/projects/foo # clean haltThe routing lives in the script (see the "single front door" block at the top of
specstride): run/start or a leading -w/-s/--flag go straight to the
orchestrator; the reserved inspection verbs and -h/--help stay in the CLI. Use
the explicit specstride run … form whenever you want to be unambiguous (or in
scripts).
Why a function and not a
symlink/PATH shim? The scripts locate their ownlib/andspecstride-lib.shviadirname "${BASH_SOURCE[0]}", which does not dereference symlinks — aln -s … /usr/local/bin/specstridewould resolve its home to/usr/local/binand fail to findspecstride-lib.sh. The function calls the real absolute path under$SPECSTRIDE_HOME, soSCRIPT_DIRstays correct. (Prefer PATH?export PATH="$SPECSTRIDE_HOME:$PATH"also works, and because the script owns its own run-vs-inspect routing, barespecstride run …starts a loop that way too — no function needed. The function is just the tidiest way to pin$SPECSTRIDE_HOME.)
The rest of this README uses the unified specstride command — specstride run … (or
a leading specstride -w …) to start, specstride <verb> … to inspect. Because the
script owns the routing, you don't even need the function: call
"$SPECSTRIDE_HOME"/specstride … directly and both specstride run … and the inspection
verbs work the same way.
orchestrator.sh (derives the current phase N from disk; reads the feature's tasks.md)
│
├─(1) PROPOSER — run a headless coding-agent loop for phase N until it writes
│ .specstride/gates/GATE<N>-EVIDENCE.md (written atomically), then the loop exits.
│ On the attempt right after a NEW diagnostician hint this is the
│ ACCELERATOR instead: the same proposer.sh, --role accelerator, with a
│ prompt narrowed to the unmet criteria, the feedback, the hint and the
│ evidence to splice. Once per hint, never twice in a row (see below).
│
├─(2) CRITIC — lib/critic.py reads phase N's acceptance criteria + the evidence,
│ does a read-only grounding pass over the files the evidence cites (byte
│ budget scaled to the critic backend's real context window), and asks an
│ LLM for a strict verdict:
│ APPROVED → writes an empty .specstride/gates/GATE<N>-APPROVED marker
│ REJECTED → writes .specstride/gates/GATE<N>-FEEDBACK.md (the specific gaps)
│
├─(3a) APPROVED → git-checkpoint the workdir, N := N+1, back to (1).
└─(3b) REJECTED → compute the phase's UNMET-CRITERIA SIGNATURE (the T### IDs the
feedback names, or a hash of its prose), then:
NEW signature → run the DIAGNOSTICIAN (lib/critic.py --diagnose: the
same critic backend, the FULL untruncated files, no
grounding budget) → .specstride/gates/GATE<N>-HINT.md.
The next attempt is the ACCELERATOR.
SAME signature → archive the rejected evidence and re-run the wide
PROPOSER with the feedback + hint. If the previous
attempt was the accelerator, its GATE<N>-ACCELERATION.md
(what it changed) is in the prompt so it is not redone.
Bounded by MAX_REJECTS (accelerator attempts count); on exceed, halt and
leave everything on disk for a human.
The same loop as a UML sequence — the five roles (orchestrator, proposer,
accelerator, critic, diagnostician) each on their own lifeline, the
approve/reject branch, and the stuck-loop path (new signature → diagnostician →
accelerator), all mediated by the .specstride/gates/ files rather than direct calls:
sequenceDiagram
autonumber
actor Human
participant O as orchestrator.sh<br/>(orchestrator)
participant P as proposer.sh<br/>(proposer · coding-agent CLI)
participant A as proposer.sh --role accelerator<br/>(accelerator · same tools, narrowed prompt)
participant C as lib/critic.py<br/>(critic · LLM gate)
participant D as lib/critic.py --diagnose<br/>(diagnostician · same backend, no budget)
participant FS as .specstride/gates/<br/>(on-disk contract)
Human->>O: run -w WORKDIR -s specs/<feature>/tasks.md
O->>FS: derive phase N from GATE* markers
Note over O: no stored counter — phase is derived
loop until all phases APPROVED (or halt)
alt attempt right after a NEW diagnostician hint<br/>(once per signature, never twice in a row)
O->>A: narrowed prompt: ONLY the unmet criteria<br/>+ critic feedback + hint (primary instruction)
activate A
A->>FS: read GATE<N>-HINT.md + the archived (rejected) evidence
A->>A: fix ONLY the unmet criteria<br/>(footprint rule: touch just the files they cite)
A->>FS: write GATE<N>-EVIDENCE.md<br/>(previous evidence spliced, atomic)
A-->>O: pass exits (test -f passes)
deactivate A
O->>FS: write GATE<N>-ACCELERATION.md<br/>(files this pass changed)
else ordinary attempt
O->>P: full phase prompt<br/>(+ feedback, hint, acceleration note if present)
activate P
loop until evidence exists
P->>P: read PROGRESS.md, do the work
P->>FS: write GATE<N>-EVIDENCE.md (atomic)
end
P-->>O: loop exits (test -f passes)
deactivate P
end
O->>C: judge phase N (criteria + evidence)
activate C
C->>FS: read-only grounding pass over cited files<br/>(byte budget scaled to the backend's context window)
C->>C: LLM verdict, nonce-bound
alt APPROVED
C->>FS: write GATE<N>-APPROVED (empty marker)
C-->>O: VERDICT nonce: APPROVED
O->>O: git checkpoint · N := N+1
else REJECTED (attempt < MAX_REJECTS)
C->>FS: write GATE<N>-FEEDBACK.md (the gaps)
C-->>O: VERDICT nonce: REJECTED
O->>O: unmet-criteria signature<br/>(task IDs the feedback names, or a prose hash)
alt signature is NEW for this phase
O->>D: diagnose phase N<br/>(full rejection history, same critic backend)
activate D
D->>FS: read the FULL, untruncated cited files<br/>(no grounding budget)
D->>D: classify the stall:<br/>CASE: GROUNDING (restage the proof)<br/>or CASE: REAL-GAP (the concrete fix)
D->>FS: write GATE<N>-HINT.md
D-->>O: hint written (advisory — never approves or rejects)
deactivate D
Note over O,A: next attempt = ACCELERATOR
else same signature as last time
Note over O,P: next attempt = wide PROPOSER<br/>(reads feedback + hint + acceleration note)
end
O->>FS: archive stale evidence<br/>(+ feedback, hint, acceleration note)
else MAX_REJECTS exceeded (accelerator attempts count too)
C-->>O: still REJECTED
O->>Human: halt (exit 2) — arbitrate
end
deactivate C
end
O->>Human: all phases approved (exit 0)
There is no file-watcher. Detection is deterministic: the proposer loop's
gate is a plain test -f .specstride/gates/GATE<N>-EVIDENCE.md, and because that loop has already
exited when control returns, the orchestrator hands the critic the exact path —
no race, no half-written file.
A large phase can cite more files than the critic's grounding snapshot can fit in
one budget (GROUNDING_TOTAL_CAP), so the same file gets degraded to a head/tail
excerpt — or elided — on every attempt. When that's the actual cause, the
criterion never converges: the judge isn't wrong about what it can see, it
just can't see enough, and a plain retry burns a full proposer+critic pass to
learn nothing new.
The orchestrator tracks each phase's unmet-criteria signature (the T### IDs a
rejection names). The FIRST time a new signature appears, it runs lib/critic.py --diagnose: one extra pass, same critic backend (SPECSTRIDE_CRITIC), given the
full rejection history and the FULL, untruncated content of the cited files —
no grounding budget. It classifies the stall as CASE: GROUNDING (the code is
fine, the critic just couldn't see it — and says what to restage) or
CASE: REAL-GAP (a genuine gap — and what to fix), and writes
.specstride/gates/GATE<N>-HINT.md, which the next proposer prompt reads alongside
the critic's own feedback.
It never re-fires on an unchanged signature (no point paying for the same
answer twice), never blocks or replaces the normal retry, and never approves or
rejects anything itself — it's advisory. Disable with SPECSTRIDE_DIAGNOSTICIAN=false.
The diagnostician knows the fix but cannot touch the tree (it is a tool-free critic call). Left alone, the next proposer pass re-reads the FULL phase prompt (every task, contract and verification obligation) to re-derive a change the hint already spelled out — on a real 12-task phase that was a 16-minute pass to apply a two-file diff.
The accelerator is that retry pass, narrowed. It is proposer.sh with the
same tools and (by default) the same backend, run with --role accelerator and a
prompt that carries only:
- the unmet criteria (by the diagnostician's
T###signature; the whole phase for prose-only specs) and only their verification obligations, - the critic feedback, and the hint as the primary instruction,
- the archived evidence, with the order to splice only the rejected criteria's sections and leave everything else byte-for-byte,
- a footprint rule: touch only files the unmet criteria and the hint cite. Confirmed criteria are pinned (W9) to a content hash of their backing files; a wide edit drops those pins and re-opens criteria that already passed.
Sequencing is strict — one writer at a time, never a parallel agent:
reject → diagnostician (new signature → GATE<N>-HINT.md)
→ attempt N+1 = ACCELERATOR (once per signature) → verification → critic
APPROVED → done
REJECTED, same signature → attempt N+2 = wide PROPOSER, whose prompt
carries GATE<N>-ACCELERATION.md (files the accelerator
changed + its own notes) so it does not redo that work
REJECTED, new signature → diagnostician → accelerator again
An accelerator attempt counts toward MAX_REJECTS like any other, and it never
runs twice in a row. Its prompt is written to accelerator-prompt.phase<N>.txt
next to the proposer's; its invocation artifacts land under .../accelerator/.
Disable with SPECSTRIDE_ACCELERATOR=false; point it at a different model with
SPECSTRIDE_ACCELERATOR_BACKEND.
GROUNDING_TOTAL_CAP (and the diagnostician's own, larger budget) is not one
flat number for every provider — each real backend has a different context
window, and using one number for all of them either wastes headroom on a
bigger window (this is the SAME starvation failure the diagnostician exists
for, just from under-sizing instead of an oversized phase) or risks
overflowing a smaller one. Specstride resolves the actual window per call — the
Claude/OpenAI providers call those vendors' APIs directly, so they're keyed
against the vendors' own published windows (Claude Opus 4.8: 1,000,000; GPT-5:
400,000), not a local fleet's internal operational settings for an unrelated
routing path. DSH/bebop genuinely do route through this fleet's local
infrastructure, so THEIR numbers come from there instead: a locally-served
Qwen3.8 at whatever this fleet measured it can actually load (229,376 on the
reference 24 GB card), GLM-5.3 at its declared 128,000. The byte budget scales
from whichever of these actually applies, instead of guessing or hardcoding
one provider's number for all of them.
Override with SPECSTRIDE_CRITIC_CONTEXT_TOKENS for any backend the built-in
table doesn't know (in particular Prime, whose backing model isn't visible to
critic.py at all) or to correct a host-specific deployment.
Clone, set your key, alias, run:
cp .env.example .env # then edit: set ANTHROPIC_API_KEY
# one-time setup (see "Install it permanently" above): add the thin specstride()
# pointer to ~/.bashrc, pointing SPECSTRIDE_HOME at this clone, then reload:
source ~/.bashrc
mkdir -p /tmp/specstride-demo/specs/001-greeting
cp examples/speckit-tasks.example.md /tmp/specstride-demo/specs/001-greeting/tasks.md
specstride run -w /tmp/specstride-demo --verification off # discovers specs/001-greeting/tasks.md(Not set up yet? The one-off equivalent calls the script directly:
"$SPECSTRIDE_HOME"/specstride run -w /tmp/specstride-demo --verification off — or ./specstride run -w /tmp/specstride-demo --verification off from inside the clone. The flag is explained under
Pre-loop test automation.)
Pick the backends with SPECSTRIDE_PROPOSER and SPECSTRIDE_CRITIC (in .env or the
environment): claude, codex, dsh, or bebop (see Configuration). The two roles are
independent, so a cheap model can propose while a stronger one judges. The bundled
examples/speckit-tasks.example.md is a small, verifiable Spec Kit task list, so you can
watch the whole loop end to end. In a real project, generate the feature with Spec Kit
(/speckit.specify, /speckit.plan, /speckit.tasks) and point Specstride at it.
A backgrounded loop is not a black box. Every meaningful step emits one structured event, and a presenter renders it in real time — in full color, with zero containers. Two views over the same event stream:
-
Inline timeline (
--live, auto-on at a TTY): when you launch the orchestrator in a terminal, it streams a clean, colored, scrolling timeline right there — like watching a coding agent work — while the noisy raw proposer/critic output goes torun.logonly. The proposer's agent stream is narrated as it happens: each tool call gets its own color and glyph (Read◎, Write✚, Edit✎, Bash❯, Grep/Glob❍, …), timestamps recede to gray so the action carries the color, and every end-of-pass line shows the cost / tokens / duration / turns each tinted. Whenever the agent goes quiet for more than a couple of seconds, an animated heartbeat (spinner + pulse bar) keeps the current activity and running totals visibly moving. No second terminal, nospecstride watch. Color auto-strips when stdout isn't a TTY, or force it off with--no-color; force the whole view off with--no-live(restores the raw tee'd output).14:02:41 ◎ Read proposer.sh 14:02:43 ❯ Bash npm test 14:02:44 ⠹▄ proposer working · 3s · run $0.04 · 12.1k tok out · 1 pass · 41s 14:03:02 ⏺ pass done · $0.12 · 5.1k tok · 18s · 12 turns 14:03:03 ✓ evidence → GATE2-EVIDENCE.md (1 iter) 14:03:19 ✗ REJECTED phase 2 (attempt 1) — criterion 3: no passing testVerbosity is
SPECSTRIDE_LIVE_DETAIL(milestones | tools | full; defaulttools). Set it in.envor inline per run.fulladds each assistant thinking/narration line (💬) on top of the tool calls — the most detailed view:SPECSTRIDE_LIVE_DETAIL=full specstride run -w ~/projects/foo --live(
milestonesis the sparsest — only coarse loop milestones, no per-tool lines.) -
Live status card:
specstride watch— a compact header (phase progress + current activity + heartbeat) over a scrolling recent-activity feed, so the latest message is always visible in place without scrolling. Attach to a backgrounded run. HonorsSPECSTRIDE_LIVE_DETAILthe same way:SPECSTRIDE_LIVE_DETAIL=full specstride watch -w ~/projects/foo
Everything is read-only except stop, resume, and learn --apply|--revert|--off
— those are the only subcommands that mutate anything (learn's writes are
confined to its own learning/applied.json decision log and a knob_adjusted
event; see Learning below).
All inspection subcommands take --feature SLUG and default to the last run's
feature (from .specstride/last-run.conf); status --all spans every feature.
| Command | Shows / does |
|---|---|
specstride status [-w DIR] [-s SPEC] [--feature S] [--all] |
one-screen state: a run-state headline (RUNNING / STOPPED / HALTED / DONE) + current phase + a ✓/✗ table of which contract files exist. --all lists every feature with its approved/total phase counts |
specstride phases [-w DIR] [-s SPEC] [--feature S] |
phases parsed from the spec + each one's state (also lints the spec) |
specstride tail [-w DIR] [--feature S] |
tail -f the orchestrator run.log (the raw log) |
specstride events [-w DIR] [--feature S] [-f|--follow] [--json] |
the raw event stream ("RPC view"): every milestone and every agent tool call / message as HH:MM:SS event key=value… lines. --follow streams; --json emits the raw JSONL |
specstride verdicts [-w DIR] [--feature S] [N] |
the critic's full reply(ies): prompt + response + parse decision |
specstride feedback <N> [-w DIR] [--feature S] |
GATE<N>-FEEDBACK.md |
specstride watch [-w DIR] |
the live status card (with heartbeat + run totals) |
specstride stop [-w DIR] [--now] |
(mutates) request a clean stop — writes stop.flag; the run finishes its current pass and exits 6. --now also kill-trees the in-flight proposer pass so it stops within seconds. Stops the single running run regardless of feature |
specstride resume [-w DIR] [--feature S] [overrides…] |
(mutates) relaunch the orchestrator from the saved config of the last run (.specstride/last-run.conf, or a feature's own with --feature); refuses if a run is already active. Extra args override the saved flags (last-wins) |
specstride learn [-w DIR] [--feature S] [--show|--apply|--revert <run-id>|--off|--summarize|--evaluate] |
the self-tuning loop over this feature's telemetry — see Learning. --show (default), --summarize (the metric JSON over every run) and --evaluate (what evaluation would conclude now; always a dry run) are read-only; --apply/--revert/--off mutate learning/applied.json |
Specstride can suggest — and, opt-in, apply — per-phase knob values derived from what
that phase has actually measured, instead of one global setting sized for the
worst phase in the project. This is deliberately narrow: suggest is the
default, nothing is ever applied silently, and the set of knobs it may ever touch
is a locked allowlist that can never include anything the critic reads. Full
design: roadmap/research/self-improvement-loops/ (the loop design, §5).
In practice:
# 1. Let runs record per-phase observations (measurement only, changes nothing)
SPECSTRIDE_LEARNING=suggest specstride run ...
# 2. After a few runs, see what the loop would change and the evidence behind it
specstride learn --show
# 3. Apply one decision you agree with (append-only, with provenance)
specstride learn --apply --knob proposer_timeout --phase 8 --default 1800
# 4. Launch runs that read applied values
SPECSTRIDE_LEARNING=apply specstride run ...
# Undo one decision, or all of them
specstride learn --revert <run-id>
specstride learn --off
# Read-only: the metric JSON, and what evaluation would conclude now
specstride learn --summarize
specstride learn --evaluate-
Suggest by default.
specstride learn --show(or barespecstride learn) prints a suggested value per phase plus the sample count behind it — it writes nothing. Every knob needs at least 3 samples of the phase evidence it is derived from before a value is shown at all; a phase killed for futility (repeat_stall,progress_stall) never contributes a sample to any of the three, because that failure's duration means nothing about how long the work (or the wait) actually takes. Both allowlisted knobs have a suggestion engine:proposer_timeout— from the phase's measuredwork_sec(wall time minus declared wait/sleep), clamped to[900 s, 2×default]and to no more than a ±50% step from whatever value is currently in effect.yield_poll_interval— from the phase's measured yield job durations (yield_resume), targeting roughly a tenth of the job's own length so a missed tick wastes a small fraction of it rather than a fixed cost; clamped to[10 s, 300 s]and the same ±50%-per-step window, stepped from the poll interval's own default (SPECSTRIDE_YIELD_POLL, else 30 s).
specstride learn --show/--applypass the default of the knob you ask about when you give no--default:SPECSTRIDE_PROPOSER_TIMEOUT(else 1800) forproposer_timeout,SPECSTRIDE_YIELD_POLL(else 30) foryield_poll_interval. -
What can never move. The adjustable-knob allowlist is exactly
proposer_timeoutandyield_poll_interval— nothing else, ever. (The design's third knob,inject_yield_hint, was removed: the yield contract is already appended to every proposer and accelerator prompt, so it had nothing to switch.) Grounding caps, the critic's backend/timeout,--max-rejects, anything inverification-commands.json, and every breaker setting (SPECSTRIDE_PROPOSER_MAX_ERRORS/MAX_NOPROGRESS/MAX_CAPS,REPEAT_LIMIT) are permanently out of scope: a breaker must never be able to relax itself, and nothing that changes what a verdict means may be tuned.lib/test_learn.pyasserts this set literally, so adding a name to it means deliberately editing a test that explains why that must not happen. A second test locks wherelearn.pymay be invoked at all: exactly the tworesolvecalls and theobservehook inorchestrator.sh, and thespecstride learndispatcher — and never fromlib/critic.pyorlib/verification_plan.py. -
Bounded and reversible. Every numeric suggestion is clamped to its hard bound (above) and to no more than a ±50% step from whatever value is currently in effect — a self-tuner cannot run away in one step even if the telemetry that produced the suggestion was noisy.
-
Keyed on the phase's shape. Every sample and every decision is filed under the phase's shape: a digest of its number, title and criteria text (
python3 lib/specstride_spec.py shape N --specs <spec>), recorded on eachphase_start. Ticking a checkbox or reflowing whitespace does not change it; editing a criterion or the title does. After an edit, samples from the old phase stop counting toward the 3-sample floor and a decision learned for it stops applying — the run resolves the default again. A decision or a run recorded before shapes existed matches no shape: it is never silently applied, andresolve/--show/--applyprint a one-line notice saying so (the run's notice goes to itsrun.log).specstride learncomputes the shapes from the feature's spec (-s, else the savedSPECS=). -
Applying, and undoing it.
specstride learn --apply --knob <knob> --phase N [--default S] [--force]appends one decision to.specstride/features/<slug>/learning/applied.json(a separate, append-only file from the plain observations below — decisions and observations are never conflated) with full provenance: the phase shape, the run ids the samples came from, the sample count, the previous value, and a timestamp (specstride.learn.applied/2;/1entries are still read); it also emits aknob_adjustedevent.specstride learn --revert <run-id>undoes exactly that one decision, restoring the value from just before it (refused if a later decision has already superseded it — reverting a stale one would silently clobber the newer one).specstride learn --offreverts every currently-applied knob for the feature at once. None of this takes effect at run time unless the run itself is launched withSPECSTRIDE_LEARNING=applyin its environment — unset,off, or any other value is a total no-op: the applied log isn't even opened, and behaviour is byte-identical to a project that has never usedlearnat all. -
What a live run reads. Under
SPECSTRIDE_LEARNING=applya run reads both knobs back, once per phase:proposer_timeoutas thelearnedsource ofresolve_proposer_timeout, andyield_poll_intervalas the poll interval the phase's proposer passes run with (resolve_yield_poll). An operator's own setting still wins over either:--proposer-timeout-phase/SPECSTRIDE_PROPOSER_TIMEOUT_PHASE_<N>for the ceiling, and a setSPECSTRIDE_YIELD_POLL(set at all, even to 30) for the poll. Both values and their sources are on everyproposer_capevent. -
The integration point. The one place a learned value can reach a live run is
lib/learn.py resolve, a small shell-callable entry point — not a change toorchestrator.shitself:python3 lib/learn.py resolve --knob proposer_timeout --phase <N> --default <S>prints one integer to stdout: the applied value for that phase if
SPECSTRIDE_LEARNING=applyand one has been applied, else<S>unchanged.
Separately from the applied-decisions log above, lib/learn.py observe writes
the §5.4 per-phase observation document — exactly that phase's entry from
summarize's metric set, with a schema tag and a generation timestamp — to
.specstride/features/<slug>/learning/phase-<N>.json. It is the measurement half of
§5.4's storage table; it never reads or writes applied.json, and nothing
reads it automatically today — it exists so a phase's own cross-run history is
sitting on disk for a human, specstride learn --show, or a future run of the same
phase to read, instead of being re-derived from every run's raw events.jsonl
each time. It is idempotent: the same input always overwrites --out with the
same "observation" content, so it is safe to call unconditionally.
python3 lib/learn.py observe --events <events.jsonl|run-dir|runs-dir> --phase <N> \
[--out <feature-dir>/learning/phase-<N>.json]
The exact line an orchestrator's phase_done hook would call (one line, using
the orchestrator's own $LIB_DIR, $FEATURE_DIR, and current-phase $n):
python3 "$LIB_DIR/learn.py" observe --events "$FEATURE_DIR/runs" --phase "$n" \
--out "$FEATURE_DIR/learning/phase-$n.json"
--events "$FEATURE_DIR/runs" (a directory of run directories) is deliberate,
not the current run's events.jsonl alone: a phase's observation should reflect
every run that has ever touched it, the same cross-run view summarize's
phases map already gives attempts_to_approval and runs_seen. The hook IS wired at
phase_done (see the next section).
- Observations are written where the phase closes. With
SPECSTRIDE_LEARNINGset to anything butoff, an approved phase writes one observation of itself —python3 lib/learn.py observeover that run'sevents.jsonlinto.specstride/features/<slug>/learning/phase-<N>.json— and emits onelearning_observedevent naming it. Observations and decisions stay separate files on purpose (design §5.4). The hook is best-effort in every direction: it is skipped when the layer is off, skipped silently when the installedlearn.pyhas noobservesubcommand, and a failure is logged to the run log and dropped. Measuring an approved phase may never un-approve it.
An applied decision is checked, not trusted. specstride learn --apply records a
baseline beside the decision: the per-pass cost and wall-clock of exactly the samples
that produced it, the phase shape, the backend label of the most recent source run
(run_start.backend, e.g. prime:sol; no event records a model version, so the label is
the reset key), and the counts the guardrails need. At every phase_done, beside
observe and under the same discipline (layer on, evaluate --help probe, a failure is
logged and never costs the phase), learn.py evaluate reads every run of the feature and
labels each active decision for the phase:
- The sample unit is the billed, non-futility pass, clustered by phase episode (one
run's window on the phase). A futility-killed or unbilled pass counts in neither arm.
The applied arm is the runs whose
proposer_capsays they ran under the decision. - Below 6 samples per arm the label is
insufficientand nothing else happens. - Otherwise, for log cost and log wall-clock per pass:
r= the difference of means (applied − baseline), and the minimum detectable effect isMDE = 2.80 · s · sqrt(1/n_a + 1/n_b) · sqrt(1 + (m̄ − 1)·0.5), withsthe pooled sd of the logs floored at 0.50 andm̄the passes per episode.r ≤ −MDEishelped,r ≥ +MDEisregressed, anything between isneutral;regressedon either primary isregressed. There is no significance test: at Specstride's run counts none is attainable. The effect is always printed with its MDE beside it —neutralmeans "nothing this large could be seen", and at 6 passes per arm that is roughly a factor of two to three in cost. - A changed phase shape or backend label resets the comparison (
reset: shape|backend, labelinsufficient); a decision applied before baselines existed isinsufficient. - For a shorter
proposer_timeoutthe entry also records a wall-clock-only counterfactual from the recorded pass durations (never cost: a truncated pass's cost is unobserved); a longer one is markedcensored, since a killed pass does not show how long it would have run.
Each result that differs from the last one for the same decision appends an evaluate
entry to applied.json (the apply entry is never mutated, and replay ignores evaluate
entries) and emits knob_evaluated. learn.py evaluate --dry-run prints without writing;
--report prints the last recorded evaluation of every active decision. Design:
roadmap/research/self-improvement-loops/05-evaluate-design.md.
Guardrails veto, even when a primary improved. Each is computed per arm and is
ok, breach or unknown — below its minimum n it is unknown, never ok:
| Guardrail | From | Breach |
|---|---|---|
grounding_gap rate per verdict |
grounding_gap |
exact one-sided binomial tail p < 0.01 against the baseline rate, ≥ 10 applied verdicts |
| MALFORMED verdict rate | verdict.result |
same, ≥ 10 applied verdicts |
| declared verification | verification_failed |
any failure in the applied arm where the baseline had none, at any n |
| diagnostician GROUNDING share | diagnostician_done.case |
binomial, ≥ 10 applied cases |
| critic input size | the PROMPT section of verdicts/phase<N>.attempt<A>.<ts>.txt (no event carries it) |
exact rank test p < 0.01, ≥ 6 prompts per arm |
| human arbitration — a proxy | run_stop.reason ∈ max_rejects, gate_oscillation, critic_config, proposer_no_progress (no event records arbitration) |
binomial, ≥ 10 applied episodes |
| first-attempt approval rate | verdict on attempt 1 |
an alarm in both directions (either tail p < 0.01), ≥ 10 verdicts; never a reward and never part of a label |
False-MISSING rate is not a guardrail: no event carries it, and parsing the critic's prose
into a score would make its output an optimization input. It stays an offline release check
on critic.py.
Auto-revert and quarantine. On a breach evaluate re-reads applied.json and reverts
the decision through the normal revert path — the only automatic write to a decision the
loop makes, since it only moves a knob back toward its default — and emits
knob_auto_reverted naming the guardrail. If a manual --apply overtook the decision
meanwhile, the revert is skipped (knob_evaluated action=skipped_superseded) and never
retried; after --off it is a silent no-op. The auto-revert entry quarantines the reverted
value for its (knob, phase, shape, backend): --show still prints the suggestion, marked
QUARANTINED, and --apply refuses a value within ±10 % of it (exit 4, naming the auto-revert)
until 6 new billed non-futility samples have accrued since; --apply --force overrides and is
recorded (force: true). A shape or backend change clears it.
The arm and the tamper rule. learn.py resolve prints <value>\t<arm>, and every
proposer_cap carries arm and yield_poll_arm: applied while a decision is in effect for
that phase and shape — even one whose value equals the default — else baseline. Evaluation
compares before with after; it does not alternate arms (at these run counts alternation never
reaches its own floor, and a resumed run would split one episode across both arms), which is
why only a revert is ever automatic. Because the proposer runs unsandboxed in the project and
appends to the same events.jsonl, the orchestrator — the parent process — brackets every
proposer launch while the layer is on: it records the size and SHA-256 prefix of
events.jsonl, learning/applied.json and learning/phase-<N>.json before the pass, and if
afterwards any is shorter or its old bytes differ, it emits events_tampered and that run is
excluded from every evaluation (appends are normal and never flagged; a run whose arm was
recorded without the bracket is excluded, not trusted). With the layer off no bracket is taken.
The rule does not catch forged appended events, writes outside a pass (detached long jobs),
edits to verdict transcripts, verification records or evidence files, or spec edits (which the
shape detects); see the design note.
Where it is shown. specstride learn --show prints each active decision's latest
evaluation after the suggestions, and every run prints the same lines on exit — on
run_end and on each run_stop path — when the layer is on.
The Spec Kit feature (its tasks.md, generated from spec.md and plan.md) is the one
input you write; it can live anywhere (-s, or discovered — see Spec resolution).
Everything else Specstride generates lives under .specstride/, namespaced per feature,
so the workdir root stays clean — only your real project artifacts sit there.
| File | Written by | Meaning |
|---|---|---|
specs/<feature>/tasks.md |
you (via Spec Kit) | Ordered phases and their tasks (the input). A legacy SPECS.md is still read. |
.specstride/features/<slug>/PROGRESS.md |
proposer | Durable state; read first each iteration. |
.specstride/features/<slug>/gates/GATE<N>-EVIDENCE.md |
proposer | Evidence phase N's criteria are met. Written atomically. |
.specstride/features/<slug>/gates/GATE<N>-APPROVED |
critic | Empty marker; unblocks phase N+1. |
.specstride/features/<slug>/gates/GATE<N>-FEEDBACK.md |
critic | Present after a REJECT; the gaps to fix. |
.specstride/ |
orchestrator | State dir (see below). The current phase is derived from the GATE* markers, never stored. |
Feature-scoped state. Durable state hangs off .specstride/features/<slug>/ so
multiple Spec Kit features can build into one repo without their gates,
evidence, and verdicts colliding. <slug> is the feature-dir basename when the
spec lives inside a .specify project (001-reverse-engineering-analysis), and
default otherwise — which is also the back-compat identity of every pre-v2
.specstride/gates/ on disk (a native workdir keeps its state, transparently migrated
once on the next run).
| Path | Scope | Holds |
|---|---|---|
.specstride/features/<slug>/gates/ (+ gates/proofs/) |
per-feature | all the phase-control files above — where to look for what the loop produced |
.specstride/features/<slug>/runs/<run-id>/{run.log,events.jsonl} |
per-feature | each run isolated |
.specstride/features/<slug>/learning/phase-<N>.json |
per-feature | the phase's latest observation (Learning) |
.specstride/features/<slug>/learning/applied.json |
per-feature | the append-only decision log: apply entries (value, shape, provenance, baseline), revert entries, and evaluate entries (label, r/MDE/n per primary). Design: roadmap/research/self-improvement-loops/ (the loop design §5, and 05-evaluate-design.md) |
.specstride/features/<slug>/{verdicts,attempts,debug}/ |
per-feature | critic transcripts, archived rejected attempts (attempts/phase<N>/attempt<M>/), debug dumps |
.specstride/features/<slug>/debug/invocations/<run-id>/<role>/phase-<N>/attempt-<M>/iter-<I>/<invocation-id>/ |
per-feature | one reconstructable proposer/critic invocation: metadata.json (contract specstride-invocation/v1) + a terminal result.json (specstride-invocation-result/v1), and — only when raw capture is explicitly enabled — prompt.txt / provider.jsonl / events.jsonl / response.txt. Every field is routed through lib/observability_policy.py first: secrets redacted, thinking dropped, oversized payloads truncated with truncated=true. Raw content expires after 7 days; redacted metadata + terminal result are kept 30 (the summary always outlives the raw it describes) |
.specstride/features/<slug>/PROGRESS.md, last-run.conf |
per-feature | proposer notes; that feature's resume config |
.specstride/lock, .specstride/stop.flag |
workdir | one run per repo, ever — concurrency is per-workdir, not per-feature |
.specstride/run.log, .specstride/events.jsonl |
workdir | symlinks retargeted into the active feature's newest run, so specstride tail/watch/events work with no flags |
.specstride/last-run.conf |
workdir | the active-feature pointer + last launch config; what bare specstride resume replays |
.specstride/features/<slug>/proposer.pid |
per-feature | in-flight proposer pass, so specstride stop --now can kill the tree |
One run per workdir. The lock stays at the .specstride/ root: a second
specstride run in the same workdir exits E_LOCK (5) even for a different
feature, because the workdir is the repo and two features mutating one source
tree concurrently is a corruption, not a feature. Sequence features with the
operator; specstride status --all makes the sequence visible.
Every meaningful step appends one JSON object (one per line) to
.specstride/events.jsonl; specstride events and the live views render it. Lifecycle
events come from the orchestrator/proposer; the agent_* and evidence_writing
events come from the proposer's stream-json tap (lib/agent_stream.py, gated by
SPECSTRIDE_AGENT_STREAM).
| Event | Emitted by | Meaning |
|---|---|---|
run_start / run_end |
orchestrator | a run begins / all phases approved (outcome); run_start names critic_model and critic_window |
critic_over_budget |
critic | a critic prompt was still over its window after shrinking (model, window, notes); specstride status warns |
run_stop |
orchestrator | run halted early — reason (stop_flag, wall_budget, max_rejects, proposer_max_iter, proposer_consecutive_errors, proposer_cap_exhausted, proposer_yield_budget, proposer_yield_timeout, proposer_no_progress, proposer_no_evidence, critic_config) + phase |
phase_start / phase_done |
orchestrator | phase N entered / approved. phase_start carries shape, the phase-shape digest learned state is keyed on |
learning_observed |
orchestrator | a per-phase observation was written at phase_done — phase, path (learning/phase-<N>.json). Only under SPECSTRIDE_LEARNING; best-effort, and never fails the phase |
knob_evaluated |
learn.py (at phase_done) |
an active decision was labelled — knob, phase, label (helped | neutral | regressed | insufficient), cost_r/cost_mde, wall_r/wall_mde, n_applied/n_baseline, reset, action (evaluated | auto_reverted | skipped_superseded), evaluates_run_id. Emitted only when the result changed |
knob_auto_reverted |
learn.py (at phase_done) |
a guardrail breached, so the decision was reverted — knob, phase, guardrail, from, to, reverts_run_id. The only automatic write the loop makes to a decision |
events_tampered |
orchestrator | with the learning layer on, a proposer pass shortened a bracketed file or changed bytes that predated it (events.jsonl, learning/applied.json, learning/phase-<N>.json) — phase, attempt, file. The run is excluded from every evaluation |
proposer_start |
orchestrator | a proposer pass for phase N begins |
proposer_cap |
orchestrator | the pass ceiling this attempt runs under — seconds + source (override | learned | declared | global), and the yield poll interval its passes use — yield_poll + yield_poll_source (override | learned | default). Both are resolved once per phase; the event repeats them per attempt. arm / yield_poll_arm (applied | baseline) say which side of an evaluation the pass is on, and tamper_bracket=on that the tamper rule's bracket was taken. An unsourced budget is what makes budget archaeology expensive six hours in |
iter_cap |
proposer | a pass was killed at the ceiling — reason (hard_cap), elapsed, consec/max against SPECSTRIDE_PROPOSER_MAX_CAPS. A budget signal, not an error |
pass_cost_unknown |
proposer | a killed pass reports NO usage or cost (the kill severs the provider stream); this says "unmeasured", never "cheap" |
pass_yield |
proposer | a pass ended cleanly while a job it depends on runs — reason, predicate_kind, deadline_sec, job_mode, job_log, yield_index |
yield_job_start |
proposer | the job specstride now owns — pid, argv, log, sid (a session of its own: no pass kill can reach it) |
yield_wait |
proposer | sampled while waiting with no model session open — elapsed, predicate_kind |
yield_resume |
proposer | the predicate is satisfied — waited_sec, job_rc, job_duration_sec. waited_sec is what finally separates "how long the model worked" from "how long the loop was blocked" |
yield_timeout / yield_invalid |
proposer | the wait ran out (deadline | wall_budget), or the artifact was refused (schema violation, a disabled predicate, an adopt pid in the pass's own session, a watchdog-killed pass) |
prompt_block_dropped |
orchestrator | a prompt block did not fit the ASSEMBLED prompt budget (SPECSTRIDE_PROMPT_MAX_BYTES) — said out loud, never silently omitted |
iter_start / iter_done |
proposer | one headless proposer iteration |
evidence_written / evidence_present |
proposer | GATE<N>-EVIDENCE.md was just written / already existed |
attempt_archived |
orchestrator | a rejected evidence file was archived before retry |
verdict |
critic | the critic's APPROVED/REJECTED decision |
reject |
orchestrator | phase N rejected (attempt M) with feedback |
verification_infra |
orchestrator | the reverse source guard failed only on paths no tool call of the attempt named: something outside the run changed them. Not a reject; the run stops (exit 4, run_stop reason=verification_infra) |
preflight_warning |
orchestrator | a non-fatal start-up check (plugin_unignored_writes | concurrent_claude_session) with its detail |
diagnostician_trigger / diagnostician_start / diagnostician_done / diagnostician_error |
orchestrator / critic | a NEW unmet-criteria signature: one tool-free full-file pass wrote GATE<N>-HINT.md. diagnostician_done carries bytes and case — the diagnostician's declared first line, grounding (the proposer cites badly), real_gap (the work is incomplete) or unknown; a read-only metric, counted per phase by learn.py summarize and never read by the gate |
accelerator_start |
orchestrator | attempt M is an accelerator pass (narrowed prompt) acting on the hint for criteria |
acceleration_note |
orchestrator | GATE<N>-ACCELERATION.md written: what the accelerator pass changed |
git_checkpoint / gates_migrated |
orchestrator | per-phase commit / one-time relocation of pre-v2 state into features/default/ |
agent_observability |
agent tap | the capability this invocation begins with — mode (structured | degraded | raw-text) + supported_signals + reason + provider_format + role. Re-emitted if a fatal schema diagnostic degrades structured→degraded mid-stream, so a loss of fine-grained capture is explicit, never silent |
agent_init |
agent tap | once per pass: model + tool count |
harness_config |
agent tap | once per claude pass, from the child's own init record: plugins, mcp_servers, skills/slash_commands/tools counts, setting_sources, inherit_plugins, and a 16-hex fingerprint over them. Absent when the init reports no plugin fields. learn.py flags evaluations across fingerprints as confounded |
agent_tool |
agent tap | every proposer tool call: tool name + compact target |
agent_text |
agent tap | first line of each assistant message (thinking/narration) |
agent_diagnostic |
agent tap | a bounded parse warning (code, e.g. malformed_json / unsupported_schema / absent_schema) — capped, never a flood; schema-fatal codes drive the structured→degraded transition above |
agent_result |
agent tap | end of pass: cost, tokens, duration, turns, and a terminal reason_code — success on a clean pass, or one of the failure codes (timeout, provider_error, provider_auth, malformed_stream, missing_terminal, unsupported_schema, producer_nonzero, …) that the failure breaker counts |
evidence_writing |
agent tap | first Write/Edit/Bash of the pass that touches a GATE<N>-EVIDENCE.md |
_reopen |
presenter | synthetic, not on disk: the events.jsonl symlink retargeted (a new run after stop+resume), so a following viewer prints a divider and keeps narrating |
Specstride parses the spec through a single pluggable layer (lib/specstride_spec.py —
the one source of truth both the bash side and the critic call). Spec Kit tasks.md is
the input to use. OpenSpec changes are also supported, and the older hand-written
SPECS.md format still works but will be deprecated soon. The format is
auto-detected, or forced with --spec-format / SPECSTRIDE_SPEC_FORMAT.
speckit-tasks (recommended) — a GitHub Spec Kit
tasks.md. Each ## Phase N: heading becomes a Specstride phase, and every - [ ]
task line under it becomes a required deliverable the critic gates on (the task's
cited file paths are exactly what the grounding pass verifies):
## Phase 2: User Story 1 - <title> (Priority: P1)
### Implementation for User Story 1
- [ ] T003 [US1] Implement greet(name) in src/greet.py
- [ ] T004 [US1] Add a __main__ block to src/greet.pySpecstride also accepts Spec Kit implementations that group executable tasks under
priority headings such as ## P0 — Safety, ## P1 — Contracts, and repeated
## P1 — Security sections. Each task-bearing priority section becomes an
ordered phase with a unique gate id; the priority label remains in the title.
Trailing shared sections such as ## Dependency order and
## Definition of done are included in every normalized phase's context.
When the tasks.md lives inside a Spec Kit project (a .specify/ directory above
it), the feature's full design-doc set is injected into both the proposer prompt
and the critic as read-only context — they explain the why/how and are the
documents a grounding claim is verified against, but only the tasks are gated. The
set, in descending gating value (the order the context budget truncates from the
tail):
constitution.md → spec.md → plan.md → every contracts/*.md →
data-model.md → research.md → quickstart.md → every checklists/*.md.
Each is optional (included only when present). The total injected context
respects SPECSTRIDE_CONTEXT_BUDGET (default ~24000 chars), allocated across docs in
that priority order with per-doc floors — so a large plan.md cannot starve
contracts/ — and truncation is line-clean and code-fence-safe (never mid-line,
never a dangling ```), marked explicitly in the prompt.
A file named tasks.md, or any doc whose ## Phase N: or task-bearing ## P<N>
headings carry - [ ] task lines and no ### Acceptance criteria, is detected as
speckit-tasks unless it has the canonical OpenSpec change path described below.
A runnable Spec Kit example lives at
examples/speckit-tasks.example.md:
mkdir -p /tmp/specstride-speckit && cp examples/speckit-tasks.example.md /tmp/specstride-speckit/tasks.md
specstride run -w /tmp/specstride-speckit -s /tmp/specstride-speckit/tasks.md --verification offopenspec-change — an active
OpenSpec change at
openspec/changes/<change>/tasks.md. Each numbered level-2 task group becomes a
Specstride phase and its dotted checkbox items become required deliverables:
## 1. Domain contract
- [ ] 1.1 Add the export requirement.
- [ ] 1.2 Add empty and populated-log scenarios.
## 2. Implementation
- [ ] 2.1 Implement the exporter in `src/audit/export.py`.The change name becomes the feature-scoped Specstride state slug. Specstride injects the
change's proposal.md, every delta specs/**/spec.md, design.md, and matching
current openspec/specs/**/spec.md documents into both proposer and critic as
read-only context. The task list remains the gate; Specstride does not sync or archive
the OpenSpec change.
Canonical OpenSpec paths are detected before the generic tasks.md filename rule.
The numbered task shape is also content-detected when the file has another name.
A standalone example is available at examples/openspec-tasks.example.md.
native (legacy; to be deprecated soon) — a hand-written SPECS.md where each
phase is a level-2 heading whose text starts with Phase <N>, containing an
### Acceptance criteria block. It is still the fallback when nothing else is detected, so
existing SPECS.md projects keep running; new work should use a Spec Kit feature:
## Phase 0 — <title>
<description of the work>
### Acceptance criteria
- [ ] criterion one
- [ ] criterion twoInside a Spec Kit or OpenSpec project you rarely need -s. When it is omitted,
Specstride resolves the spec in this order (never picking silently between candidates):
<workdir>/SPECS.md, if a legacy one exists — checked first so existing projects are unaffected; remove it once the work has moved to a Spec Kit feature.<workdir>/.specify/feature.json→ itsfeature_directory→<dir>/tasks.md.- discover
<workdir>/specs/*/tasks.mdand<workdir>/openspec/changes/*/tasks.md— exactly one match is used; two or more with no--featureexitsE_SPEC(3), listing every candidate with the-sand--featureforms to disambiguate. - none of the above → an error naming every location tried.
So a single-feature project starts with just specstride run -w <project>:
specstride run -w ./ # resolves specs/001-.../tasks.md, no -sSpec Kit numbers every feature's tasks.md from 1 and builds them all into one
repo. Specstride keeps each feature's gates independent under
.specstride/features/<slug>/ (above), selected with --feature SLUG (or
SPECSTRIDE_FEATURE) — which also disambiguates step 3 of resolution. The inspection
CLI is feature-aware: specstride status/phases/verdicts/feedback/tail/events
take --feature and default to the last run's feature; specstride status --all lists
every feature with its approved/total phase counts; specstride resume --feature X
replays that feature's saved config (preserving its SPEC_FORMAT).
specstride run -w ./ --feature 001-login # run one feature to completion
specstride run -w ./ --feature 002-billing # then the next — independent gates
specstride status -w ./ --all # see both, side by sideWrite the work as a Spec Kit feature and let Spec Kit own its tasks.md; it is generated
from the feature's spec.md and plan.md. Never keep a hand-written SPECS.md beside
it for the same work: SPECS.md would be checked first and become a second,
un-reconciled source of truth. SPECS.md remains readable for existing projects, but it
will be deprecated soon; move non-feature work (migrations, refactors, ops roadmaps) into
a Spec Kit feature too.
Gate approvals live in .specstride/features/<slug>/gates/, which is what "is phase N
done" means. When the critic approves a phase, Specstride also ticks that phase's task
checkboxes in tasks.md, so the task list reads done as the feature completes
(SPECSTRIDE_TICK_TASKS=false turns that off). Checkbox state is outside the plan hash
and never re-plans anything.
Runtime is bash + python3 stdlib — no pip, no dependency manager, clone-and-run. Contributors run the test suite with the stdlib runner:
python3 -m pytest lib/.
Everything is set in .env (copy from .env.example; the real .env is
gitignored). Precedence: built-in defaults < .env < CLI flags.
Pick a backend per role; each role has its own accepted list:
Proposer backends: dsh[:provider/model] | claude | codex | bebop[:name] | prime[:variant]
Critic backends: dsh[:provider/model] | claude | codex | bebop | prime[:variant]
Both lists are generated from lib/backends.py, the single backend registry, and
checked against it by lib/test_backend_registry.py; to add a new backend, add it
to that registry first, then to its dispatch arm and these doc lines.
dsh— DeepSeek Harness'sheadlessprofile, using the provider/model in$DSH_HOME/settings.yamlunless a model override is supplied. Use backend refs such asdsh:zai/glm-5.3ordsh:qwen3.8-27b, or setSPECSTRIDE_DSH_MODEL=zai/glm-5.3. Bareglm-*model ids map to providerzai;qwen3.8-27bmaps to the LiteLLM-backedlocal-high/qwen3.8-27b-q5route. It is the default proposer; as critic it runs with model-facing tools disabled. The proposer may request persistent profile plugins whenSPECSTRIDE_DSH_PLUGIN_ALLOWLISTnames exact approvedpackage@semverspecs.claude— Anthropic. Claude Code CLI (proposer) + Messages API (critic).codex— OpenAI. Codex CLI (proposer) + Chat Completions (critic). Ships, but UNVERIFIED (no Codex CLI on the author's host to test against).bebop— a local selector → Compass/qwen via a shim (host-specific).prime[:variant]— bareprimeuses the standardprime-agentCLI and its configured default model, so no custom variants are required. If an optionalprime <variant>fleet launcher is installed, select it withprime:sol,prime:judge, etc. Proposer passes are fresh; Prime critics run without tools.
Allowlisted DSH plugin installation. Set SPECSTRIDE_DSH_PLUGIN_ALLOWLIST to a
comma-separated list of exact registry package@semver specs. When a DSH proposer
cannot complete a phase with existing tools, it may write the documented
specstride-dsh-plugin-request/v1 artifact and stop. Specstride validates the request,
runs dsh plugin --profile <profile> add --save-exact without a shell, archives
an audit receipt, emits plugin_installed, and restarts a fresh pass. Unlisted or
non-pinned specs halt visibly. Installed plugins persist in the DSH profile; the
DSH critic remains tool-free and cannot request them. See
Configuration.
Finding plugins safely. There is no curated DSH marketplace whose contents
are automatically trusted. Prefer, in order: bundles shipped or explicitly linked
by the official DeepSeek Harness repository,
official @deepseek-ai npm packages, or
internally reviewed bundles published to your private npm registry. Before adding
an exact version to the allowlist, inspect its repository, package.json,
dsh.bundle.patch, lifecycle scripts, dependencies, and cordis.patch.yml; use
npm pack package@version to review the tarball without installing it. Test new
plugins in a disposable DSH profile first. An npm listing, download count, or
allowlist entry is not proof that third-party code is safe—especially on hosts
where the DSH proposer runs with danger-full-access.
Observability parity. A Prime invocation emits the same signal classes as
Claude — init, text, tool, evidence, result — when its structured JSON
schema (prime-v3) is recognized; the invocation-start agent_observability
event announces mode=structured up front. If the schema is unavailable the
capability degrades explicitly (mode=raw-text, only text,result) rather
than pretending fine-grained events exist; a schema that parses but then breaks
mid-stream transitions structured→degraded (only the terminal result
stays trustworthy). The last-resort escape hatch is SPECSTRIDE_AGENT_STREAM=false,
which turns off structured capture entirely and restores the legacy raw
tee'd output — no per-tool events, and the redaction/payload policy no longer
applies, so use it only when you accept raw provider text in run.log.
Adding a stream format. A new structured schema is three edits: (1) write the
adapter module — a class built as Adapter(policy, *, expected_evidence=None) with
consume(record) (and optionally consume_raw/finish), importing from
lib/stream_seam.py, never from the tap; (2) add one row to the marked
"stream formats (the seam)" table in lib/agent_stream.py (one import + one
StreamFormat(...) row with its capability); (3) set the backend's registry entry's
stream value to the new format name in lib/backends.py — routing, the
invocation artifacts and the finalizer hand-off follow from that value alone, and any
failure (unknown format, SPECSTRIDE_AGENT_STREAM=false, missing tap) degrades to the
announced raw-text path.
Key knobs (see .env.example for all of them): SPECSTRIDE_MAX_REJECTS (3),
SPECSTRIDE_MAX_ITER, SPECSTRIDE_PROPOSER_TIMEOUT (1800s),
SPECSTRIDE_CRITIC_TIMEOUT (300s), SPECSTRIDE_CRITIC_MALFORMED_LIMIT (3),
SPECSTRIDE_PROPOSER_MAX_ERRORS (2), SPECSTRIDE_PROPOSER_MAX_CAPS (3),
SPECSTRIDE_PROPOSER_TIMEOUT_PHASE_<N> (per-phase pass ceiling; also
--proposer-timeout-phase N=SECONDS and a "phaseTimeouts" map in the
--verification-commands document — first of those three wins, else the global
value. --proposer-timeout and the overrides now round-trip through
last-run.conf, so specstride resume keeps the budget the run was planned for),
SPECSTRIDE_YIELD_POLL (30s; setting it at all overrides a learned poll interval),
SPECSTRIDE_YIELD_MAX_PER_ATTEMPT (4), SPECSTRIDE_YIELD_ALLOW_COMMAND (false), SPECSTRIDE_YIELD_COUNTS_AS_ITER (false),
SPECSTRIDE_PROMPT_MAX_BYTES (180000),
SPECSTRIDE_MAX_WALL_MIN (0 = unlimited),
SPECSTRIDE_CRITIC_GROUNDING (on), SPECSTRIDE_GIT_COMMITS (auto),
SPECSTRIDE_LEARNING (unset, which is off: off | suggest | apply; see
Learning), SPECSTRIDE_LEARNING_THROUGH (unset; set by a
mixture-of-loops contract to the run_id of the last decision it bound: resolve ignores
later apply entries until a re-derivation, while later reverts still count, so a bound
value is an upper bound on what runs, not a promise).
An unattended approve-your-own-work loop invites specific failure modes; each is guarded, all cheap:
-
Nonce-bound verdict. The critic must end with
VERDICT <nonce>: APPROVED|REJECTED, where<nonce>is random per call. The verdict is parsed only from the critic's reply, so a proposer can't approve its own gate by writingVERDICT …: APPROVEDinto the evidence. Missing/duplicate/wrong-nonce/ambiguous → REJECTED (fail-safe: never auto-approve on doubt). -
Grounded critic. Before the LLM call, the critic verifies the files the evidence cites (exists/size/mtime + bounded excerpt) and appends that snapshot, so claims about missing/empty files are visible. Read-only — never executes.
-
Stale-evidence rule. On REJECT the rejected
GATE<N>-EVIDENCE.mdis archived before the retry, so the proposer's file-existence gate isn't instantly satisfied by the old file (which would make "retry" a no-op). -
Single-run lock, timeouts, wall budget,
stop.flag. One orchestrator per workdir; per-pass and per-critic-call timeouts; an optional whole-run wall-clock budget; a manual clean halt. -
Pass watchdogs — stuck and futile.
--timeoutis only an absolute backstop: elapsed time cannot distinguish "still working" from "hung", and a bigger number just delays the same failure. Three signals actually end a bad pass, and each kill writes a checkpoint (reason, elapsed, the pass's last tool calls and words) to.specstride/features/<f>/pass-checkpoints/that the next pass's prompt carries forward, so a killed hour degrades into a note instead of vanishing. A kill is accounted by CLASS, not as one thing: a futility or hang kill (repeat_stall,progress_stall,idle_timeout) counts as an erroring pass and trips the failure breaker; a budget kill (hard_cap) says only that the work did not fit the pass, so it has its own bounded counter (SPECSTRIDE_PROPOSER_MAX_CAPS, default 3) and its own halt. Both surface to you (exit 4) rather than repeating for hours, with different remedies — see the exit-code table. Everypass_killedevent carries a stableclassfield, and a capped pass also emitspass_cost_unknown: a kill severs the provider stream, so the most expensive passes of a run report no cost at all. The accounting is backend-neutral: the structured Prime path reaches the same two breakers from the invocation's durableresult.json, which records the kill askill_reason+kill_classbeside its reason code.Signal Fires when Knob (default) idle no cpu-time growth anywhere in the pass's process tree — a genuinely hung pass, not a slow one (a busy docker execchild counts as progress)--idle-timeout/SPECSTRIDE_PROPOSER_IDLE_TIMEOUT(900s)disk stall nothing created or modified under the workdir, however busy the tree is ( .git/.specstride/node_modules/.venvexcluded — the harness and a detached long job write there on their own)--progress-timeout/SPECSTRIDE_PROPOSER_PROGRESS_TIMEOUT(1800s, 0 = off)repetition the same tool call (identical tool + target) issued N times in one pass and still the agent's most recent action — a retry loop, invisible to any cpu or wall-clock measure. A pass that retried something and moved on is untouched --repeat-limit/SPECSTRIDE_PROPOSER_REPEAT_LIMIT(5, 0 = off)repetition, process level the same child command line re-spawned N times in one pass — the same detector one level down, so it also covers backends that emit no tool events at all ( dsh,codex). Counted per agent tool call x command line: an agent that OCRs twelve screenshots runs one command line twelve times, once per image, from twelve different tool calls, and that is twelve pieces of work, not a retry loop (project B 003 phase 14, 2026-09-13 — a pass killed on the twelfth image). The same command line re-spawned under one tool call is still a retry loop and is still killed.SPECSTRIDE_PROPOSER_REPEAT_IGNOREis an extended regex of command lines never counted; its default covers the usual test runners, linters and type checkers (pytest,ruff,mypy,go test,make test, …) plus the per-file batch tools that have one command line and N inputs by construction (tesseract,convert,magick,compare,ffmpeg,pdftotext,identify, anchored at the command name). Set it to `` (empty) to count everything exceptsleep--repeat-limit/SPECSTRIDE_PROPOSER_REPEAT_LIMIT(5, 0 = off),SPECSTRIDE_PROPOSER_REPEAT_IGNORE -
Yield/resume — a pass may end cleanly while its job keeps running. A pass boundary and a measurement boundary are independent. When a phase's evidence needs a job that cannot finish inside one pass, the proposer writes one JSON artifact (
specstride-pass-yield/v1) to.specstride/features/<f>/yield/and exits normally; specstride launches or adopts the job in its own session (so no pass kill can reach it), waits for a declared predicate with no model session open, and resumes the phase with the job's exit code, duration and a bounded head+tail slice of its log in the next prompt. The orchestrator prints the contract into the proposer prompt — without that no agent will ever use it, and the alternative is what actually happened: an agent hand-rollingsetsid nohupwrappers so its work would survive the pass. A yield is not an error, not a stall, and by default does not burn an iteration.Piece Value predicates exit_code_file|pid|file_exists|file_stable|grep(+command, disabled unlessSPECSTRIDE_YIELD_ALLOW_COMMAND=true: it is execution with no pass running, and takes fixed argv only)required deadline_sec— the run holds the workdir lock for the whole waitbounds SPECSTRIDE_YIELD_MAX_PER_ATTEMPT(4) then exit 9;SPECSTRIDE_YIELD_POLL(30s);SPECSTRIDE_YIELD_COUNTS_AS_ITER(false)during a wait stop.flagis honoured every tick (exit 6, job left running);SPECSTRIDE_MAX_WALL_MINis checked in the loop, not only at phase boundaries -
Crash-safe resume. The current phase is derived from the
GATE*markers on start, not from a stored counter. Kill it anywhere, rerun the same command, it continues.--start-phase Noverrides. -
Per-phase git checkpoint. After each
GATE<N>-APPROVED, if the workdir is a git repo with changes, the orchestrator commitsspecstride: phase <N> approved — <title>. Never inits, never pushes.
| Code | Meaning |
|---|---|
0 |
all phases approved |
1 |
unexpected/internal error, and the critic-outage breaker: SPECSTRIDE_CRITIC_MALFORMED_LIMIT consecutive MALFORMED verdicts (default 3). MALFORMED is how the critic fails safe when it times out, is unreachable, or answers without a verdict line — the feedback it writes is contentless, so every further proposer attempt runs blind and the phase can never be approved. Without the breaker the run spends its whole MAX_REJECTS budget on a critic that is simply down (check_oscillation cannot catch it: it keys on criterion IDs, which a contentless feedback has none of). It emits run_stop reason=critic_unavailable; raise SPECSTRIDE_CRITIC_TIMEOUT, point --critic at a reachable backend, or raise SPECSTRIDE_CRITIC_MALFORMED_LIMIT, then specstride resume |
2 |
MAX_REJECTS exceeded — a human needs to arbitrate |
3 |
invalid spec/config |
4 |
budget exceeded — wall clock, MAX_ITER without evidence, or one of the two proposer breakers. The failure breaker (SPECSTRIDE_PROPOSER_MAX_ERRORS consecutive passes ending in an agent error: crash, timeout, auth/model error, malformed output, no terminal record, or a futility/hang watchdog kill — repeat_stall, progress_stall, idle_timeout; default 2) emits run_stop reason=proposer_consecutive_errors; raise --timeout / SPECSTRIDE_PROPOSER_MAX_ERRORS or fix the phase harness, then specstride resume. The cap breaker (SPECSTRIDE_PROPOSER_MAX_CAPS consecutive passes killed at the absolute pass ceiling, hard_cap; default 3) emits run_stop reason=proposer_cap_exhausted — that is a budget signal, not a failure: the passes may have been productive the whole time and the phase's work simply does not fit one pass. Make the long step outlive the pass (--long-job-phase / --long-job-cmd) or split the phase; raise SPECSTRIDE_PROPOSER_TIMEOUT only when the work genuinely is one indivisible pass. The yield budget (a declared yield's deadline_sec, the run's wall clock, or more than SPECSTRIDE_YIELD_MAX_PER_ATTEMPT yields in one attempt) emits run_stop reason=proposer_yield_budget — the job is left alone, never killed, so read its log under .specstride/features/<f>/yield-jobs/ first |
5 |
lock held by another run |
6 |
stopped via stop.flag (clean; specstride resume or rerun continues). Now also produced when the stop lands mid-proposer — specstride stop --now — which earlier versions mislabeled as 4 |
Off by default; the loop is fully legible with zero containers. When you want a dashboard too, specstride has two independent telemetry backends — enable either or both at once (dual-ship):
| Backend | Flag | URL flag (its own — never crossed) | Default | Ships to |
|---|---|---|---|---|
| Loki | --telemetry |
--loki-url |
:3100 |
Loki push API directly |
| OpenTelemetry | --otel |
--otel-url |
:4318 |
the OTLP Collector (which then feeds Loki + Prometheus) |
The two are wired separately: --loki-url only configures the Loki sink and
--otel-url only configures the OTEL sink. They do not share a URL — pointing
--otel-url at your Loki push port (or --loki-url at the Collector) will not work.
--telemetry ships the event stream straight to Loki's push API:
(cd "$SPECSTRIDE_HOME/telemetry" && docker compose up -d) # Grafana :3010, Loki :3110 (both free here)
specstride --telemetry --loki-url http://localhost:3110 -w ./myproject
# open http://localhost:3010 → the bundled Specstride dashboardThis is an independent deployment on its own ports (the defaults deliberately
avoid the common :3000/:3100). Every port is an .env variable.
--otel ships the same event stream over OTLP/HTTP+JSON to the bundled OTEL
Collector, which forwards logs to the same Loki (so the bundled dashboard is
unchanged) and turns cost/tokens/duration into first-class Prometheus metrics
(ralph_cost_usd_total, ralph_tokens_total, ralph_iter_duration_ms, …). Like
--telemetry, it's stdlib-only — no OTEL SDK, no pip:
(cd "$SPECSTRIDE_HOME/telemetry" && docker compose up -d) # + otel-collector :4318, Prometheus :9091
specstride --otel --otel-url http://localhost:4318 -w ./myprojectThe OTEL sink is driven only by --otel / --otel-url (env SPECSTRIDE_OTEL_URL) —
never by --loki-url. Note --otel-url points at the Collector on :4318, not
at Loki: the Collector is what fans OTLP out to Loki (logs) and Prometheus (metrics).
So a --loki-url change never affects OTEL, and vice versa.
--telemetry and --otel are independent: run either alone, or both at once
to dual-ship (Loki push and OTLP in parallel) — handy while migrating. To send
telemetry over OTEL only, pass --otel without --telemetry:
specstride --otel --otel-url http://localhost:4318 -w ./myproject # OTEL only
specstride --telemetry --loki-url http://localhost:3110 \
--otel --otel-url http://localhost:4318 -w ./myproject # both (dual-ship)The OTLP shipper mirrors the Loki shipper's add()/flush() seam and is covered
by unit, characterization, and old-vs-new parity tests
(lib/test_telemetry_parity.py and the shipper's own test module under lib/).
If Specstride built your project, say so with this badge:
[](https://github.com/mairp/specstride)Projects carrying it: agentic-netops, agentic-netops-srl and mixture-of-loops.
Specstride itself was built with the first version of the Ralph loop it grew out of.
Code is provider-agnostic and lives entirely on main. Branches differ only in
.env defaults:
main— defaults the proposer todshand the critic toclaude.bebop— overlay; defaults both roles tobebop compass(author's host).codex-demo— overlay; defaults both roles tocodex(OpenAI-only demo).
