You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
[Tracking] Solheim provider smoke and archived review qualification #1896
Current decision — smoke only; review research stopped (2026-10-05)
qwen3.8-27b, under the tested Zoo-Code review configurations, has not demonstrated incremental independently verified review value. Do not enable it for review judgment, parameter extraction, probe selection, token routing or evidence-search dispatch. This is specific to the tested configurations, not a general conclusion about small models or interactive code investigation.
Stage 1B pilot completed: STOP. No retries, parser repairs, extra diagnostic requests or full-qualification expansion. The working diff has been reduced to Project A: real VS Code/provider smoke. Keep this tracking issue open for landing and the first live run of that simplified workflow.
Active scope / landing checklist
Land the smoke-only implementation (currently local, uncommitted/unpushed).
Verify protected environment final-vscode-review-smoke is main-only and requires approval; secret name remains FINAL_SMOKE_OPENAI_API_KEY. YAML does not establish repository policy.
Run the first simplified live smoke on main and record its metadata-only verdict.
Workflow: .github/workflows/solheim-provider-smoke.yml; implementation/docs: scripts/solheim-smoke/.
Manual main-only dispatch checks out the exact triggering SHA, builds trusted code, starts real VS Code with isolated storage and an empty temporary workspace, sends one correlated IPC no-tools task to https://api.solheim.ai/v1 using qwen3.8-27b, and requires actual output-token usage plus a nonempty nonpartial completion for the accepted task. One credential-bearing step; read-only GitHub permissions; no PR inputs, review prompts, delegation, shadow checks, judging or posting. Only whitelisted verdict.json status metadata is uploaded—not answers, prompts, host logs or credentials. Child environment/storage isolation is not a sandbox for malicious extension code; CI runs trusted main only. Code QA runs offline smoke tests/types, never live qualification.
Earlier question-lane success proved the old integration, not this refactor's pending live-CI behavior. No commit, push, deployment or GitHub environment changes are implied.
Smoke-only cleanup verification (2026-10-05)
Removed research harness/package commands and artifact docket from the active diff by moving them into the private local archive below. Retained smoke workflow/driver, correlated IPC acknowledgment and completion support, secret-safe startup logging and focused tests. Removed indentation-only lint-suppression churn; counts did not change. No research code was destroyed; the archive also contains the pre-cleanup patch/files and copies of the frozen pilot admission and comparison directories. Local /tmp preservation is recoverable now but is not durable remote storage.
Offline verification after cleanup: 82 tests passed (30 smoke/workflow, 37 event/IPC schema, 15 extension IPC/logging/acknowledgment). Smoke type-check, touched script/types lint, formatting and git diff --check passed. No live model calls, commit or push during cleanup. The first live simplified smoke remains pending.
Research chronology and capability boundary
Open-ended review / narrow judgment: earlier retrospective work found no gold recovery (0/56 historical comments; 0/13 reachable theme findings; 0 true positives among five QUESTION candidates, four false and one unjudged). Real quoted text did not prevent semantic misreading, including confident reversed claims. Prior one-rater, tuning and revision caveats remain. Homogeneous keys did not provide independent votes.
Template 1 / deterministic integration: secret-free generation of mechanically established Certain rows survived merge/report paths, stayed separate from Questions and Smoke, and remained shadow-only. Tightened literal-coverage rules produced four retrospective rows after same-set tuning; this is not prospective precision.
Template 2: conservative static migration matcher plus a separate seed-and-check executable layer. Static claim requires mechanically complete proof; unsupported complexity skips. Static migration rule produced zero rows on 23 saved heads. Registry covered four invariants: disabled repetition limit; global rate transfer with profile override preservation; legacy Host transfer/clear with header preservation; todo default with explicit opt-out preservation. Actual PR feat: add tiered tool-repetition detection (soft warning + hard stop) #1829 execution recovered hard=0/soft=2 once; other three probes passed. Repeated detection is one logical defect.
Diff-only template plus parameter extraction: five feat: add tiered tool-repetition detection (soft warning + hard stop) #1829 repetitions: correct template 5/5, correct parameters 0/5. Packet lacked the legacy-field/hard-soft relationship available to deterministic routing, so this rejected configuration does not isolate extraction capability.
ID-only semantic dispatch: trusted registry retained parameters/oracles. Fifteen requests: eight calibration, seven held out; four registered probes. Qwen returned [] on all held-out cases. Exact selection 5/7 in both lanes was negative-driven. Relevant pairs: deterministic 1/4, Qwen 0/4; misses 3 vs 4; irrelevant selections 1 vs 0; incremental relevant selection and additional Certain findings zero. Separate calibration repetition missed known probe 5/5; do not conflate with 5/5 template classification above. Only matching feat: add tiered tool-repetition detection (soft warning + hard stop) #1829 checkout was executed; other cases were routing-only. 18,429 tokens, median 406 ms; no known currency spend.
Stage 1A first-token plumbing: 24 synthetic requests, 3,192 reported tokens; no held-out PR transfer. Without label bias, full A/B/C coverage 0/6 at each temperature; equal +20 bias gave 5/6 at temperature 0 and 6/6 at 0.7. Rotating label meanings changed relevance-margin sign: temperature 0 +2.875 / -1.688 / -3.000; temperature 0.7 +4.107 / -2.411 / -4.286. First value had one incomplete repeat. Repeat stability did not cure semantic label sensitivity. Measurement failed admission; stopped token router. Logprobs were ranking diagnostics, not calibrated probabilities or defect confidence. Provider tokenizer/processing-order explanations were not established. Research-paper claims in the earlier conversation are not independent evidence for these local results.
Stage 1B amendment: remove model routing; share deterministic concern allocation and tools; compare scripted search against investigator/alternative/falsifier trajectories with one aggregate budget and an independent hidden verifier. No majority-vote confidence.
Stage 1B dataset and admission
Primary cases required untouched logical defects, independently supported narrow labels, deduplication, exact revision receipts, discriminatory verifier controls, secret-screened/hash-frozen packets and explicit human signoff. Hashes prove integrity, not label truth. Model/reviewer approvals were preserved separately and did not satisfy human gates. Human defect/control approvals were obtained before admission.
Passing-head control: Buffer/Uint8Array decode as UTF-8, preserving view boundaries, BOM and replacement behavior; PR itself contains a fix
4f2d586e82a6bb9c3b0d4d9b25fb7471db90d322
Verifier receipts included defective-revision failing arms, meaningful same-revision controls and fixed observations. #1852: partial burst 20 posts versus complete-message 3; fixed 3/3. #1713: inherited UTF-8 override while intended fallback controls remain. #1834 six passing arms; #1836 fifteen UTF-8 arms. No broad lifecycle/cache/encoding claims beyond these invariants.
#1831 surrogate defect excluded from primary scoring due substantive previous experimental exposure. #1724 native MCP policy bypass remained exploratory with unresolved logical-defect contamination. #1815 was explicitly previously reviewed and excluded. Audits covered available specified Claude/scratch histories, not inaccessible histories or pretraining. Case contamination is disqualifying; hidden evaluator knowledge of gold for verifier construction is allowed. Rejection inventory/provenance is preserved.
Frozen execution and isolation
Execution-v1: 12 aggregate charged action attempts per lane/case, not 12 per role; 180-second cap; Qwen maximum 6,000 completion tokens total and 500/call; temperature 0.3; three roles/four fixed rounds; at most two candidate records. Duplicates, invalid and failed actions charge budget; identical requests cache. Scripted search costs zero model tokens. Same tool schemas, snapshots, grounded target eligibility and bounded retrieval results (8,192 bytes / 20 hits / 80 lines). No retries, parallel extra opportunities, final reflection or parser normalization.
Hidden gold, decisive-location criteria, expected observations and receipts were excluded from both search lanes. Generic probes exposed observations, not gold-revealing case names/oracles; per-lane grounding kept one lane from using another's evidence. PR code executed only in a credential-free OS-isolated child: readonly source, private network/PID/mount/user namespaces, dropped capabilities, no-new-privileges, bounded output and terminable process groups. Trusted parsing used separately terminable workers. Provider credentials stayed in the controller memory, never in PR execution or artifacts.
Internal outcomes distinguished concern allocated → decisive evidence reached → verifier invoked with sufficient inputs → independently verified finding. Initial packet evidence was tracked separately from newly acquired evidence. Candidates required independent scope review and exact-evidence provenance; logical duplicates counted once. Unsupported candidates were not surfaced as verified findings.
457 tests / 93 suites passed before admission; hashes bound 96 code/config files plus 41 artifacts/runtime dependencies. This is test/admission evidence, not live review quality.
Frozen decision: full qualification only for ≥2 distinct incremental verified defects and ≤1 additional false surfaced finding. Zero or one incremental defect stops, even with richer evidence. Mechanical repair after held-out output requires an explicit new decision/version/new pilot, not silent adaptation.
Authorized pilot result (2026-10-05)
One explicitly authorized comparison, four cases, five successful Solheim completions. No retries or additional requests.
Scripted #1852's two evidence clusters deduplicate to one narrow defect, established from decisive initial-packet references plus the isolated 20-versus-3 verifier. No navigation advantage/crash claim is credited. Scripted #1713/#1836 stopped with no generic entry; control clusters did not allege defects.
Qwen issued five charged attempts, but completed only one evidence action (a #1834 source read). Of five returned content responses: one valid action JSON, one fenced JSON, three prose. Strict JSON parsing stopped all four lanes before any candidate. No falsifier turn ran. Scripted used 24 charged actions and zero model tokens. Provider reported 21,201 prompt + 688 completion = 21,889 total tokens; monetary cost unknown.
Interpretation limits: This is an action-interface conformance failure of the frozen configuration, not proof that interactive investigation cannot work. Archived responses do not retain all provider message fields/error bodies; serving-side explanations are unknown. All four initial packets already contained decisive evidence, so missing-cross-file search benefit was not demonstrated. With two positives and scripted recovering one, only one incremental defect remained available—below the frozen threshold of two. Do not lower the threshold or call the pilot a qualification pass. Zero false findings after early stops does not establish precision.
Preservation / handoff
The research harness, registry, frozen packets/labels, contamination and rejection inventories, receipts, role prompts, baseline, tests and artifact-update docket are retained outside the active smoke diff, not destroyed. Local private archive: /tmp/zoo-solheim-research-archive-HDTiBw/. This is local scratch, not a durable GitHub attachment. Raw model outputs/keys/private session histories are not published here. Original evidence directories remain available locally.
Project B deterministic review remains separate and inactive: diff → cheap routing → static rule/executable probe → Certain finding. It is not part of the smoke PR; no automatic template 3 or CI deployment is scheduled. A stronger model may be considered only with new authorization/versioned qualification against preserved baselines, not because it passes provider smoke. No further current-Qwen review experiments are authorized/scheduled.
The Claude artifact has not been edited by this agent. Its prior docket is preserved in the archive; a future artifact update must include Stage 1A/1B and STOP, not just the earlier ID-dispatch result.
Prior tracking snapshot — preserved historical evidence; its paths/roadmap/status are superseded by the update above
Decision: Project A is provider smoke only
qwen3.8-27b, under the tested Zoo-Code review configurations, has not demonstrated incremental review value over deterministic routing. Do not enable it for review selection or judgment.
Stop the current model's reviewer, narrow judge, parameter extractor, and semantic dispatcher roles. Do not keep inventing narrower Qwen roles. This conclusion is specific to the tested model/configurations; it is not a general claim about small models or semantic dispatch.
Project A — Solheim integration
Keep exactly one active Solheim integration: a lightweight real VS Code/provider smoke test, explicitly infrastructure validation, not review.
The working branch has been simplified to .github/workflows/solheim-provider-smoke.yml and scripts/solheim-smoke/:
One manual-only job, restricted to main and the exact triggering SHA. No PR input or CodeRabbit approval trigger.
Build the trusted extension/webview, activate real VS Code in an empty temporary workspace, start one no-tools task through correlated IPC, observe actual provider output and a completion result.
No review prompts, delegation, Qwen routing/judgment, shadow theme lane, deterministic review jobs, merged findings report, or PR posting.
One credential-bearing step; read-only GitHub permissions. Retain the existing protected environment/secret names for compatibility; repository policy must restrict deployments to main and require approval.
Upload only whitelisted verdict.json status metadata, not model answers, host logs, credentials, or review claims.
Code QA checks the smoke contracts and types; it does not run qualification experiments.
Status: implemented locally, still uncommitted/unpushed. The simplified workflow's first live-provider run is pending after landing on main and configuring environment approval. Earlier question-lane success does not prove this refactor's live CI behavior. No new Solheim calls were made during this cleanup.
Latest experiments
Diff-only template + free-form parameter extraction: rejected. Five PR 1829 repetitions selected the template 5/5 but extracted parameters correctly 0/5. The required legacy-field/relationship information was missing from Qwen's packet but available to deterministic routing. This does not isolate extraction capability. It does reject that exact configuration.
The cleaner replacement exposed only IDs and behavioral descriptions for four registered executable instances. Seeds and oracle values remained trusted and hidden from the model:
Disabled hard repetition limit must keep the soft limit disabled.
Global rate-limit transfer must preserve profile overrides.
Legacy Host conversion must clear the transferred field and preserve explicit headers.
Todo defaults must preserve explicit opt-out.
Both routers used the same production patches and eligible IDs; selected IDs bind to the same registry and deterministic runner. There were 15 live requests: eight calibration requests and seven held-out cases.
ID-only held-out measure
Deterministic
Qwen
Exact case selections
5/7
5/7
Relevant probe/case pairs selected
1/4
0/4
Relevant pairs missed
3
4
Irrelevant pairs selected
1
0
Incremental relevant selections by Qwen
—
0
Qwen returned [] on every held-out case. Equal exact accuracy comes from negatives, not useful positive selection. It avoided one deterministic false positive but added zero relevant selections and zero additional Certain findings.
In this separate ID-only experiment, Qwen missed the known repetition probe in all five calibration repetitions; it selected the other three direct migration introductions correctly. Do not conflate this with the earlier 5/5 template-classification result.
All four registered probes ran successfully at the actual PR 1829 head. The repetition probe reproduced the known hard-0/soft-2 defect; the other three passed. The matching-head comparator reused one deterministic execution for repeated selections, yielding one unique finding, not five defects. Other historical revisions were routing-only and explicitly not executed in that comparator.
Cost: 18,429 tokens, median request latency 406 ms, summed request latency 8.86 seconds. Deterministic routing took 0.95 ms across 11 distinct packets. Currency spend was not returned.
Limits: small manually chosen, single-author-labeled dataset; file-only patches; broad shared-decoding relevance labels do not assert defects; historical code may be in training data. No prospective/generalized performance claim is made. The conservative static migration rule emitted zero findings on the 23 saved heads and skips the complex task-history anchor it cannot prove.
Project B — separate deterministic validation/review
No model selects the invariant or defines success. Prioritize rules separately by prospective usefulness, maintenance cost, and runtime. Some can later become CI checks; others can remain advisory findings. No Project B jobs are enabled in Project A or Code QA by this cleanup, and template 3 is not the next task. Untrusted PR code must run only in an unprivileged job without provider credentials or secret-bearing artifacts.
Model qualification archive
The former pipeline was moved outside GitHub's workflow directory to scripts/model-qualification/. Preserve the defect-review benchmark, parameter-extraction harness, ID dispatcher, deterministic baselines, registry/oracles, comparison harnesses, historical fixtures, labels, and frozen packets. The archived posting CLI is disabled; live review qualification requires explicit opt-in and rejects the retired qwen3.8-27b.
The archive preserves 57 exact input records with SHA-256 verification: all 11 ID-dispatch diffs plus their labeled manifest, 22 safe legacy diff packets, and 23 legacy base references. One credential-shaped legacy packet remains private in original scratch, with only its hash/provenance archived. Raw live model outputs remain private; nothing was automatically published.
Reopen review architecture only when Solheim offers a materially stronger model and an authorized replay of the frozen suite demonstrates incremental useful signal over deterministic routing:
new model → frozen defect benchmark → parameter extraction → ID dispatch
Report misses, irrelevant selections, runtime/tokens, and executed additional findings; keep calibration separate from held-out cases and the oracle independent of the model. A provider smoke pass or model upgrade alone is not admission.
Verification and next steps
446 focused unit/workflow assertions passed across active smoke and preserved qualification infrastructure.
Both script projects type-check; shared-config ESLint passes without increasing extension suppressions.
Exact frozen bytes verified; the ID dataset exported and replayed offline with its original hash.
Keep this issue open for landing and the first live run of the simplified Project A smoke.
The Claude artifact handoff is docs/solheim-project-a-handoff.md; its agent must inspect the existing artifact and apply this separation without exposing keys or raw outputs.
The history below is retained for provenance. Its old reviewer roadmap and deployment instructions are superseded, not current plans.
Superseded research history and original roadmap
Goal
Use the free Qwen 3.8 27B model on Solheim (one model, 21 keys) in the review process as a reader of facts that scripts build. It supplements CodeRabbit and the Stryker mutation step. It must keep precision high, because one maintainer reads every row.
Status
Branch: feat/final-vscode-review-smoke. The work is not committed or pushed at the time of writing. No PR exists yet.
Built so far:
An agentic reviewer that runs in VS Code over IPC, with a findings format, triage, and a posting job (earlier work).
Map and scenario review prototypes (map-review.mts) and a gold set for PR 1829.
A verifier for the AGENTS.md persisted-setting checklist (checklist-review.mts), with a gate for the kind of setting.
A theme tagger with two checks (theme-review.mts): text against behavior, and state that several code paths write. It has a grounding check that verifies quotes (stage 5a), a base-commit filter, and a cap of 5 rows (stage 6).
A question lane (question-lane.ts). It runs each question as a short task through the real extension with disabledTools. It writes lane-verdict.json, and a failed question becomes a lane_error row apart from review rows.
Shadow-mode steps at the end of the smoke job. The report passes through redactForPublish, the upload needs a marker, and nothing posts to GitHub.
A shared model client with key rotation (solheim.ts).
Tests: 279 unit tests and 23 workflow tests pass.
Measured results
Recall on the six maintainer comments on PR 1829 is 0 of 6 for every model-based approach. A hand-written migration probe finds one comment with certainty in about 4 seconds, with no model.
The model gives confident wrong answers (80 to 99 percent) and repeats the same misreading across runs.
The 21 keys agree: answer probabilities differ by about 0.01, and text at temperature 0 is identical. More keys add speed and no independent opinion.
A verdict-first one-token judge picks "not enough text" on every instance. When the reason comes first, the verdict token is saturated.
A sweep of 103 past review comments on 24 PRs gave 57 defect claims. Checks built for PR 1829 reach 5 of its 6 comments and 1 of the other 51. Test strength is the largest theme at about 40 percent.
The lane ran 12 questions with 0 failures in 115 seconds. With an unreachable provider it reported FIRST_BYTE_TIMEOUT and 12 lane_error rows.
The disabledTools setting removes the tools that stalled runs (follow-up question and new_task).
The grounding check (stage 5a) checks that a quote exists. It does not check that the model read the quote correctly. One false claim survives with a real quote: that normalizeSoftLimit(0) raises the value.
The thesis that Qwen adds value beyond the scripts is a hypothesis, not a result. The one clear success so far, the migration probe, uses no model.
Decisions made
Keep the VS Code lane, so the check doubles as a smoke test.
The shadow step is a second reference to the provider secret. A workflow test now allows exactly two named steps (the driver and the shadow theme review). Revisit this if the secret must stay in one step.
Open decisions
Add a provider-profile checklist and list rules to AGENTS.md. A draft exists, and the priority is low. Two rules fail on main today: mimoApiKey and poeApiKey are missing from SECRET_STATE_KEYS. Nobody has confirmed whether these are real gaps.
Enable the VS Code lane in the CI shadow step. It needs the extension artifact path and xvfb.
When to turn on posting, and which token to use.
Retrospective baseline (2026-10-03)
The theme review ran on 23 of the 24 swept PRs. All 23 have exact diff fidelity: the reconstructed base and the changed files match GitHub. PR 1758 is excluded, because the shallow clone cannot reach its fork point. This is a retrospective set, so the numbers are baseline data and not evidence of generalization.
Rows: 73 in total. 5 were QUESTION rows and 42 were ok.
Candidate precision: of the 5 QUESTION rows, 0 are true positives, 4 are verified false positives, and 1 is unjudged. All 5 had a verified quote, so a real quote did not make the reading correct.
Gold recall: 0 of 56 defect comments, and 0 of the 13 that the two implemented themes (T4 and T7) can reach.
Deterministic upper bound from the sweep: 23 of 56 comments (3 state probes, 14 unit probes, 6 static checks). None was recovered, because no deterministic check exists in this pipeline yet. The scripts-only against scripts-plus-Qwen ablation cannot run until it does.
Funnel: 9 rows were removed for a muddled reason and 17 for low confidence. The grounding check removed 0. The base-commit filter ran on all 5 QUESTION rows and removed 0.
Tag routing: the tagger found T4 in 4 of 5 and T7 in 6 of 8 gold comments, but hints force those two tags. Without hints it found T10 in 3 of 3, T2 in 3 of 10, T6 in 1 of 4, and nothing for T1, T3, T8, T9 and T12.
Confident wrong answers: at least two ok rows at 98 percent assert the opposite of a gold comment (PR 1756 and PR 1829).
Caveats: one rater and no second judge. 20 of the 56 gold comments were anchored at earlier commits and may already be fixed. Five QUESTION rows is a very small base.
Decision (2026-10-03): suspend model judgment for T4 and T7
The retrospective baseline produced 0 of 13 reachable gold findings and 0 of 5 true-positive QUESTION rows. It also produced at least two high-confidence conclusions that contradict gold findings. Model judgment for T4 and T7 is suspended, not queued for better prompting.
The VS Code question lane stays. It is an end-to-end smoke lane for the provider and the extension. It is not a review-quality mechanism.
Deterministic and static checks come first, then probe templates.
Executable corroboration of Qwen claims is parked. It solved the wrong bottleneck: it verifies claims that the model almost never proposes correctly.
Qwen is revisited only after the deterministic system has measurable signal, and only for narrow roles where being wrong is cheap: choosing among known probes, extracting probe parameters, or routing into a deterministic checker. It is not asked whether something is a defect.
Default path: diff, then a deterministic candidate or check, then an executable proof, then a finding.
Optional model branch: diff, then Qwen proposes routing or parameters, then a deterministic proof, then a finding.
Solheim role: the open question
The goal is still that Solheim takes part in review. The question is now narrower: which review task can Qwen 3.8 27B own where its output can still be verified by a machine? The model proposes, and a machine proves or disproves. Candidate roles, in the order to test them:
Probe or check selection. Qwen picks which family from a fixed catalog of deterministic checks applies to a diff.
Parameter extraction. Qwen names the changed symbols, defaults, migration keys, and boundary values that a chosen check needs.
Bounded probe proposal in a small DSL, never arbitrary code.
Constraints for all three roles:
The oracle never comes from the model. A probe fails only against an invariant that comes from a registry, an existing test, or a maintainer comment. If the model writes both the probe and the expected result, this is judgment again, which the baseline suspended.
The baseline to beat is running every check on every PR. At the current catalog size (9 families), run-all has no routing error, and most families have a deterministic trigger (changed locale keys, a migration flag in the diff). Routing is worth building only for families whose applicability is a semantic decision, or when the catalog grows large or each probe is expensive.
Success criterion for the router role: Qwen selects checks that deterministic triggers miss, or extracts parameters that a deterministic parse cannot, and an executed probe confirms the finding, with few invented parameters.
If this also gives no signal, Qwen stays only in the smoke lane.
Result of the selection-only experiment (2026-10-03)
Retrospective, 23 evaluated PRs, 9 check families, 3 runs per PR (69 calls, no parse failures). 11 PRs have a gold family (13 pairs). One rater, no execution, and the triggers were written after reading the gold, so they are optimistic.
Selection recall: by majority vote, Qwen hit 8 of 13 gold pairs. The triggers hit 11 of 13. Run-all hits 13 of 13.
Volume and stability: Qwen selected 1.2 families per PR and nothing on 3 PRs. The triggers fire 1.5 per PR. The Qwen selections were stable across runs (mean pairwise Jaccard 0.86).
Beyond the triggers: 9 selections had no trigger. One is a gold pair (PR 1818, a changed literal passed to t(), which the narrow trigger does not read). Two are facts_for_a_decision, which has no trigger. The rest are unverified until the checks exist.
Parameter extraction: literals were extracted correctly (63 names, 0 invented), but a deterministic parse gives the same names. For the migration check on PR 1829, all three runs routed correctly and then extracted the wrong parameters: they named toolRepetitionSoftLimit as both the legacy and the new field, and seeded the soft limit. The needed seed is consecutiveMistakeLimit: 0. Exact match: 0 of 3. On PR 1664 the migration check was never selected.
Cost: 22.5 s per call on average (62 s at most), about 8.8k input tokens. A trigger costs nothing.
Reading: the success criterion is not met. Qwen found no gold family that a reasonable trigger would not find. Where extraction needs meaning, it failed in the same way as the judgment task: it misread which field plays which role. Where applicability is a fact about the diff, a trigger or run-all wins. Proposal for the maintainer: Qwen stays in the smoke lane for the current model. Revisit the router and extraction roles if Solheim offers a stronger model, because the verification framework does not change. The third role (bounded probe proposal) is untested, and its oracle problem remains.
Design rules
A model answer is never evidence of a defect by itself.
A row has one of three levels. Certain: a machine check establishes the complete claim that the row states. Supported: a machine check establishes important supporting facts or a counterexample, but some of the inference still comes from the model. Question: the substantive claim depends on model interpretation. The levels name what was established, not whether a model took part. Only certain and supported rows may post automatically later. Questions stay in shadow mode or go to the author only.
A check that runs PR code runs in the unprivileged build job, never in the job that holds the provider secret.
1. Deterministic static and list rules. Six historical comments look statically provable. Add the list rules (SECRET_STATE_KEYS, dynamicProviders, provider identifiers, locale parity). These give the first real Certain rows.
Template 1 built (2026-10-03): changed literal not in tests. Files: literal-coverage.ts, literal-coverage.test.ts, rules-review.mts (no model, about 1 second per PR). It extracts new English locale keys, quoted ns:key literals, union and enum members, exports, and model ids, and reports each one that no test mentions. 308 unit tests pass.
On the 23 evaluated PRs the tight rules give 4 rows in total (PR 1718: 1, PR 1726: 1, PR 1818: 2). The base mode gave 26 rows, of which only 2 were true, so the rules were tightened by sampling 31 distinct rows. The tuning used the same rows, so the 4-of-4 precision is optimistic.
Target comments, at their anchored commits: comments 5 and 6 (one defect at two commits, PR 1818) and comment 14 (PR 1755, gpt-6-sol and gpt-6-luna) are caught as Certain rows. Comment 11 (PR 1760) is not caught, because other specs mention the union members and the gap is a boundary spec for a consumer file. At the final heads these gaps are fixed, so the rule finds nothing there.
It cannot see: coverage at consumer boundaries, a test that mentions a literal without asserting behavior, literals built at run time, and non-English locale parity.
Wired into the shadow report (2026-10-03). Two new jobs, rules and report, hold no secret, have contents: read only, run plain node with no install, and post nothing. rules runs rules-review.mts and uploads a redacted artifact. report downloads the rules, theme shadow, and smoke artifacts and merges them with merge-report.mts into one shadow report with three separate sections: Certain findings (deterministic), Questions (model interpretation, not defects), and Smoke status (provider and extension, not review quality). A missing input shows as "Not available" and never as a pass. Locally, PR 1818 produced two Certain rows in the merged report. 323 unit tests and 27 workflow tests pass.
Untested until a real CI run (the workflow must reach the base branch first): a missing artifact with continue-on-error, the job-level success() || failure() result when a needed job fails or is skipped, origin/<base> after a fetch-depth: 0 checkout, and the live theme-rows.json write.
Remaining for step 1: the other static rules and the list rules (SECRET_STATE_KEYS, dynamicProviders, provider identifiers, locale parity).
2. Cluster the 14 unit-probe comments by proof pattern. Examples: boundary behavior, migration or default behavior, wiring to every consumer, integration through the real component, negative or error path. Do not write 14 bespoke tests first. Find out whether a small library of probe templates covers several comments per template.
Analysis done (2026-10-03), 23 comments (14 unit probes, 6 static, 3 state), one reader. A script could run on a new PR with no human-written test for 10 of 23 (9 distinct concepts of 22, because two comments are one defect at two commits). 1 needs a one-time invariant registry. 2 give facts for a decision and never a finding by themselves. 1 needs a person to record one real payload. 9 are bespoke or need intent. None of the 23 needs the real extension host.
Of the 10 script-only comments, 4 were already fixed at the final PR head, so a rule would only have caught them at the anchored commit. 6 are still true at the head.
Templates, by comments covered over effort: changed literal not in tests (4 comments, small), migration seed-and-check (2, medium), raw i18n key leak (1, small), workflow limits (1, small), spy restore (1, small), cross-surface flag parity (1, medium), fetcher contract fixture (1, medium), lifecycle (2, large). Facts for a decision cover 2 more.
Build first: changed literal not in tests, then migration seed-and-check, then raw i18n key leak. Together they cover 7 of 23. The raw i18n key leak rule has the highest false-positive risk: 89 of 1583 key literals in non-test source do not follow t() directly, so it needs a flow check into t(). The spy restore rule matches 28 of 803 spec files today.
Only one of the six static comments overlaps a list rule (B7). The other five are new rules. Full analysis: probe-templates.md (kept in the working notes, not in the repository).
3. The three state probes. They are more expensive and lifecycle-specific, so they come after step 2 shows what the probe abstraction should look like.
4. Keep the VS Code lane as the smoke lane. Enable it in the shadow step.
5. Gold sets and evaluation. Five more gold sets, a weekly regression run, and a prospective holdout kept apart from the 24 swept PRs. Leave-one-out is not a clean holdout here, because the checks were designed after reading these PRs.
6. Revisit Qwen against the deterministic baseline for routing and parameter extraction only (see the section above). The scripts-only against scripts-plus-Qwen ablation runs here, and run-all is the baseline to beat.
7. Shadow ratings, then posting. Proposed gates: do not leave shadow mode if fewer than 15 percent of the first 40 rated rows are useful. After posting is enabled, return to shadow mode after 3 rows in one week that the maintainer rates wrong.
8. Commit the branch and open a PR.
Parked: executable corroboration of model-derived rows (the evaluator for arithmetic and clamp expressions). Review memory is also parked until Qwen has a role that it can prove.
Risks
False confidence from the model.
Overfitting to one PR.
Prompt injection through diff text, comments, and locale strings.
Provider speed (about 20 tokens per second) and variation inside one instance.
Model drift when Solheim changes the served model.
Upkeep of rules and probes.
VS Code lane flakiness: stalled approvals and host memory kills in earlier runs.
Qwen may add nothing beyond the scripts. The ablation (step 4) answers this before more model infrastructure is built.
Current decision — smoke only; review research stopped (2026-10-05)
qwen3.8-27b, under the tested Zoo-Code review configurations, has not demonstrated incremental independently verified review value. Do not enable it for review judgment, parameter extraction, probe selection, token routing or evidence-search dispatch. This is specific to the tested configurations, not a general conclusion about small models or interactive code investigation.Stage 1B pilot completed: STOP. No retries, parser repairs, extra diagnostic requests or full-qualification expansion. The working diff has been reduced to Project A: real VS Code/provider smoke. Keep this tracking issue open for landing and the first live run of that simplified workflow.
Active scope / landing checklist
final-vscode-review-smokeis main-only and requires approval; secret name remainsFINAL_SMOKE_OPENAI_API_KEY. YAML does not establish repository policy.Workflow:
.github/workflows/solheim-provider-smoke.yml; implementation/docs:scripts/solheim-smoke/.Manual main-only dispatch checks out the exact triggering SHA, builds trusted code, starts real VS Code with isolated storage and an empty temporary workspace, sends one correlated IPC no-tools task to
https://api.solheim.ai/v1usingqwen3.8-27b, and requires actual output-token usage plus a nonempty nonpartial completion for the accepted task. One credential-bearing step; read-only GitHub permissions; no PR inputs, review prompts, delegation, shadow checks, judging or posting. Only whitelistedverdict.jsonstatus metadata is uploaded—not answers, prompts, host logs or credentials. Child environment/storage isolation is not a sandbox for malicious extension code; CI runs trusted main only. Code QA runs offline smoke tests/types, never live qualification.Earlier question-lane success proved the old integration, not this refactor's pending live-CI behavior. No commit, push, deployment or GitHub environment changes are implied.
Smoke-only cleanup verification (2026-10-05)
Removed research harness/package commands and artifact docket from the active diff by moving them into the private local archive below. Retained smoke workflow/driver, correlated IPC acknowledgment and completion support, secret-safe startup logging and focused tests. Removed indentation-only lint-suppression churn; counts did not change. No research code was destroyed; the archive also contains the pre-cleanup patch/files and copies of the frozen pilot admission and comparison directories. Local
/tmppreservation is recoverable now but is not durable remote storage.Offline verification after cleanup: 82 tests passed (30 smoke/workflow, 37 event/IPC schema, 15 extension IPC/logging/acknowledgment). Smoke type-check, touched script/types lint, formatting and
git diff --checkpassed. No live model calls, commit or push during cleanup. The first live simplified smoke remains pending.Research chronology and capability boundary
[]on all held-out cases. Exact selection 5/7 in both lanes was negative-driven. Relevant pairs: deterministic 1/4, Qwen 0/4; misses 3 vs 4; irrelevant selections 1 vs 0; incremental relevant selection and additional Certain findings zero. Separate calibration repetition missed known probe 5/5; do not conflate with 5/5 template classification above. Only matching feat: add tiered tool-repetition detection (soft warning + hard stop) #1829 checkout was executed; other cases were routing-only. 18,429 tokens, median 406 ms; no known currency spend.Stage 1B dataset and admission
Primary cases required untouched logical defects, independently supported narrow labels, deduplication, exact revision receipts, discriminatory verifier controls, secret-screened/hash-frozen packets and explicit human signoff. Hashes prove integrity, not label truth. Model/reviewer approvals were preserved separately and did not satisfy human gates. Human defect/control approvals were obtained before admission.
8cac2ac9705984df037caae347ed09aea669e2de177e7a8eb956180430d54f1375a0b7fee8ed955ed2c05e3a34d180aea3cf5022f7f160e82bbff6eb4f2d586e82a6bb9c3b0d4d9b25fb7471db90d322Verifier receipts included defective-revision failing arms, meaningful same-revision controls and fixed observations. #1852: partial burst 20 posts versus complete-message 3; fixed 3/3. #1713: inherited UTF-8 override while intended fallback controls remain. #1834 six passing arms; #1836 fifteen UTF-8 arms. No broad lifecycle/cache/encoding claims beyond these invariants.
#1831 surrogate defect excluded from primary scoring due substantive previous experimental exposure. #1724 native MCP policy bypass remained exploratory with unresolved logical-defect contamination. #1815 was explicitly previously reviewed and excluded. Audits covered available specified Claude/scratch histories, not inaccessible histories or pretraining. Case contamination is disqualifying; hidden evaluator knowledge of gold for verifier construction is allowed. Rejection inventory/provenance is preserved.
Frozen execution and isolation
Execution-v1: 12 aggregate charged action attempts per lane/case, not 12 per role; 180-second cap; Qwen maximum 6,000 completion tokens total and 500/call; temperature 0.3; three roles/four fixed rounds; at most two candidate records. Duplicates, invalid and failed actions charge budget; identical requests cache. Scripted search costs zero model tokens. Same tool schemas, snapshots, grounded target eligibility and bounded retrieval results (8,192 bytes / 20 hits / 80 lines). No retries, parallel extra opportunities, final reflection or parser normalization.
Hidden gold, decisive-location criteria, expected observations and receipts were excluded from both search lanes. Generic probes exposed observations, not gold-revealing case names/oracles; per-lane grounding kept one lane from using another's evidence. PR code executed only in a credential-free OS-isolated child: readonly source, private network/PID/mount/user namespaces, dropped capabilities, no-new-privileges, bounded output and terminable process groups. Trusted parsing used separately terminable workers. Provider credentials stayed in the controller memory, never in PR execution or artifacts.
Internal outcomes distinguished concern allocated → decisive evidence reached → verifier invoked with sufficient inputs → independently verified finding. Initial packet evidence was tracked separately from newly acquired evidence. Candidates required independent scope review and exact-evidence provenance; logical duplicates counted once. Unsupported candidates were not surfaced as verified findings.
457 tests / 93 suites passed before admission; hashes bound 96 code/config files plus 41 artifacts/runtime dependencies. This is test/admission evidence, not live review quality.
Frozen decision: full qualification only for ≥2 distinct incremental verified defects and ≤1 additional false surfaced finding. Zero or one incremental defect stops, even with richer evidence. Mechanical repair after held-out output requires an explicit new decision/version/new pilot, not silent adaptation.
Authorized pilot result (2026-10-05)
One explicitly authorized comparison, four cases, five successful Solheim completions. No retries or additional requests.
Incremental independently verified defects: 0. Additional false surfaced findings: 0. Decision: STOP.
Scripted #1852's two evidence clusters deduplicate to one narrow defect, established from decisive initial-packet references plus the isolated 20-versus-3 verifier. No navigation advantage/crash claim is credited. Scripted #1713/#1836 stopped with no generic entry; control clusters did not allege defects.
Qwen issued five charged attempts, but completed only one evidence action (a #1834 source read). Of five returned content responses: one valid action JSON, one fenced JSON, three prose. Strict JSON parsing stopped all four lanes before any candidate. No falsifier turn ran. Scripted used 24 charged actions and zero model tokens. Provider reported 21,201 prompt + 688 completion = 21,889 total tokens; monetary cost unknown.
Interpretation limits: This is an action-interface conformance failure of the frozen configuration, not proof that interactive investigation cannot work. Archived responses do not retain all provider message fields/error bodies; serving-side explanations are unknown. All four initial packets already contained decisive evidence, so missing-cross-file search benefit was not demonstrated. With two positives and scripted recovering one, only one incremental defect remained available—below the frozen threshold of two. Do not lower the threshold or call the pilot a qualification pass. Zero false findings after early stops does not establish precision.
Preservation / handoff
The research harness, registry, frozen packets/labels, contamination and rejection inventories, receipts, role prompts, baseline, tests and artifact-update docket are retained outside the active smoke diff, not destroyed. Local private archive:
/tmp/zoo-solheim-research-archive-HDTiBw/. This is local scratch, not a durable GitHub attachment. Raw model outputs/keys/private session histories are not published here. Original evidence directories remain available locally.Admission hash:
e78974923e55cbca97cf5ad08a8fc0210ce92a0f769b384772ebed478cedf519.Score hash:
69f3026f7b3d943603a29c59a25db48da98debfdb856772aadffd7e96b4cdf44.Execution-v1 hash:
fc059dab9eaf41d2daa828cfca366594769971e6402a4a1016b42ead5b193da5.Decision-v1 hash:
fcb9fb70dbdeaed7ef9b0445de48240c0ee07ebe1ea7a11b9ed936386b84c259.Project B deterministic review remains separate and inactive:
diff → cheap routing → static rule/executable probe → Certain finding. It is not part of the smoke PR; no automatic template 3 or CI deployment is scheduled. A stronger model may be considered only with new authorization/versioned qualification against preserved baselines, not because it passes provider smoke. No further current-Qwen review experiments are authorized/scheduled.The Claude artifact has not been edited by this agent. Its prior docket is preserved in the archive; a future artifact update must include Stage 1A/1B and STOP, not just the earlier ID-dispatch result.
Prior tracking snapshot — preserved historical evidence; its paths/roadmap/status are superseded by the update above
Decision: Project A is provider smoke only
Stop the current model's reviewer, narrow judge, parameter extractor, and semantic dispatcher roles. Do not keep inventing narrower Qwen roles. This conclusion is specific to the tested model/configurations; it is not a general claim about small models or semantic dispatch.
Project A — Solheim integration
Keep exactly one active Solheim integration: a lightweight real VS Code/provider smoke test, explicitly infrastructure validation, not review.
The working branch has been simplified to
.github/workflows/solheim-provider-smoke.ymlandscripts/solheim-smoke/:verdict.jsonstatus metadata, not model answers, host logs, credentials, or review claims.Status: implemented locally, still uncommitted/unpushed. The simplified workflow's first live-provider run is pending after landing on main and configuring environment approval. Earlier question-lane success does not prove this refactor's live CI behavior. No new Solheim calls were made during this cleanup.
Latest experiments
Diff-only template + free-form parameter extraction: rejected. Five PR 1829 repetitions selected the template 5/5 but extracted parameters correctly 0/5. The required legacy-field/relationship information was missing from Qwen's packet but available to deterministic routing. This does not isolate extraction capability. It does reject that exact configuration.
The cleaner replacement exposed only IDs and behavioral descriptions for four registered executable instances. Seeds and oracle values remained trusted and hidden from the model:
Both routers used the same production patches and eligible IDs; selected IDs bind to the same registry and deterministic runner. There were 15 live requests: eight calibration requests and seven held-out cases.
Qwen returned
[]on every held-out case. Equal exact accuracy comes from negatives, not useful positive selection. It avoided one deterministic false positive but added zero relevant selections and zero additional Certain findings.In this separate ID-only experiment, Qwen missed the known repetition probe in all five calibration repetitions; it selected the other three direct migration introductions correctly. Do not conflate this with the earlier 5/5 template-classification result.
All four registered probes ran successfully at the actual PR 1829 head. The repetition probe reproduced the known hard-0/soft-2 defect; the other three passed. The matching-head comparator reused one deterministic execution for repeated selections, yielding one unique finding, not five defects. Other historical revisions were routing-only and explicitly not executed in that comparator.
Cost: 18,429 tokens, median request latency 406 ms, summed request latency 8.86 seconds. Deterministic routing took 0.95 ms across 11 distinct packets. Currency spend was not returned.
Limits: small manually chosen, single-author-labeled dataset; file-only patches; broad shared-decoding relevance labels do not assert defects; historical code may be in training data. No prospective/generalized performance claim is made. The conservative static migration rule emitted zero findings on the 23 saved heads and skips the complex task-history anchor it cannot prove.
Project B — separate deterministic validation/review
Preserve the useful substrate:
diff → cheap routing → static rule / executable probe → Certain findingNo model selects the invariant or defines success. Prioritize rules separately by prospective usefulness, maintenance cost, and runtime. Some can later become CI checks; others can remain advisory findings. No Project B jobs are enabled in Project A or Code QA by this cleanup, and template 3 is not the next task. Untrusted PR code must run only in an unprivileged job without provider credentials or secret-bearing artifacts.
Model qualification archive
The former pipeline was moved outside GitHub's workflow directory to
scripts/model-qualification/. Preserve the defect-review benchmark, parameter-extraction harness, ID dispatcher, deterministic baselines, registry/oracles, comparison harnesses, historical fixtures, labels, and frozen packets. The archived posting CLI is disabled; live review qualification requires explicit opt-in and rejects the retiredqwen3.8-27b.The archive preserves 57 exact input records with SHA-256 verification: all 11 ID-dispatch diffs plus their labeled manifest, 22 safe legacy diff packets, and 23 legacy base references. One credential-shaped legacy packet remains private in original scratch, with only its hash/provenance archived. Raw live model outputs remain private; nothing was automatically published.
Reopen review architecture only when Solheim offers a materially stronger model and an authorized replay of the frozen suite demonstrates incremental useful signal over deterministic routing:
new model → frozen defect benchmark → parameter extraction → ID dispatchReport misses, irrelevant selections, runtime/tokens, and executed additional findings; keep calibration separate from held-out cases and the oracle independent of the model. A provider smoke pass or model upgrade alone is not admission.
Verification and next steps
docs/solheim-project-a-handoff.md; its agent must inspect the existing artifact and apply this separation without exposing keys or raw outputs.The history below is retained for provenance. Its old reviewer roadmap and deployment instructions are superseded, not current plans.
Superseded research history and original roadmap
Goal
Use the free Qwen 3.8 27B model on Solheim (one model, 21 keys) in the review process as a reader of facts that scripts build. It supplements CodeRabbit and the Stryker mutation step. It must keep precision high, because one maintainer reads every row.
Status
Branch:
feat/final-vscode-review-smoke. The work is not committed or pushed at the time of writing. No PR exists yet.Built so far:
map-review.mts) and a gold set for PR 1829.AGENTS.mdpersisted-setting checklist (checklist-review.mts), with a gate for the kind of setting.theme-review.mts): text against behavior, and state that several code paths write. It has a grounding check that verifies quotes (stage 5a), a base-commit filter, and a cap of 5 rows (stage 6).question-lane.ts). It runs each question as a short task through the real extension withdisabledTools. It writeslane-verdict.json, and a failed question becomes alane_errorrow apart from review rows.smokejob. The report passes throughredactForPublish, the upload needs a marker, and nothing posts to GitHub.solheim.ts).Tests: 279 unit tests and 23 workflow tests pass.
Measured results
FIRST_BYTE_TIMEOUTand 12lane_errorrows.disabledToolssetting removes the tools that stalled runs (follow-up question andnew_task).normalizeSoftLimit(0)raises the value.The thesis that Qwen adds value beyond the scripts is a hypothesis, not a result. The one clear success so far, the migration probe, uses no model.
Decisions made
packages/types: tracked in Add packages/types to Stryker mutation testing #1892.Open decisions
AGENTS.md. A draft exists, and the priority is low. Two rules fail on main today:mimoApiKeyandpoeApiKeyare missing fromSECRET_STATE_KEYS. Nobody has confirmed whether these are real gaps.xvfb.Retrospective baseline (2026-10-03)
The theme review ran on 23 of the 24 swept PRs. All 23 have exact diff fidelity: the reconstructed base and the changed files match GitHub. PR 1758 is excluded, because the shallow clone cannot reach its fork point. This is a retrospective set, so the numbers are baseline data and not evidence of generalization.
ok.okrows at 98 percent assert the opposite of a gold comment (PR 1756 and PR 1829).Decision (2026-10-03): suspend model judgment for T4 and T7
The retrospective baseline produced 0 of 13 reachable gold findings and 0 of 5 true-positive QUESTION rows. It also produced at least two high-confidence conclusions that contradict gold findings. Model judgment for T4 and T7 is suspended, not queued for better prompting.
Default path: diff, then a deterministic candidate or check, then an executable proof, then a finding.
Optional model branch: diff, then Qwen proposes routing or parameters, then a deterministic proof, then a finding.
Solheim role: the open question
The goal is still that Solheim takes part in review. The question is now narrower: which review task can Qwen 3.8 27B own where its output can still be verified by a machine? The model proposes, and a machine proves or disproves. Candidate roles, in the order to test them:
Constraints for all three roles:
Result of the selection-only experiment (2026-10-03)
Retrospective, 23 evaluated PRs, 9 check families, 3 runs per PR (69 calls, no parse failures). 11 PRs have a gold family (13 pairs). One rater, no execution, and the triggers were written after reading the gold, so they are optimistic.
t(), which the narrow trigger does not read). Two arefacts_for_a_decision, which has no trigger. The rest are unverified until the checks exist.toolRepetitionSoftLimitas both the legacy and the new field, and seeded the soft limit. The needed seed isconsecutiveMistakeLimit: 0. Exact match: 0 of 3. On PR 1664 the migration check was never selected.Reading: the success criterion is not met. Qwen found no gold family that a reasonable trigger would not find. Where extraction needs meaning, it failed in the same way as the judgment task: it misread which field plays which role. Where applicability is a fact about the diff, a trigger or run-all wins. Proposal for the maintainer: Qwen stays in the smoke lane for the current model. Revisit the router and extraction roles if Solheim offers a stronger model, because the verification framework does not change. The third role (bounded probe proposal) is untested, and its oracle problem remains.
Design rules
Next steps
Work in this order:
SECRET_STATE_KEYS,dynamicProviders, provider identifiers, locale parity). These give the first real Certain rows.literal-coverage.ts,literal-coverage.test.ts,rules-review.mts(no model, about 1 second per PR). It extracts new English locale keys, quotedns:keyliterals, union and enum members, exports, and model ids, and reports each one that no test mentions. 308 unit tests pass.gpt-6-solandgpt-6-luna) are caught as Certain rows. Comment 11 (PR 1760) is not caught, because other specs mention the union members and the gap is a boundary spec for a consumer file. At the final heads these gaps are fixed, so the rule finds nothing there.rulesandreport, hold no secret, havecontents: readonly, run plainnodewith no install, and post nothing.rulesrunsrules-review.mtsand uploads a redacted artifact.reportdownloads the rules, theme shadow, and smoke artifacts and merges them withmerge-report.mtsinto one shadow report with three separate sections: Certain findings (deterministic), Questions (model interpretation, not defects), and Smoke status (provider and extension, not review quality). A missing input shows as "Not available" and never as a pass. Locally, PR 1818 produced two Certain rows in the merged report. 323 unit tests and 27 workflow tests pass.continue-on-error, the job-levelsuccess() || failure()result when a needed job fails or is skipped,origin/<base>after afetch-depth: 0checkout, and the livetheme-rows.jsonwrite.SECRET_STATE_KEYS,dynamicProviders, provider identifiers, locale parity).t()directly, so it needs a flow check intot(). The spy restore rule matches 28 of 803 spec files today.probe-templates.md(kept in the working notes, not in the repository).Parked: executable corroboration of model-derived rows (the evaluator for arithmetic and clamp expressions). Review memory is also parked until Qwen has a role that it can prove.
Risks