Before submitting
Problem
On a committed-range native review of a large documentation-heavy diff (26 paths, +858/-137, risk tier high), two of the four reviewer lenses — review-readability and review-reliability — deterministically fail with reviewer-empty-output (stopReason: length) while the other two (review-risk, review-resilience) complete normally on the same ~190 KB materialized prompt. The lineage can therefore never reach 4/4 lenses, so the review can never complete, on any model the host has available.
Observations across 10 capture attempts (single-slot captures, fresh STATUS before each):
| lens |
attempts |
stopReasons |
model / effort |
| review-readability |
4 |
length ×3, stop ×1 |
nan/deepseek-v4-flash, effort high |
| review-readability |
2 |
length ×2 |
nan/glm5.3-flash, effort high |
| review-readability |
1 |
length |
nan/glm5.3-flash, effort medium |
| review-readability |
1 |
length |
openai-codex/gpt-6-astra, effort medium |
| review-reliability |
1 |
length |
nan/deepseek-v4-flash, effort high |
| review-reliability |
1 |
length |
nan/glm5.3-flash, effort high |
Every failing run lasts ~55–60 s and then reports stopReason: length with no text at all — the fixed output budget is consumed entirely (apparently by reasoning) before any artifact text is emitted. review-risk and review-resilience on the identical prompt size complete in one to three attempts with 2–4 KB results.
Related but distinct reports: #1167 (default thinking level can end a review — ours fails even at medium effort), #1083 (deterministic zero bytes — ours is truncation via length, not a silent child), #1259 (empty-output refusal discards model/token evidence — agreed, that made the diagnosis slower).
Two secondary observations from the same investigation:
Steps to reproduce
- Repo with a reviewable committed range of ~25–30 mostly-Markdown/Shell paths (~1000 changed lines), mostly normative prose.
gentle-ai review start committed-range (full base ref, committedOnly), high risk tier, four lenses, in-process host relay (--materialize=true slots).
- Capture
review-risk and review-resilience — they admit normally.
- Capture
review-readability (order 2): observe reviewer-empty-output, stopReason: length, ~55–60 s elapsed, no text produced.
- Repeat with any enabled model/effort routed via
subagents.json model_profiles — same result. review-reliability (order 3) behaves the same.
Expected and actual behavior
Expected: every lens slot can eventually produce its reviewer artifact on a reasonable model, or the failure envelope names the binding constraint (output bound reached, prompt size, lens) so the operator can act.
Actual: review-readability and review-reliability never produce text on any of three model families and two effort levels; the envelope only says stopReason: length, and the review is permanently uncompletable — it had to be abandoned with the audited operator_disposition flow.
gentle-pi version
3.7.0 (provider binary: dev-main build 3a19dbf8961bea9771b4e867cf3872e47ebe2699)
Pi version
0.87.1
Operating system
Windows (WSL)
Relevant logs or error output (optional)
{"tool":"gentle_review_capture","status":"blocked","outcome":"pi-host-relay-transport-failure",
"failure":{"kind":"reviewer-empty-output","stage":"pi","exit_code":null,"timed_out":false,
"elapsed_ms":58598,"timeout_ms":1062764,"reviewer":{"stopReason":"length"}},
"reason":"Reviewer produced no text for review-readability (stopReason: length).",
"mutation_performed":false,"mutation_outcome":"none"}
(Representative of all 10 attempts; only lens, stopReason and model/effort routing varied.)
Before submitting
Problem
On a committed-range native review of a large documentation-heavy diff (26 paths, +858/-137, risk tier
high), two of the four reviewer lenses —review-readabilityandreview-reliability— deterministically fail withreviewer-empty-output(stopReason: length) while the other two (review-risk,review-resilience) complete normally on the same ~190 KB materialized prompt. The lineage can therefore never reach 4/4 lenses, so the review can never complete, on any model the host has available.Observations across 10 capture attempts (single-slot captures, fresh STATUS before each):
Every failing run lasts ~55–60 s and then reports
stopReason: lengthwith no text at all — the fixed output budget is consumed entirely (apparently by reasoning) before any artifact text is emitted.review-riskandreview-resilienceon the identical prompt size complete in one to three attempts with 2–4 KB results.Related but distinct reports: #1167 (default thinking level can end a review — ours fails even at
mediumeffort), #1083 (deterministic zero bytes — ours is truncation vialength, not a silent child), #1259 (empty-output refusal discards model/token evidence — agreed, that made the diagnosis slower).Two secondary observations from the same investigation:
subagents.jsonmodel_profiles, but theagents/review-*.mdfiles carry amodel:frontmatter that managed-asset sync reverts on every capture — editing them looks like a supported knob but silently has no durable effect. A doc pointer to the real knob would save users a long detour.Steps to reproduce
gentle-ai review startcommitted-range (full base ref,committedOnly), high risk tier, four lenses, in-process host relay (--materialize=trueslots).review-riskandreview-resilience— they admit normally.review-readability(order 2): observereviewer-empty-output,stopReason: length, ~55–60 s elapsed, no text produced.subagents.jsonmodel_profiles— same result.review-reliability(order 3) behaves the same.Expected and actual behavior
Expected: every lens slot can eventually produce its reviewer artifact on a reasonable model, or the failure envelope names the binding constraint (output bound reached, prompt size, lens) so the operator can act.
Actual:
review-readabilityandreview-reliabilitynever produce text on any of three model families and two effort levels; the envelope only saysstopReason: length, and the review is permanently uncompletable — it had to be abandoned with the auditedoperator_dispositionflow.gentle-pi version
3.7.0 (provider binary: dev-main build
3a19dbf8961bea9771b4e867cf3872e47ebe2699)Pi version
0.87.1
Operating system
Windows (WSL)
Relevant logs or error output (optional)
(Representative of all 10 attempts; only
lens,stopReasonand model/effort routing varied.)