Skip to content

bug(review): readability/reliability lenses deterministically exhaust the relay output bound (stopReason length) on any model — review uncompletable #1483

Description

@jparradog

Before submitting

  • I searched open and closed issues and did not find a report of this problem.
  • I reviewed this report and removed credentials, tokens, private paths, hostnames, and other sensitive data.

Problem

On a committed-range native review of a large documentation-heavy diff (26 paths, +858/-137, risk tier high), two of the four reviewer lenses — review-readability and review-reliability — deterministically fail with reviewer-empty-output (stopReason: length) while the other two (review-risk, review-resilience) complete normally on the same ~190 KB materialized prompt. The lineage can therefore never reach 4/4 lenses, so the review can never complete, on any model the host has available.

Observations across 10 capture attempts (single-slot captures, fresh STATUS before each):

lens attempts stopReasons model / effort
review-readability 4 length ×3, stop ×1 nan/deepseek-v4-flash, effort high
review-readability 2 length ×2 nan/glm5.3-flash, effort high
review-readability 1 length nan/glm5.3-flash, effort medium
review-readability 1 length openai-codex/gpt-6-astra, effort medium
review-reliability 1 length nan/deepseek-v4-flash, effort high
review-reliability 1 length nan/glm5.3-flash, effort high

Every failing run lasts ~55–60 s and then reports stopReason: length with no text at all — the fixed output budget is consumed entirely (apparently by reasoning) before any artifact text is emitted. review-risk and review-resilience on the identical prompt size complete in one to three attempts with 2–4 KB results.

Related but distinct reports: #1167 (default thinking level can end a review — ours fails even at medium effort), #1083 (deterministic zero bytes — ours is truncation via length, not a silent child), #1259 (empty-output refusal discards model/token evidence — agreed, that made the diagnosis slower).

Two secondary observations from the same investigation:

Steps to reproduce

  1. Repo with a reviewable committed range of ~25–30 mostly-Markdown/Shell paths (~1000 changed lines), mostly normative prose.
  2. gentle-ai review start committed-range (full base ref, committedOnly), high risk tier, four lenses, in-process host relay (--materialize=true slots).
  3. Capture review-risk and review-resilience — they admit normally.
  4. Capture review-readability (order 2): observe reviewer-empty-output, stopReason: length, ~55–60 s elapsed, no text produced.
  5. Repeat with any enabled model/effort routed via subagents.json model_profiles — same result. review-reliability (order 3) behaves the same.

Expected and actual behavior

Expected: every lens slot can eventually produce its reviewer artifact on a reasonable model, or the failure envelope names the binding constraint (output bound reached, prompt size, lens) so the operator can act.

Actual: review-readability and review-reliability never produce text on any of three model families and two effort levels; the envelope only says stopReason: length, and the review is permanently uncompletable — it had to be abandoned with the audited operator_disposition flow.

gentle-pi version

3.7.0 (provider binary: dev-main build 3a19dbf8961bea9771b4e867cf3872e47ebe2699)

Pi version

0.87.1

Operating system

Windows (WSL)

Relevant logs or error output (optional)

{"tool":"gentle_review_capture","status":"blocked","outcome":"pi-host-relay-transport-failure",
 "failure":{"kind":"reviewer-empty-output","stage":"pi","exit_code":null,"timed_out":false,
 "elapsed_ms":58598,"timeout_ms":1062764,"reviewer":{"stopReason":"length"}},
 "reason":"Reviewer produced no text for review-readability (stopReason: length).",
 "mutation_performed":false,"mutation_outcome":"none"}

(Representative of all 10 attempts; only lens, stopReason and model/effort routing varied.)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions