Skip to content

Test: verify retention release and restore against independent model (#128) - #159

Merged
flyingrobots merged 4 commits into
mainfrom
test/128-retention-release-restore
Oct 3, 2026
Merged

flyingrobots merged 4 commits into
mainfrom
test/128-retention-release-restore

Conversation

@flyingrobots

@flyingrobots flyingrobots commented Oct 2, 2026 •

Copy link
Copy Markdown
Owner

Missing evidence and outcome

The retention model omitted release and restore. Its accepted-publication update also copied generation and anchors from the candidate, which could conceal an incorrect operation recipe. This PR adds both operations and advances the reference state from the requested operation, initial fixture and prior model instead.

Change kind: correction of missing verification and oracle improvement. No production behavior changes. Every history uses fresh migrated storage; after every step, fenced observations must match the complete namespace/generation map, anchor sets and liveness. Exact typed stale/retry refusal checks remain. No harness-count assertion is introduced.

Evidence

The operation alphabet now contains both namespace initials, successor, release, restore, retry and stale initial. Every three-operation history is explored, including initial → release → restore and interleavings with the other namespace. Precondition no-ops are disclosed; the 343-history figure describes a bound, not proof of correctness.

There is no claimed production bug or fabricated RED against working main. A copied production mutation that hides empty selected roots survives the old suite and fails the new one. Separate oracle calibrations make release retain anchors or restore remain empty; both fail actual fenced anchor-set comparisons against the unchanged model. Unmodified production passes in debug/release. Mutants use separate source/build directories and are not committed.

Formatting, source structure and all-feature/minimal-feature warnings-denied Clippy pass in copied Docker source. Full workspace debug/release tests and doctests pass on candidate f8ab8128c0b591e24110dfc883905d6995bde6c3. All four required hosted CI jobs pass on that exact head (run 37057259727). The first broad local run exhausted inodes in the shared 2 GiB test filesystem; its log is retained. The filesystem was expanded to 4 GiB without deleting existing evidence, and the full debug/release run then passed. CodeRabbit was rate-limited; its successful status is not review approval. The evidence record identifies the owning contracts, replay commands, calibration results, fixture limitations and deletion criteria.

Current landing evidence

Integrated candidate 64bbbf915d87e43fa5c902ddeb0d893f246f9af5 passed a full fresh copied-Docker command chain and all four hosted jobs. The independent review found no implementation defect and identified one finite calibration-evidence gap. Three additional isolated production-reader mutations each failed the intended namespace-map, liveness or selected-root-generation assertion; restored model laws passed debug/release. Patches and normalized raw RED/GREEN receipts are now linked from the existing evidence document.

Final successor e768c8fcf769f25c422762cf67fa93384363ea68 changes documentation/receipts only; runtime code and test expectations are unchanged. Markdown lint, whitespace validation and replay patch applicability pass. The initial receipt formatting failure is retained in commit history and corrected by the successor. Exact-head independent review APPROVES and closes E1; CodeRabbit also approves with no actionable findings. All four jobs in final hosted run 37149868953 pass on this exact head. Older green checks are not substituted for final-head acceptance.

Scope, compatibility and limits

No format, identity, dependency, performance, synchronization, recovery or API changes. This is not a new fault campaign or proof of production platform eligibility; the existing model fixture deliberately isolates post-admission behavior. Histories beyond three steps and arbitrary anchor sets remain outside this model. Existing golden, corruption and crash owners remain in place.

The branch started from main 6051abb, after prerequisite #99 merged, and integrated main 1325841cbcd27f4c504870728e3ca87d87a2c4cc through normal merge 64bbbf915d87e43fa5c902ddeb0d893f246f9af5. The only conflict was the changelog; all entries were preserved. It does not depend on #156, #157 or #158. The prepared historical version was adapted selectively: newer exact refusal laws were retained, the candidate-derived oracle was replaced, and no case-count assertion was carried over. Original roadmap checkboxes are unchanged; the integration commit remains the final completion evidence.

Closes #128. Refs #131, #132.

Alternatives rejected: candidate-derived expected state and harness-count assertions do not independently establish the retention contract. Security implications: production authorization, identity checks and storage behavior are unchanged; these tests improve detection of incorrect retained state.

@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, add credits to your account and enable them for code reviews in your settings.

@coderabbitai

coderabbitai Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: ASSERTIVE
  • Plan: Advanced
  • Run ID: 12a86022-a960-48da-8cbc-b07c4035943f
📥 Commits

Reviewing files that changed from the base of the PR and between 1325841 and e768c8f.

📒 Files selected for processing (11)
  • CHANGELOG.md
  • docs/formats/segment-store-v2/requirements.md
  • docs/testing-evidence/retention-release-restore-model.md
  • docs/testing-evidence/retention-release-restore-model/liveness-red.txt
  • docs/testing-evidence/retention-release-restore-model/liveness.patch
  • docs/testing-evidence/retention-release-restore-model/manifest-red.txt
  • docs/testing-evidence/retention-release-restore-model/manifest.patch
  • docs/testing-evidence/retention-release-restore-model/restored-green.txt
  • docs/testing-evidence/retention-release-restore-model/root-generation-red.txt
  • docs/testing-evidence/retention-release-restore-model/root-generation.patch
  • src/adapters/retention/retention_model_tests.rs

Included review availability: This review used your included allowance. Your plan provides up to 1 included review per hour; 0 remain after this review.

📜 Recent review details
🔇 Additional comments (1)
src/adapters/retention/retention_model_tests.rs (1)

15-15: LGTM!

Also applies to: 26-27, 44-47, 54-54, 58-59, 218-219, 228-310, 320-320, 333-333, 350-353, 359-359, 366-370, 371-374, 400-400, 403-405, 447-460


Summary by CodeRabbit

  • Documentation
    • Updated retention documentation with coverage for release and restore operations, including how namespace generations, anchor sets, and liveness are checked across operation histories.
  • Tests
    • Expanded retention-model validation from 125 to 343 three-operation histories, adding release- and restore-starting sequences.
    • Recorded successful debug and release test runs, along with additional checks for detecting discrepancies in retention state. Production behavior is unchanged.

Walkthrough

The retention model now includes release and restore in its three-operation histories. It derives expected state from operations and compares it with fenced observations. New evidence documents the expanded histories, mutation calibrations, and passing debug and release test runs. Production behavior is unchanged.

Changes

Retention release/restore model coverage

Layer / File(s) Summary
Model release and restore transitions
src/adapters/retention/retention_model_tests.rs
The operation schedule adds release and restore. Their recipes set anchors to empty or template anchors, and advance namespace generation. The model derives expected state from each operation and prior state.
Verify expanded histories
src/adapters/retention/retention_model_tests.rs
Verification compares fenced observations with the model, includes operation schedules in assertion diagnostics, and adds tests for histories starting with release or restore.
Record coverage and calibration evidence
docs/testing-evidence/retention-release-restore-model*, docs/formats/segment-store-v2/requirements.md, CHANGELOG.md
The evidence documents 343 length-three histories, mutation calibrations, and passing debug and release runs. The requirements and changelog record the expanded coverage.

Priority: ⬇️ Low

Estimated code review effort: 3 (Moderate) | ~20 minutes

Change: Other

Merge Risk: ⚪ Minimal · up to e768c

The expanded retention tests and evidence leave no identified issue requiring a fix before merge.

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Linked Issues check ✅ Passed #128 asks for permanent release/restore model coverage and updated evidence. The model tests include release and restore in the seven-operation alphabet, explore all 343 three-step histories, create a…
Out of Scope Changes check ✅ Passed The test changes, changelog, requirements entry, evidence record, and calibration receipts support #128's model coverage and evidence objectives. The inspected test and evidence document report no pro…
Docstring Coverage ✅ Passed Docstring coverage is 80.00% which is sufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 10 functions across 1 files. (10 skipped: 1…
Title check ✅ Passed The title clearly and concisely identifies the main change: verifying retention release and restore against an independent model.
Description check ✅ Passed The description is substantially complete. It explains the problem, approach, change kind, alternatives, test evidence, scope, compatibility, and security impact. Some template items are not stated un…
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Three steps trace the roots in flight
Release clears anchors from the chart
Restore brings the template back
A model checks each fenced fact
Debug and release runs turn green
The evidence records what tests have seen

Comment @coderabbitai help to get the list of available commands.

@flyingrobots

Copy link
Copy Markdown
Owner Author

Validation receipt for exact candidate f8ab8128c0b591e24110dfc883905d6995bde6c3 (local and pushed heads match; worktree clean).

  • Full copied-source Docker workspace tests and doctests pass in debug and release. Formatting, source-structure checks, all-feature/minimal-feature warnings-denied Clippy, and Markdown lint pass.
  • All required hosted jobs are green on this SHA: CI run 37057259727, covering Rust quality gates, runtime fuzz smoke, dependency policy, and documentation/workflow integrity.
  • The initial broad local run failed with ENOSPC in namespace-capacity tests. The private shared ext4 filesystem had only 7,756 free inodes after cleanup of the failing fixtures. Expanded that filesystem from 2 GiB to 4 GiB without deleting prior work or evidence, then reran the complete debug/release workspace command successfully. Original full-validation.log and subsequent full-validation-expanded-filesystem.log are both retained in the author’s issue-128 scratch evidence.
  • The empty-root reader mutation passes the parent model and fails the expanded model. Incorrect release/restore candidate mutations fail fenced anchor comparisons against the independent expected state. These are calibrated coverage/oracle evidence, not a claim of a production bug on main.

CodeRabbit reports a review limit, and the hosted Codex reviewer reports exhausted review usage. Neither supplied an independent approval of this candidate. Green CI is not substituted for review or mainline integration.

@flyingrobots

Copy link
Copy Markdown
Owner Author

Code Lawyer evidence finding

Candidate 64bbbf915d87e43fa5c902ddeb0d893f246f9af5.

Severity File/lines Verified gap Acceptance
P3 / verification src/adapters/retention/retention_model_tests.rs:269–304,350–375; docs/testing-evidence/retention-release-restore-model.md The oracle now advances generations independently, but retained calibration receipts directly falsify missing roots and anchor comparisons only. They do not demonstrate the manifest-map, liveness or selected-root-generation checks failing at their own assertions. In separate copied production variants, hide manifest output, hide retention-head output, and return an otherwise canonical root with a wrong generation after normal read verification. Each focused run must compile and fail its intended assertion; retain exact patches/source coordinates and one consolidated restored-GREEN receipt.

This is an evidence gap, not a demonstrated Keep runtime bug. No candidate test expectation will be changed. The bounded calibration slice will run focused tests; current complete validation remains applicable to unchanged runtime/test code. @codex: independent second-opinion input is welcome; quota/status comments are not approval.

@chatgpt-codex-connector

Copy link
Copy Markdown

To use Codex here, create an environment for this repo.

@flyingrobots

Copy link
Copy Markdown
Owner Author

PR 159 independent review

Repository flyingrobots/keep; branch test/128-retention-release-restore; exact reviewed head 64bbbf915d87e43fa5c902ddeb0d893f246f9af5; tracked tree ee9b32733e943eb40e364822ae88aa83c5490899; target main at 1325841cbcd27f4c504870728e3ca87d87a2c4cc. Original change commit f8ab8128c0b591e24110dfc883905d6995bde6c3 starts from 6051abb25a9fd33ae7ee0de5614514b709a4d82a.

This is an independent Codex fallback applying the entire ULTRA STRICT agy-review protocol, authorized as a binding review gate. It is not a claim that agy or CodeRabbit performed the review. No source edits, tests, commits, pushes, comments, merges, configuration changes or delegation were performed by this reviewer. Only this report was written. The checkout remained clean and exact-head coordinates were verified locally and against live GitHub.

Findings

No demonstrated production or test implementation defect was found in the four-file target-to-head change or its merge interactions. The review is REQUEST CHANGES for the bounded evidence gate below, separately identified from a code bug.

E1 — generation calibration evidence required durable recording

Locations: src/adapters/retention/retention_model_tests.rs:278, :300, :306, :350, :357, :366; binding requirement docs/Testing Standards.md:48; evidence declaration docs/testing/enforcement.md:22.

The change replaces candidate-derived root generations with operation/model-derived generations and checked liveness advancement. Namespace-map equality, liveness equality and selected-root-generation equality are load-bearing observations supporting that model contract. The retained release/restore calibrations execute and fail the anchor equality assertion at :373; the hidden-empty-root calibration fails before root decoding at :366. They do not reach and fail the three generation observations. The supplied earlier exact-refusal calibrations falsify rejection diagnostics, not these observations. Rule 4 requires each new or materially changed load-bearing assertion to execute and fail for its intended reason; a passing assertion before another assertion fails does not establish that calibration.

This is missing evidence, not a verified defect in the current generation calculation. Suggested bounded closure: in isolated copies only, supply a missing manifest projection, missing retention-head projection, and a selected-root projection with a valid but incorrect generation after ordinary selection admission. Observe the intended namespace-map, liveness, and root-generation assertion failures separately, retain exact source/tree coordinates and mutation diffs, and consolidate the results in one evidence update. Leave candidate runtime/test behavior intact. Compilation, setup and unrelated failures cannot discharge this gate. Obtain a fresh exact-head review after any evidence commit.

Before this report was finalized, the parent completed precisely those three finite calibrations against archived 64bbbf9, without changing candidate production or tests. This reviewer independently inspected all mutation diffs, the traced launcher and actual failures: 159-calibration/manifest.log:49-52 fails the complete namespace map with observed empty versus expected A generation one; liveness.log:49-52 fails liveness with zero versus one; root-generation.log:49-52 fails selected root generation with two versus one. Every copy compiled, executed the named initial-A law, and exited 101 as verified by execution.log. original.rs compares byte-identically to candidate snapshot source; the three mutants each change only the appropriate production result projection. The generation mutant runs after normal root selection/digest admission and re-encodes a valid successor rather than causing decoder/setup failure. restored-green.log then executes all seven unmutated model partitions successfully in debug and release. The empirical gap is therefore closed. Remaining revision is one durable linked evidence record/receipt update under docs/testing-evidence/, after which the exact successor head needs a bounded documentation review. This report preserves the original finding and its attempted/completed closure; it does not ask for further calibration or production changes.

Verification Checklist

Scope, contracts and live feedback

  • Read AGENTS.md, the complete binding docs/Testing Standards.md, and docs/testing/enforcement.md. This review applies Keep's rules, including typed refusals, oracle separation, observed calibration, resource/gap disclosure and truthful durability evidence. Read the agy-review skill and all mandatory protocol clauses.
  • Read the complete four-file target-to-head diff and original change. The only Rust change is the test-only retention_model_tests.rs; all production files are byte-identical to target main. File size 460 lines remains below the hard 500-line ceiling; its review threshold was explicitly handled by inspecting every branch. Changed counters use checked arithmetic, BTreeMap ordering is deterministic, and no public boolean API, new dependencies, codecs, formats or async behavior are added.
  • Read live PR body and issue Integrate retention release/restore model coverage #128. Scope is missing verification/oracle correction, not a production bug fix. Acceptance requires fresh migrated storage, complete fenced namespace/generation/anchor/liveness agreement after every operation, lawful release and successor restore, exact stale/retry behavior, debug/release execution and preservation of the original roadmap contract. Existing production behavior already permits empty roots. The original ROADMAP/audit files are absent from this candidate; the diff does not change those historical files or their checkboxes. Issue Integrate retention release/restore model coverage #128 remains OPEN, appropriately pending integration.
  • Inspected the exhaustive queue 159-queue.json, its pagination implementation, and independently refreshed live GraphQL connections. There are 3 global comments, 0 reviews and 0 inline threads; every connection ends with hasNextPage=false, so no thread-comment page is omitted. The live CodeRabbit body now names the reviewed merge head and still reports rate limiting; the historical validation comment is pinned to f8ab812, not current-head approval. Codex usage-limit and CodeRabbit SUCCESS statuses are not independent approvals. No actionable review finding is hidden in those bodies.

Every changed path and its actual production boundary

Path traced Verified behavior and parallel-path agreement
Enumeration and fresh scenes retention_model_tests.rs:54 → :386 → :420-459; fixture filesystem_retention_test_fixture.rs:58 → :247-255
Initial A/B and repeated initial retention_model_tests.rs:145-152,178-195 → Recipe::publish:83-93 → fixture :90-107 / :145-158 → transition_preflight.rs:77-98 → publication_preparation.rs:22-60 → successor_manifest.rs:9-46,97-111
Successor, release and restore retention_model_tests.rs:215-217 → :228-274 → Recipe::publish:94-104 → fixture :112-127 → transition_planner.rs:92-124 / closure_verifier.rs:30-38 → publication_preparation.rs:35-59
Independent expected state retention_model_tests.rs:278-310 → apply :320-323
Exact retry retention_model_tests.rs:219-222 → Recipe::publish:79-104 → publication_execution.rs:25-30 → filesystem_retention_storage.rs:26-90 → filesystem_retention_current.rs:135-176,185-212
Stale initial retention_model_tests.rs:196-214 → fixture :90-107 → filesystem_retention_current.rs:135-176 → retention_model_refusal.rs:84-111
Revalidated durable publication Recipe::publish:104 → publication_execution.rs:21-47,50-147 → filesystem_retention_storage.rs:26-90 and filesystem stage capabilities
Complete fenced observations retention_model_tests.rs:130-135,333-378 → filesystem_retention_snapshot.rs:122-161,171-180,194-249 → retention_view_collector.rs:99-117 / filesystem_retention_current.rs:101-127,307-325
Refusal state preservation retention_model_tests.rs:314-329 → :398-401
#99 recovery compatibility filesystem_retention_storage.rs:31-46 → filesystem_retention_recovery.rs:119 → recovery_planner.rs:28-77 → recovery_execution.rs:98-120

The model holds no long-lived reader fence across publication: each snapshot is local and dropped after candidate construction/verification. Authority remains writer-locked through each history. No new shutdown, async cancellation, interruption, raw concurrent mutation, device/config transition or restart behavior is introduced. Existing crash/recovery paths remain separate and unchanged; reading the state-machine paths is not physical power-loss testing.

Every merge against both parents and retained integration invariants

All eight merges reachable since original base were inspected against both parents. Incoming runtime/test/evidence subsystems are byte-identical to current target main; this review verifies their interaction and retention, and does not reopen or claim a fresh independent audit of every unchanged imported subsystem. Earlier binding receipts were treated as constraints and checked against current code/merge coordinates, not transferred as blanket approval.

Merge SHA Both-parent diff and resolution Retained invariants
64bbbf915d87e43fa5c902ddeb0d893f246f9af5 Parents f8ab8128c0b591e24110dfc883905d6995bde6c3 and 1325841cbcd27f4c504870728e3ca87d87a2c4cc: first-parent diff imports 60 mainline files; second-parent diff is precisely the 4-file #128 change. Remerge diff shows only CHANGELOG conflict markers removed, with all five entries retained. New model file equals original #128 file; every incoming production file equals target. #99, #171, #146, #150 and #147 claims/invariants survive.
1325841cbcd27f4c504870728e3ca87d87a2c4cc Parents 64fafe3ddcc92bcc45a461a0161030b87559070d, 05658799fb259a445f68c8bd434135483d491e40: 4-file #147 diff against first; empty diff against second. Exact reviewed #147 integration tree; strict platform-admitted public stage laws retained.
05658799fb259a445f68c8bd434135483d491e40 Parents 5f90f22bde1582ecaa7215d85e7ce9db62975d25, 64fafe3ddcc92bcc45a461a0161030b87559070d: incoming 58-file correction diff against first; 4-file #147 diff against second. Remerge resolution only removes CHANGELOG conflict markers. Public stage checks preserve ext4 platform admission, exclusive creation, canonical sealed bytes and unsealed evidence. They do not replace the deliberately repository-admitted retention model fixture with a platform-eligibility claim.
64fafe3ddcc92bcc45a461a0161030b87559070d Parents 182e49520f98c6035828a739dcf3c224df535b84, 3626f6e2677a1d6d3af08b88245584800596ff55: 12-file #150 diff against first; empty diff against second. Exact reviewed #150 tree, including final caller-documentation correction.
657593fe50c824dd31dc328bf9e696183ef20767 Parents fa06adfde89b70be5aaf7356206b7d8c1fbce169, 182e49520f98c6035828a739dcf3c224df535b84: incoming 49-file diff against first; 12-file #150 change against second. Only CHANGELOG and rationale conflict markers removed; rationale retains both catalog-admission and sealed-observation sections. filesystem_catalog_publisher.rs:94-101 routes legacy locked-root constructor through production admission; filesystem_platform_profile.rs:61-63 checks strict profile. This catalog change does not alter model retention authority/preflight.
182e49520f98c6035828a739dcf3c224df535b84 Parents d0cff10d7c911d33d615c3aa2246ae2b4497432a, e781c0b276ec4d1f66a76668bd31893261a2e6dd: 15-file #146 diff against first; empty diff against second. Exact reviewed sealed-authority tree.
e781c0b276ec4d1f66a76668bd31893261a2e6dd Parents 15aa976a77d5b65a4b6e8f8a2f9e01d2ba3b0d58, d0cff10d7c911d33d615c3aa2246ae2b4497432a: 36-file #171 incoming diff against first; 15-file #146 change against second. CHANGELOG marker-only resolution retains both entries. sealed_segment.rs:69-90 keeps stage extraction private; observed_segment_stage.rs:25-70,79-84 preserves real write/durability effects and sealed metadata, never exposes a writable sealed stage. Retention model uses root publication, not sealed callback conversions.
d0cff10d7c911d33d615c3aa2246ae2b4497432a Parents 6051abb25a9fd33ae7ee0de5614514b709a4d82a, 07e6cc4875c05592b71bb1f8b9c80631a31bcda5: 36-file #171 change against first; empty diff against second. recovery_segment_classifier.rs:69-83 → recovery_segment_seal_framing.rs:9-66 refuses partial v1 seal framing contradictions before discard assessment. Version-two retention recovery retains its separate no-discard policy.

Constants, numerical claims and raw evidence

  • OPERATIONS length 7 (retention_model_tests.rs:54-62), fixed history depth 3 (:333,:386-402), seven distinct first-position tests (:420-459): mathematical exploration is 7 × 7 × 7 = 343. Parent alphabet is 5, giving historical 125. These agree with CHANGELOG, requirements KEEP-RETENTION-010, evidence :7, issue scope and PR body. There is no contractual test-count assertion; precondition no-ops are disclosed. Histories are not all successful release/restore witnesses.
  • Evidence :11 concrete Initial(A), Release, Restore trace: literal model generation 1 at :292, then checked successors at :298-301 give 2 and 3; anchors go fixture → empty → fixture at :291-295. Actual candidate generation independently comes from observed current successor at :253. B preservation follows complete-map comparison and A-only successor planning. Fresh debug/release logs execute all seven law partitions, identified below.
  • Unchanged fixture limits: restart byte cap 1,048,576 at :126; read limits are SegmentRecordLimit::MAXIMUM and LayoutEntryLimit::MAXIMUM; ReaderAttemptLimit::DEFAULT is 3 (reader_attempt_limit.rs:11-13). These are admitted test inputs/protocol retry bounds, not latency, memory-measurement or throughput claims. No changed limit or timing/buffer/rate constant makes a binding performance receipt stale. Sequence index is checked u32 naming only; model counters checked u64 and bounded by three steps. Namespace B is a fixed fixture identity, not caller input.
  • Pinned Rust 1.96.0 matches rust-toolchain.toml:2, evidence replay :23, and fresh log 159-validation.log:3-4. Original base SHA and Retention recovery, crash-matrix evidence, reader fence, and model-based transitions (item 6) #99 prerequisite match Git ancestry. No API/format/identity/dependency/sync/recovery/performance runtime diff exists against target. Existing golden/corruption evidence remains separate from the shared fixture/model oracle.
  • Old hidden-empty-root campaign: keep-audit/128/old-suite-mutant.log:38-49 compiles isolated /audit/keep128-cal-empty-old into /build/keep128-cal-empty-old and executes five original model laws, all PASS. Its model source SHA-256 21e18c6546782b77590f1710324a9f3a803edc0ed82f021bc24b6e225d8b8c6b equals git show 6051abb:.../retention_model_tests.rs. New campaign new-suite-mutant.log:38-84 compiles separate source/build and executes seven laws, all fail with missing selected root; Initial(B), Initial(A), Release is concretely at :71-72. Preserved source diff inserts only if root.root().anchors().is_empty() { return Ok(None); } after selected-root decode in production snapshot. The expanded copy's remaining model differences from candidate are only the later schedule-message addition; oracle and assertions match. This is calibrated coverage, not mainline runtime bug RED.
  • Wrong-release and wrong-restore source diffs were read directly from preserved Docker copies. Each changes only candidate construction in successor_recipe, leaving advance_model/verify untouched. Release changes empty candidate anchors to current anchors; restore changes fixture candidate anchors to empty. Candidate's current model source SHA-256 694dc53425d1e6520df27ff4ce6ea26fc2aa14d8c26a3e332542fdc3dc8cab19 matches historical /audit/keep128-source, whose Git HEAD is f8ab812. release-oracle-red.log:55-58,84-87,113 reaches anchor comparison with actual nonempty/expected empty. restore-oracle-red.log:55-58,84-87,113 reaches anchor comparison with actual empty/expected fixture. All seven partitions fail the intended comparison in each campaign; separate build/source prefixes exclude shared-artifact confusion. These are semantic candidate/oracle calibrations, not a requirement that publisher reject canonically lawful requests.
  • Generation assertion calibration: independently inspected the three newly supplied intended runtime failures and restored GREEN described in E1. No mandatory generation observation remains empirically uncalibrated. Durable recording of the new receipt remains pending at this reviewed head.
  • Historical focused GREEN: focused.log:38-51,90-103 executes all seven laws in debug and release; :105-159 completes warnings-denied workspace checking. Raw expanded full-validation evidence was retained alongside the original failed run. PR body/validation comment expressly pin this evidence to f8ab812; it is not relabeled as exact-merge-head execution. The external comment's 2→4 GiB/7,756-inode figures describe historical infrastructure observations, not implementation limits or correctness claims; the raw failed/full logs establish failed versus successful executions, but do not independently measure that inode/resize figure.
  • Every new evidence-document numerical figure and scoped claim was checked above. Unchanged incoming documentation/receipt figures retain their pinned historical coordinates and original bound/scope; current PR changes no parameters in them. Earlier 172/158/157/156-independent-review.md receipts were read for integration constraints and checked against exact incoming merge coordinates; no unchanged incoming numerical subsystem campaign is presented as newly rerun by this reviewer. No unrelated repository-wide numerical certification is claimed.

Errors, state machines, determinism and repository evidence standards

  • Independent model and producer use separate generation/anchor derivations; shared fixtures/decoders are explicitly limited foundations. Refusal helper retains exact typed error/source checks and expected coordinates. No generic is_err assertion substitutes for refusal precision. All unexpected success/failure dispositions terminate the law with operation/schedule diagnostics.
  • Release/restore prerequisites are honest no-ops, and verification still executes. Expected Published advances state only after production reports Published; any other result must match exact declared idempotence/refusal or fails. No model state is copied from publication candidate output. All generated histories are deterministic and finite; no random seed/fuzz campaign is claimed. Reported failing three-step schedule supports manual reduction by removing no-ops, as disclosed.
  • Medium-size filesystem designation is accurate; each fresh store owns its scratch scene. Production eligibility is intentionally outside the repository-admitted fixture and disclosed. New test notes declare oracle/deletion criteria; evidence document declares owner/change kind/profile/limits. Ordinary per-test ceilings, sandbox admission, suite SLO/latency and evolving-seed/reduction automation remain disclosed enforcement gaps, not asserted compliance. No new isolation or quantitative performance promise is added. Loops/conditionals serve systematic exploration and are not themselves defects.
  • Changed evidence prose has one physical paragraph per line; Markdown whitespace checks pass. No unrelated source refactor or roadmap completion-checkbox change is included. No new durable protocol decision requires a runtime ADR; the substantial oracle/evidence rationale is colocated under docs/testing-evidence.

Checks executed, inspected, pending and excluded

Executed by this reviewer: read-only Git source/diff/remerge/both-parent inspection; live GitHub PR/issue/queue/check queries; Docker read-only source comparisons and Git/SHA checks; exact-head/clean-status checks; whitespace checks against target and entire tracked tree. These are inspection/static checks, not host Rust execution. Only this report was written.

Inspected historical execution: focused debug/release GREEN, hidden-empty-root old/new calibration and wrong-release/wrong-restore calibration logs, their retained isolated source copies/build coordinates, earlier exact-refusal receipts, and scoped incoming review/validation receipts. Historical mutation commands and timestamps are not fully printed in the raw logs; immutable source comparisons and retained mutation diffs substantiate their source scope, while the logs establish compilation and intended runtime outcomes. No original calibration shell script was available. This limitation is not disguised as reviewer execution.

Inspected fresh integrated Docker validation: 159-validation-manifest.md and traced 159-validation.log. Log :1-10 identifies exact candidate tree, Rust 1.96.0, Linux aarch64, ext4 scratch and tmpfs negative fixture. :11-16 executes worldline, debug/optimized process-death matrix, conformance, structure and fmt; :17,:78,:83,:89 executes all/minimal feature checks and Clippy; debug model laws pass at :454-462, optimized at :2473-2480; full workspace commands are at :94,:2107; doc/doctest/MSRV/fuzz format/check/Clippy commands at :4125-4186 complete successfully in visible output. These commands were run by the parent, not this reviewer. The parent separately confirmed terminal exit 0 for session 9152 and restored focused session 75863. Restored log 159-calibration/restored-green.log:1-2 binds the same candidate tree, :7-16 records seven debug laws PASS, and :22-31 records seven optimized laws PASS. Process-death checks are not power-loss proof; fuzz compilation is not runtime fuzz evidence.

Hosted checks: independently refreshed exact-head run 37149193816, verified headSha equals 64bbbf915d87e43fa5c902ddeb0d893f246f9af5, status COMPLETED and conclusion SUCCESS, and all four jobs completed SUCCESS: Documentation/workflow integrity, Dependency policy, Rust quality gates and Runtime fuzz smoke. CodeRabbit status SUCCESS still accompanies rate limiting and supplies no review. Original run 37057259727 is historical f8ab812 evidence. Final successor-head hosted green and immediate pre-merge head/protection verification remain integration gates; no approval is issued while the requested durable E1 receipt update remains pending.

Excluded: host Rust tests, source mutations, new crash/random/concurrency/power-loss campaigns, long histories, arbitrary anchor sets, namespace exhaustion, full production platform certification from this fixture, and unrelated re-audit of byte-identical incoming subsystems. These exclusions match this evidence-only change; no green result is promoted into exhaustive correctness proof.

Verdict

The scoped code and merge review is complete with no demonstrated implementation defect. The generation-observation experiments have closed the empirical E1 gap, but the durable receipt update is still pending at this head. Record the supplied finite calibration evidence, then obtain the bounded exact-successor-head review before merge. Do not repeat the experiment or broaden the product scope.

REQUEST CHANGES

@flyingrobots

Copy link
Copy Markdown
Owner Author

PR 159 exact-successor independent review

Reviewed flyingrobots/keep, branch test/128-retention-release-restore, exact head e768c8fcf769f25c422762cf67fa93384363ea68, tracked tree a03dd192c5d7b1e2f3e3b87e11c07d07d9589200, targeting main at 1325841cbcd27f4c504870728e3ca87d87a2c4cc. Local and live GitHub coordinates agree; checkout remains clean.

This is the bounded independent review of the documentation-only successor to 64bbbf915d87e43fa5c902ddeb0d893f246f9af5. It explicitly adopts the complete original Verification Checklist, its production paths, both-parent merge analysis, invariants, numerical checks, execution evidence and limits. Adoption is justified by independently verifying the exact successor contains no runtime/test/tool/config/CI dependency changes: the delta is eight files under docs/testing-evidence/. No source review is transferred across an uninspected semantic change. This report closes the original E1 evidence obligation and supersedes its REQUEST CHANGES verdict for the exact successor named above.

Findings

No actionable findings. E1 is closed by durable, independently verified calibration receipts. No further calibration, field-by-field expansion, runtime change or test expectation change is required.

Verification Checklist

  • Entire successor diff and history: read all eight changed files and both commits. 6156770aef1cb8436d394e2d096a5fa98608667c adds the consolidated evidence section, three replay patches, three actual RED logs and restored GREEN. e768c8fcf769f25c422762cf67fa93384363ea68 changes patch context to zero-context form, trims one final empty GREEN-log line and documents both normalization/replay details. No new merge is introduced. The intermediate whitespace issue remains in history and is corrected at the reviewed head, without rewriting the earlier experiment or its outcome.
  • Unchanged production paths and merge invariants: git diff --exit-code 64bbbf9 HEAD -- src tests fuzz benchmark xtask Cargo.toml Cargo.lock rust-toolchain.toml .github AGENTS.md is empty. The complete --name-only delta contains only the eight evidence paths. Therefore the original enumeration→recipe→independent model→production preflight/preparation/execution→fenced reader→exact refusal paths, all eight both-parent merge checks, and Retention recovery, crash-matrix evidence, reader fence, and model-based transitions (item 6) #99/Refuse corrupt partial segment seals before recovery discard #171/Prevent post-seal writable stage escape through receipt mapping #146/Preserve platform admission across public catalog publisher routes (T-12.2) #150/Deliver promised public filesystem-stage integration target (T-11.3) #147 preserved invariants remain valid at this exact head. The publication state machine, recovery/effect reporting, sealed authority, platform admission, identity/format/API/dependencies and resource constants are unchanged. The patches are inert evidence artifacts, not included product mutations.
  • Exact source and model independence: compared 159-calibration/original.rs byte-for-byte with current src/adapters/retention/filesystem_retention_snapshot.rs. Reconstructed each committed zero-context patch in memory against those original lines and compared the complete result byte-for-byte with its actual executed retained manifest.rs, liveness.rs or root-generation.rs. Every comparison matched. Each experiment changes exactly one production-reader output projection; test assertions and expected model state remain unchanged. The recorded source head/tree at docs/testing-evidence/retention-release-restore-model.md:23 are the original reviewed calibration coordinates, correctly distinct from this docs-only successor.
  • Replay patches: all three committed patches pass read-only git apply --check --unidiff-zero against the exact successor. Manifest patch replaces filesystem_retention_snapshot.rs:178-180 with None; liveness patch replaces :172 with None; root-generation patch replaces the final selected-root return at :249 with canonical successor re-encoding after ordinary digest/generation selection checks at :239-247. The documented --unidiff-zero replay instruction at evidence :35 matches these artifacts. Independent fresh source/build directories are required; replay is not instructed in the live candidate checkout.
  • Raw-to-committed RED/GREEN fidelity: compared all four committed .txt files against their original raw logs. The only transformations were disclosed replacements of /build/keep159-cal-{name}-source with <calibration-source>, corresponding build target with <calibration-target>, restored build target with <restored-target>, and trimming trailing empty log lines. Every byte otherwise matches, including compilation output, assertion diagnostics, observed/expected values, law names, schedules, timings, outcomes and filtered counts. No RED was manufactured from compiler/setup failure and no failure output was omitted.
  • Manifest-map falsification: committed manifest-red.txt:38-40 establishes compilation and execution in its isolated copy/build; :49-52 reaches retention_model_tests.rs:350 and fails the complete-map assertion with observed {} versus namespace A generation one. :59-61 reports the focused law failed. This calibrates the exact modeled namespace-map observation, not a harness count.
  • Liveness falsification: liveness-red.txt:38-40 compiles/executes; :49-52 reaches retention_model_tests.rs:357 and fails liveness with observed zero versus expected one. :59-61 reports the focused law failed. It does not reuse the manifest mutation or its failing assertion.
  • Selected-root-generation falsification: root-generation-red.txt:38-40 compiles/executes; :49-52 reaches retention_model_tests.rs:366 and fails generation with observed two versus expected one. The mutation re-encodes a valid successor root after normal selected-root admission; decoder/setup errors do not explain this RED. Inspection of verify confirms manifest-map and liveness assertions execute successfully first, so the evidence claim at :29 is supported.
  • Specific failure and reduced witness: the retained traced launcher and 159-calibration/execution.log verify all three process exits are 101, rather than merely observing nonzero commands. Each run names the same focused initial-A law and reports [Initial(A), Initial(A), Initial(A)]. run_sequences verifies immediately after each operation, so the failures occur after the first initial publication; the one-operation reduced witness documented at :31 is correct. The full schedule is a replay coordinate, not a claim all three operations ran before the failure.
  • Restored product GREEN: restored-green.txt:1-2 records original tracked tree ee9b32733e943eb40e364822ae88aa83c5490899; :3-16 executes all seven unmutated model partitions in debug; :18-31 executes them in release. Both record 7 passed, 0 failed. This remains historical execution on the unchanged runtime/test tree before the docs-only addition, explicitly scoped as such. The earlier full copied-Docker validation and exact-64bbbf9 hosted run are not relabeled as execution on the successor commit.
  • Every added numerical/environment claim: three independent copies and three named comparisons match launcher/source/log inventory; exits 101 match the traced shell checks; map/root/liveness generation values 0/1/2 match the assertion diagnostics; history depth three and reduced length one match actual execution order. Seven GREEN laws and 347 filtered correspond to the restored focused selection; one RED law and 353 filtered correspond to the same 354-law library inventory. No count is asserted as correctness evidence. Rust 1.96.0, Linux aarch64 and ext4 scratch agree with the original validation environment and launcher/mount configuration already reviewed. Raw elapsed/build durations are historical measurements, not new latency guarantees. No buffer, size, timeout, rate, profile or performance threshold is altered.
  • Documentation and metadata: evidence :21-35 adds one consolidated calibration section, named failure outcomes, source coordinates, replay commands, reduced witness, restoration, normalization and scope. All seven linked patch/log/GREEN artifacts exist. Prose paragraphs use one physical line. Existing fixture, ordinary test-budget/sandbox, long-history/arbitrary-anchor/concurrency/fault/power-loss limits remain disclosed. The experiments are labeled oracle calibration, not new product bug fixes or power-loss evidence. Owner/deletion criteria and original issue/roadmap scope remain unchanged.
  • Whitespace and static checks: independently executed target-to-successor and empty-tree-to-successor git diff --check; both pass at e768c8f. All read-only patch applicability and in-memory reconstruction checks pass. The parent reports copied-Docker Markdown lint passed; this reviewer separately inspected exact-head hosted Documentation/workflow job 111281282198: it completed SUCCESS on e768c8f, with documentation/workflow and repository-whitespace steps successful. This is executed documentation validation by CI, not runtime evidence from a prose checker.
  • Feedback pagination and reconciliation: independently refreshed GraphQL connections at successor head: 6 global comments, 0 reviews, 0 threads, all hasNextPage=false; no nested thread-comment page remains. Read the original validation comment, bounded finding, usage/environment notices and updated CodeRabbit processing body; the full prior checklist report was already inspected and is adopted above. No new actionable finding is present. CodeRabbit now processes the successor and is PENDING; its status or later approval is not presumed. The earlier E1 comment is reconciled by the exact committed receipts checked above.

Execution and remaining integration gates

Executed by this reviewer: read-only Git/GitHub/source/diff inspection, exact-head and clean-status checks, complete receipt-byte comparisons, in-memory patch-to-executed-mutant reconstruction, link-existence checks, all three read-only patch applicability checks and whitespace checks. Only this scratch report was written. No host Rust test, mutation, commit, push, public comment, merge, configuration change or subagent execution occurred.

Inspected rather than rerun: original copied-Docker full validation, three compiled intended calibration failures, restored debug/release GREEN, and exact-successor hosted documentation validation. The parent's earlier full validation and restored focused sessions returned exit 0; every mutation session returned the expected 101. No additional runtime execution is necessary to review the docs-only delta. The complete prior runtime/evidence limitations remain applicable: finite length-three model, shared fixture/codec foundations, admitted repository fixture rather than platform certification, disclosed ordinary enforcement gaps, no physical power-loss proof, no arbitrary anchor sets or new concurrency/fuzz campaign from this change.

At the final status read, successor run 37149868953 has successful Documentation/workflow integrity and Dependency policy jobs; Rust quality gates and Runtime fuzz smoke remain IN_PROGRESS. This approval is the exact-head source/evidence review gate. The owning agent must separately verify all required hosted checks/protections and immediately recheck head/base/feedback before merge. The approved review does not authorize bypassing protections or attest an integration that has not occurred.

Verdict

The durable receipts faithfully preserve the intended experiments and close E1. The complete prior checklist plus the bounded checks above cover this exact successor, with no actionable findings.

APPROVE

@flyingrobots

Copy link
Copy Markdown
Owner Author

Code Lawyer activity summary

Candidate: e768c8fcf769f25c422762cf67fa93384363ea68, tree a03dd192c5d7b1e2f3e3b87e11c07d07d9589200, targeting main 1325841cbcd27f4c504870728e3ca87d87a2c4cc. Change kind: missing verification and independent-oracle improvement; production behavior unchanged.

Item Severity / source Files and commit Evidence and disposition
Release/restore histories and independent expected state Original #128 obligation src/adapters/retention/retention_model_tests.rs, integrated 64bbbf9 Full copied-Docker validation passes, including debug/release workspace tests, both process-death campaigns, feature/Clippy profiles, docs and fuzz compilation. Historical hidden-empty-root and incorrect release/restore mutations distinguish old/new coverage and falsify the anchor oracle.
E1: missing calibration for generation observations P3 verification gap / independent review Existing evidence document and seven linked receipts; 6156770, corrected e768c8f Three isolated production-reader mutations reach and fail the intended complete manifest-map, liveness and selected-root-generation comparisons. Restored model laws pass debug/release. The independent reviewer matched normalized logs and patches to executed sources and explicitly closes E1. No product defect or RED-on-main claim.
Receipt whitespace Documentation correction / local check Three patch files, restored log and replay instructions, e768c8f Initial formatting failure remains in history. Final whitespace, copied-Docker Markdown lint and all three zero-context patch applicability checks pass. Runtime code and test expectations are unchanged from fully validated 64bbbf9.
Complete review queue Code Lawyer All global comments, review bodies and threads Paginated queue exhausted. Historical quota notices are not approvals. E1 is closed; no inline thread remains. CodeRabbit now APPROVES this exact head with no actionable findings.
Exact-head independent approval GPT-6.1 high reasoning, agy-review protocol Full review, successor APPROVE Full production-path/merge/evidence checklist plus independently checked documentation delta covers the final head.

Scope remains finite length-three histories with fixed fixture anchors. Repository admission isolates post-admission behavior and is not platform certification. No new fault, random, concurrency or physical power-loss campaign is claimed; no production format/API/security/recovery/performance behavior changes. Refusal precision and #99 preservation/cooperating-writer/effect-reporting contracts remain intact. Candidate-derived expectations and count assertions were rejected because they would not independently establish runtime behavior.

Final hosted checks and immediate head/base/protection checks are separate merge gates; this activity summary alone does not declare integration.

Final run 37149868953 is SUCCESS at the exact candidate: Rust quality gates, runtime fuzz smoke, dependency policy and documentation/workflow integrity all pass. CodeRabbit's completed exact-head review approves. All finite obligations above are closed; no failed or skipped required job is represented as green. Normal merge remains subject to the immediate head/base and repository-rule recheck.

@flyingrobots
flyingrobots merged commit d08fafb into main Oct 3, 2026
5 checks passed
@flyingrobots
flyingrobots deleted the test/128-retention-release-restore branch October 3, 2026 20:06
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Integrate retention release/restore model coverage

1 participant