Audit: completed roadmap acceptance and delivery criteria (#131) - #141
Draft
flyingrobots wants to merge 58 commits into
Draft
flyingrobots wants to merge 58 commits into
flyingrobots wants to merge 58 commits into
Conversation
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: true
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The completed-task audit needs criterion-level evidence rather than a blanket green-build verdict. This draft preserves all 83 originally checked task identities, including T-22.1a, and records verdicts for 37 of the 64 tasks that remained checked after the first pass.
Refs #131 and #132. This audit is incomplete: 27 remaining checked tasks and the 19 reopened entries still require full accounting. Do not merge or close #131 until the audit is complete. Originally unchecked tasks are excluded; original checkboxes remain unchanged.
Problem, invariant, and approach
Preserve original acceptance criteria and distinguish observed implementation behavior from delivery on main. The scope manifest cites original and first-pass commits. Each verdict names inspected evidence, obligations, status, and validation limits. Existing correction issues retain ownership; newly discovered gaps receive independently executable issues.
The ledger records both historical findings and subsequent integrations. Outstanding findings include immediate-output acceptance (#71), post-seal writable stage authority (#146), the originally named public filesystem integration target (#147), platform-admission bypass (#150), generated independent catalog-model evidence (#166), runtime scan-cost evidence (#168), and writer-acquisition identity evidence (#169). An implementation in an unmerged PR is not mainline delivery. T-12.1 and T-12.3 have separate criterion verdicts; neither the model enhancement nor the source-text scan test is credited as evidence for the other contract.
Alternatives and failure modes
Reject certification from green workspace tests alone: those cannot establish every acceptance criterion. A shared Cargo target previously reused stale binaries across clones; source-specific reruns replace that evidence. The audit does not infer physical durability from mocks or integrate a fix merely because a PR exists.
Tests and evidence
Change-Kind: documentation-only audit evidence update; no runtime expectation changes.
Source-specific Docker with pinned Rust supplies debug/release evidence recorded beside each verdict. The latest T-12.3 inspection uses main
6051abb25a9fd33ae7ee0de5614514b709a4d82a: catalog generation, codec, head, encoding, location, transition, snapshot, restart and model integration targets pass in both modes. Source inspection shows the model is one hand-written history with a shared production decoding oracle; passing execution does not establish generated agreement. #166 owns that correction, without alleging a production catalog defect.The added initialization verdict and audit index pass pinned markdownlint-cli2 0.23.2 in Docker. These documentation changes require no artificial runtime regression. T-13.1 adds fresh mainline initialization/lock debug-release runs and a surviving ignored-identity-refusal mutation; the cited helper/public suites do not detect the weakened acquisition boundary. No new fuzz campaign, full-workspace clearance, power-loss or performance result is claimed.
Compatibility and operational implications
Documentation only: no implementation, API, format, content identity, benchmark, recovery or security behavior changes. Corrections retain traceability to their executable issue owners. Current checked-task verdict totals are audit coverage, not a product-completion percentage.
Recovery audit in progress
T-13.2 is not counted as a completed verdict. Its current mainline runtime findings are #171 (contradictory partial seal framing accepted for discard, corrected in unmerged #172) and #173 (proven duplicate identities become discardable when an incomplete record/seal tail follows). The latter probe preserves a no-tail duplicate-refusal control and then obtains discard plans for all three incomplete tail classes; it never executes filesystem deletion. The checked-in audit receipt preserves the actual RED. These findings have separate issue ownership and do not reopen the approved v2 retention landing scope.