Complete memory comparisons and accelerate benchmark development - #42
Merged
Conversation
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
Member
Author
|
Janitor status (2026-09-07): this draft is intentionally not merge-ready. The public protocol/code, CI, CodeQL, site, and preview are green, but no active paid-benchmark worker or the private full-protocol/ledger/result custody needed to resume and reconcile all 360 judgments was found on this host. Preserve the frozen public branch and do not merge or restart from a fresh ledger. Completion requires resuming from the original private custody, landing the scored results and reconciled costs, then rerunning the final review gate. |
Preserve the frozen generation runtime while porting the final auditor, supervisor and closure helpers into reviewed source. Bind portable context to the existing authority and first producer command, add Python CI coverage, and document safe continuation and final evidence gates.
Treat EPERM probes as possibly live, keep denied cleanup signals within bounded group polling, and retain cleanup-incomplete until disappearance is proven. Add deterministic permission-denial regressions and document that the active benchmark keeps its accepted supervisor.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Both planned memory comparisons are now complete and independently audited. In the preserved 120-family comparison, Oh fact retrieval scored 81/120 (67.5%), versus 78/120 (65.0%) for raw BM25 windows and 79/120 (65.8%) for BM25 over stored record windows. This is a small observed lead; it does not meet the predefined superiority criterion. In the separate locked reader evaluation, 96 KB of context scored 84/100 versus 78/100 with 24 KB. These are different procedures and samples, not a matched comparison of their absolute scores.
This PR completes the interrupted study, preserves its first responses and spending history, and adds a much faster development loop. It does not claim benchmark saturation, official leaderboard superiority, or an improvement to the complete product memory API. Production memory/search defaults are unchanged.
Completed evidence
The frozen primary paired counts are 16 wins / 13 losses / 91 ties against raw windows and 14 / 12 / 94 against record windows. The criterion requires at least +5 percentage points and both one-sided 97.5% finite-pool lower bounds above zero. All 360 cases remain included: 358 model-judged cases and two exact-cap reader failures, one in each baseline. The adverse calculation scores failed baseline cases as correct while holding ordinary judgments fixed. Both criteria fail.
The continuation retains mixed extraction provenance and a reader-failure scoring amendment added after execution began but before correctness inspection. Original incomplete/blocked studies remain preserved. There is no unchanged confirmatory error-control claim; Gateway aliases are not verified model snapshots. The
oh-factadapter uses real SQLite authority and keyword retrieval over records/facts, not the entire semantic or memory-agent API.The reserved pair used the same medium-reasoning GPT-5 mini profile in both arms. All 200 cases completed with zero reader failures, 198 distinct fresh reader requests and 122 fresh native judge requests. No historical judgments were reused. It took 210.73 seconds and accounted for $0.893499. Both evaluation sets are now closed to further tuning or replacement sampling.
Public evidence: frozen result, reserved reader result, and completed-run takeover guide. Full datasets, prompts, answers and raw responses remain private.
Faster development and implementation
Negative results remain documented: question-last prompts, human-turn retrieval, minimal-reasoning readers and the parallel three-source rank-fusion candidate were not promoted. The wider-context candidate first scored 79/100 versus 72/100 in development, then 84/100 versus 78/100 on the locked evaluation. The development guide records commands, boundaries, costs and rejected trials.
Custody, budget and verification
Generation stayed pinned to commit
c7ec1194ee9ac7a0e3229c2b7197e9223d3e4435, source SHA458086becf13ea3caeadae28dbd61cc80d5a01a634f9f387507fdcf0e02af88d, freeze SHAaeb1366264f7a9f968188a66117a1cbb96ab6592d653ffa529adf761d810999a.The original 5,064 attempted jobs were imported without resubmission. Failed pre-native launcher 001 made zero model calls and received a separate acceptance. Native batches 002 and 003 then completed 28 new reader requests and 223 distinct judge requests: 251 calls, all settled. One closure check initially found a numeric process-ID match; a later independent snapshot found no producer and the unchanged checker passed. The transient process identity was not recoverable, so PID reuse is not asserted as proven.
The independent final semantic audit passed with zero model, credential, network, process-dispatch or study-write calls, unchanged inputs and no stderr. It replayed the full matrix, judge ownership, raw responses, imported evidence, ledger prefixes and final inventory. Audit SHA:
97e32feb83272d19060d5fbe7077d00dfad55cb893327401817a6db509071977. Public frozen summary SHA:ca3d3dfc0f655decc3498af12ef197f6c9ee7370f0f9a35cfc48cd7d601f4f66.The final continuation added $0.141484, bringing cumulative amendment exposure to $25.885579 under the unchanged $40 cap. This includes every later development/reserved ledger and unresolved old reservations. The original-study ledger is a separate authenticated anchor; these amounts are conservative amendment accounting, not a consolidated invoice. No comparison producer remains active. Any future paid plan must account for every ledger exactly once across the descriptor and active cache, including native v6. A new descriptor changes the generic runner's cache namespace; the occupied original cache needs a compatible new namespace/runner or a reviewed migration preserving its evidence.
Validation already completed: focused runtime/transport/scoring/selection tests and strict typing; 39 Python custody tests; 25 TypeScript recovery/history tests; three global-budget tests; independent cross-language review; actual preserved failure acceptance replay; independent raw-response audits for the reader results; splitter parity and focused tests; and the final offline semantic replay. The latest code commit
788544e39c5e11125735ec829b500c9db8c2f110passed all eight checks in CI run 34358369904, tested merge95a1bb5fc5e6880756abe67c0c10b5317577b7faagainst main5dad0958328567b6226dcd5753b46ea21a9659d8. Independent review accepted the final result and documentation diff with no remaining material findings. Result commite47b0e081a0152feccf4b86d1ffdd759f529497fis preserved. Main integration12ca1998332ac2520c1bc9586a0816ae20c1a326resolves only the managed-baseline conflict and retains every task-only and main-only file unchanged from its parent; independent integration review passed. All eight checks passed for that exact head in final CI run 34363183970. The tested merge was17e1e64a4378caf99b2914a700738068ed392028against freshly verified main5dad0958328567b6226dcd5753b46ea21a9659d8, using Bun 1.3.14 and Node 24.19.0. Both complete OS checks, Site, CodeQL and Vercel checks passed; source admission is complete. The managed HRA repository baseline is refreshed while preserving unmanaged rules.Takeover
Branch:
devin/memory-superiority-20260906. The completed-run guide identifies all retained evidence and the read-only audit command. Privategateway-v3-implementation-state.jsonandsuperiority-live-state.mdrecord the latest machine checkpoint; occupied frozen outputs and first responses must remain intact.Future work should measure a shared immutable import graph against one instrumented baseline, then evaluate actual memory-API improvements on development groups and a newly fixed independent sample. Preserve fresh source/byte/inventory/ledger checks and complete prompt parity. The existing locked sets are complete, and neither current result establishes saturation. No additional paid experiment is active or implicitly authorized by an old descriptor.
Final delivery
Merged on 2026-09-09 at 14:27:52 UTC as
e7a7a1cd946c6abd82fb77321030565d67b499a6. The merged treef94221939a9d076969c19d1786a5910edb2ca127exactly equals the validated integration head. The repository requires squash merges; original source history remains available throughrefs/pull/42/headat12ca1998332ac2520c1bc9586a0816ae20c1a326, including frozen generation commitc7ec1194ee9ac7a0e3229c2b7197e9223d3e4435. Post-merge main CI 34363772037 passed both complete OS checks and Site on the exact merged SHA. No new release or production memory-default promotion was made. The benchmark run and repository delivery are complete; subsequent accuracy or import-performance experiments are separate future work.