Skip to content

Complete memory comparisons and accelerate benchmark development - #42

Merged
0thernet merged 30 commits into
mainfrom
devin/memory-superiority-20260906
Sep 9, 2026
Merged

Complete memory comparisons and accelerate benchmark development#42
0thernet merged 30 commits into
mainfrom
devin/memory-superiority-20260906

Conversation

@0thernet

@0thernet 0thernet commented Sep 6, 2026

Copy link
Copy Markdown
Member

Both planned memory comparisons are now complete and independently audited. In the preserved 120-family comparison, Oh fact retrieval scored 81/120 (67.5%), versus 78/120 (65.0%) for raw BM25 windows and 79/120 (65.8%) for BM25 over stored record windows. This is a small observed lead; it does not meet the predefined superiority criterion. In the separate locked reader evaluation, 96 KB of context scored 84/100 versus 78/100 with 24 KB. These are different procedures and samples, not a matched comparison of their absolute scores.

This PR completes the interrupted study, preserves its first responses and spending history, and adds a much faster development loop. It does not claim benchmark saturation, official leaderboard superiority, or an improvement to the complete product memory API. Production memory/search defaults are unchanged.

Completed evidence

Evaluation Result Interpretation
Frozen 120-family / 360-case continuation Oh 81; raw windows 78; record windows 79 +2.5 and +1.67 percentage points; fixed criterion failed
Adverse reader-failure sensitivity Oh lead +1.67 and +0.83 percentage points Criterion still failed
Locked 100-family reader pair GPT-5 mini medium: 96 KB 84; 24 KB 78 Six wins, zero losses, 94 ties on this deterministic reserved set

The frozen primary paired counts are 16 wins / 13 losses / 91 ties against raw windows and 14 / 12 / 94 against record windows. The criterion requires at least +5 percentage points and both one-sided 97.5% finite-pool lower bounds above zero. All 360 cases remain included: 358 model-judged cases and two exact-cap reader failures, one in each baseline. The adverse calculation scores failed baseline cases as correct while holding ordinary judgments fixed. Both criteria fail.

The continuation retains mixed extraction provenance and a reader-failure scoring amendment added after execution began but before correctness inspection. Original incomplete/blocked studies remain preserved. There is no unchanged confirmatory error-control claim; Gateway aliases are not verified model snapshots. The oh-fact adapter uses real SQLite authority and keyword retrieval over records/facts, not the entire semantic or memory-agent API.

The reserved pair used the same medium-reasoning GPT-5 mini profile in both arms. All 200 cases completed with zero reader failures, 198 distinct fresh reader requests and 122 fresh native judge requests. No historical judgments were reused. It took 210.73 seconds and accounted for $0.893499. Both evaluation sets are now closed to further tuning or replacement sampling.

Public evidence: frozen result, reserved reader result, and completed-run takeover guide. Full datasets, prompts, answers and raw responses remain private.

Faster development and implementation

  • Offline development screens now cover 36,820 query/variant rows. The lab loads each dataset once, shares corpus indexes and records context/source identities. Independent workers explored source retention, retrieval granularity and native API projection controls.
  • Lazy authority initialization preserved all 4,000 compared context/turn/metric rows. In one ordered pass, LongMemEval setup/sweep time fell from 22.17 to 1.89 seconds; LoCoMo fell from 2.21 to 1.50 seconds. Frozen defaults stayed eager.
  • Bounded paid execution continuously fills available request slots, reserves spending before dispatch, preserves first responses and authenticates exact cache reuse. The public pinned-config reader command supports the closed profiles and the locked reserved selector.
  • Extraction chunk construction now computes JSON-escaped UTF-8 widths directly. Five-sample synthetic medians show 9.52–14.84× faster chunk construction with 222,487 segmentation comparisons and 26,732 identical chunk/ID/prompt pairs. This is a stage-specific measurement, not a whole-study speedup. Measurement.
  • The old frozen continuation spent 24m32s in its two native launchers. Filesystem timestamps estimate only 1m55s inside request-processing windows, approximately 92.2% outside them. Those mutable timestamps cannot isolate pure provider or import latency. The timing report identifies preparation and closing verification as the next measurement target; the new splitter did not change the frozen runtime.

Negative results remain documented: question-last prompts, human-turn retrieval, minimal-reasoning readers and the parallel three-source rank-fusion candidate were not promoted. The wider-context candidate first scored 79/100 versus 72/100 in development, then 84/100 versus 78/100 on the locked evaluation. The development guide records commands, boundaries, costs and rejected trials.

Custody, budget and verification

Generation stayed pinned to commit c7ec1194ee9ac7a0e3229c2b7197e9223d3e4435, source SHA 458086becf13ea3caeadae28dbd61cc80d5a01a634f9f387507fdcf0e02af88d, freeze SHA aeb1366264f7a9f968188a66117a1cbb96ab6592d653ffa529adf761d810999a.

The original 5,064 attempted jobs were imported without resubmission. Failed pre-native launcher 001 made zero model calls and received a separate acceptance. Native batches 002 and 003 then completed 28 new reader requests and 223 distinct judge requests: 251 calls, all settled. One closure check initially found a numeric process-ID match; a later independent snapshot found no producer and the unchanged checker passed. The transient process identity was not recoverable, so PID reuse is not asserted as proven.

The independent final semantic audit passed with zero model, credential, network, process-dispatch or study-write calls, unchanged inputs and no stderr. It replayed the full matrix, judge ownership, raw responses, imported evidence, ledger prefixes and final inventory. Audit SHA: 97e32feb83272d19060d5fbe7077d00dfad55cb893327401817a6db509071977. Public frozen summary SHA: ca3d3dfc0f655decc3498af12ef197f6c9ee7370f0f9a35cfc48cd7d601f4f66.

The final continuation added $0.141484, bringing cumulative amendment exposure to $25.885579 under the unchanged $40 cap. This includes every later development/reserved ledger and unresolved old reservations. The original-study ledger is a separate authenticated anchor; these amounts are conservative amendment accounting, not a consolidated invoice. No comparison producer remains active. Any future paid plan must account for every ledger exactly once across the descriptor and active cache, including native v6. A new descriptor changes the generic runner's cache namespace; the occupied original cache needs a compatible new namespace/runner or a reviewed migration preserving its evidence.

Validation already completed: focused runtime/transport/scoring/selection tests and strict typing; 39 Python custody tests; 25 TypeScript recovery/history tests; three global-budget tests; independent cross-language review; actual preserved failure acceptance replay; independent raw-response audits for the reader results; splitter parity and focused tests; and the final offline semantic replay. The latest code commit 788544e39c5e11125735ec829b500c9db8c2f110 passed all eight checks in CI run 34358369904, tested merge 95a1bb5fc5e6880756abe67c0c10b5317577b7fa against main 5dad0958328567b6226dcd5753b46ea21a9659d8. Independent review accepted the final result and documentation diff with no remaining material findings. Result commit e47b0e081a0152feccf4b86d1ffdd759f529497f is preserved. Main integration 12ca1998332ac2520c1bc9586a0816ae20c1a326 resolves only the managed-baseline conflict and retains every task-only and main-only file unchanged from its parent; independent integration review passed. All eight checks passed for that exact head in final CI run 34363183970. The tested merge was 17e1e64a4378caf99b2914a700738068ed392028 against freshly verified main 5dad0958328567b6226dcd5753b46ea21a9659d8, using Bun 1.3.14 and Node 24.19.0. Both complete OS checks, Site, CodeQL and Vercel checks passed; source admission is complete. The managed HRA repository baseline is refreshed while preserving unmanaged rules.

Takeover

Branch: devin/memory-superiority-20260906. The completed-run guide identifies all retained evidence and the read-only audit command. Private gateway-v3-implementation-state.json and superiority-live-state.md record the latest machine checkpoint; occupied frozen outputs and first responses must remain intact.

Future work should measure a shared immutable import graph against one instrumented baseline, then evaluate actual memory-API improvements on development groups and a newly fixed independent sample. Preserve fresh source/byte/inventory/ledger checks and complete prompt parity. The existing locked sets are complete, and neither current result establishes saturation. No additional paid experiment is active or implicitly authorized by an old descriptor.

Final delivery

Merged on 2026-09-09 at 14:27:52 UTC as e7a7a1cd946c6abd82fb77321030565d67b499a6. The merged tree f94221939a9d076969c19d1786a5910edb2ca127 exactly equals the validated integration head. The repository requires squash merges; original source history remains available through refs/pull/42/head at 12ca1998332ac2520c1bc9586a0816ae20c1a326, including frozen generation commit c7ec1194ee9ac7a0e3229c2b7197e9223d3e4435. Post-merge main CI 34363772037 passed both complete OS checks and Site on the exact merged SHA. No new release or production memory-default promotion was made. The benchmark run and repository delivery are complete; subsequent accuracy or import-performance experiments are separate future work.

@vercel

vercel Bot commented Sep 6, 2026

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
oh-computer Ready Ready Preview Sep 9, 2026 2:23pm UTC

Request Review

@0thernet

0thernet commented Sep 7, 2026

Copy link
Copy Markdown
Member Author

Janitor status (2026-09-07): this draft is intentionally not merge-ready. The public protocol/code, CI, CodeQL, site, and preview are green, but no active paid-benchmark worker or the private full-protocol/ledger/result custody needed to resume and reconcile all 360 judgments was found on this host. Preserve the frozen public branch and do not merge or restart from a fresh ledger. Completion requires resuming from the original private custody, landing the scored results and reconciled costs, then rerunning the final review gate.

@0thernet 0thernet changed the title Freeze a 120-family memory comparison with faithful BM25 controls Run the frozen memory comparison through Claude Code Sep 7, 2026
@0thernet 0thernet changed the title Run the frozen memory comparison through Claude Code Add Claude subscription benchmarks and offline memory stress tests Sep 7, 2026
@0thernet 0thernet changed the title Add Claude subscription benchmarks and offline memory stress tests Add Claude subscription benchmarks with preserved first responses Sep 7, 2026
@0thernet 0thernet changed the title Add Claude subscription benchmarks with preserved first responses Add preserved-response memory studies and budgeted Gateway execution Sep 7, 2026
@0thernet 0thernet changed the title Add preserved-response memory studies and budgeted Gateway execution Preserve benchmark responses across Claude and Gateway continuations Sep 7, 2026
Preserve the frozen generation runtime while porting the final auditor, supervisor and closure helpers into reviewed source. Bind portable context to the existing authority and first producer command, add Python CI coverage, and document safe continuation and final evidence gates.
Treat EPERM probes as possibly live, keep denied cleanup signals within bounded group polling, and retain cleanup-incomplete until disappearance is proven. Add deterministic permission-denial regressions and document that the active benchmark keeps its accepted supervisor.
@0thernet 0thernet changed the title Accelerate memory benchmark development and preserve comparison evidence Complete memory comparisons and accelerate benchmark development Sep 9, 2026
@0thernet
0thernet marked this pull request as ready for review September 9, 2026 14:27
@0thernet
0thernet merged commit e7a7a1c into main Sep 9, 2026
8 checks passed
@0thernet
0thernet deleted the devin/memory-superiority-20260906 branch September 9, 2026 14:27
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant