You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Close the benchmark program for v0.6.0: the OVHC-AGENCY Graph500 ladder harness, its parity with the retired legacy orchestration, and published, unaudited scale scorecards for all four GDC suites that prove how many edges GraphForge loads and queries correctly. This is M10's only open issue. It blocks readiness #1096 and is blocked by #900.
Status — 2026-10-01
Two outcomes remain:
GDC scorecards (see GDC scorecards). Not started, and not blocked by M5. This is the M10 critical path.
Graph500-compliant input. Vertex scrambling ported from graph500-3.0.0 and checked against the pinned C reference (fix(bench): generate compliant scrambled Graph500 input #1111 / fix(bench): generate compliant Graph500 input #1117, dd93aec8a), before the current S18–S22 evidence on fa6447cc. Published identity: generator source and executable digests in the host run result (schemas/progressive-host-run-result.json), edge factor 16 and seed 13907095936298285200 in the profile (runners/certify/src/lib.rs), and scale/live_edges per rung, with live_edges == 16 · 2^scale enforced because one tuple stores one edge.
Work-root teardown is implemented and tested. reclaim_rung_workspace / inventory_work_root, schema host-work-root-inventory.json; test_progressive_host_run, test_native_ladder_bundle. The final-run proof is part of the ladder item below.
make -C benchmarks ingest-ladder-bundle SOURCE=<#900 output dir> validates the native bundle.
make -C benchmarks parity-gate reports full_ladder_evidence_complete: true. That requires a canonical S18→S26 prefix, every rung compared against legacy as match or a declared accepted difference, and a teardown inventory with empty: true.
Goal: prove the edges. For each of the four GDC suites, publish an unaudited scorecard showing how many nodes and edges GraphForge loaded through the ordinary public product surface, and how fast and how correctly it answered that workload's queries. Each suite climbs its own scale ladder until the first typed failure, like the Graph500 ladder. The card headlines the largest passing rung and names the failure that stopped the next one. All runs happen on OVHC-AGENCY under local-linux-cgroups-v2.
Card
GraphForge <version> (<commit>) — GDC <suite>, <dataset/SF>, unaudited
These are not LDBC Benchmark Results. <one-line variance from the spec>
Hardware: OVHC-AGENCY — 16 cores, 125 GiB, NVMe ext4, Ubuntu 26.04
Graph: <nodes> nodes / <edges> edges loaded (LDBC published: <nodes> / <edges>; exact match)
Load time: <BenchExec wall for gf import-session> (CSV→Parquet conversion reported separately)
On disk: <published project bytes>
Coverage: <supported>/<total> queries; refused: <ids> (<typed cause>)
Throughput: <read-only ops/s or queries/hour, single client>
Latency: p50 <x> p95 <y> (per query in the evidence; all-query on the card)
Peak RAM: <BenchExec peak process RSS>, load and query phases separately
Correctness: <matched>/<checked> supported results match <reference source>
Next rung: <dataset/SF> — <passed | typed failure | not admitted (reason)>
Graphalytics reports its own metrics instead of p50/p95: per-algorithm load time Tl, processing time Tp (mean of 3 runs), makespan and EVPS.
Ladders
Each rung gets the Graph500 per-rung envelope: four-hour wall, 4 GiB process memory limit, disk admission with the declared reserve, and first-failure stop. A rung that exceeds the memory limit fails typed. Raising the limit is a maintainer decision recorded here, not a harness default. Graphalytics algorithms also use Graphalytics' per-class timeouts (S: 15 min). v0.6.0 floor:
The SF1-class rung passes: SNB BI, SNB Interactive and FinBench at SF1; Graphalytics on wiki-Talk and cit-Patents.
The SF10-class rung is attempted: SF10, or graph500-22 for Graphalytics. Its result, pass or typed failure, is on the card.
Later rungs run if admitted.
The rows below are LDBC-published node/edge counts. The loaded graph must reconcile to them exactly.
Reference outputs shipped in each dataset archive. Matching rules: exact (BFS, CDLP), equivalence (WCC), 1e-4 epsilon (PR, LCC, SSSP).
SNB BI
SF1 → SF3 → SF10 → SF30 → SF100
17.2M → … → 170.3M → … → 1.70B
Umbra validation output, published for SF10 only (bi-pre-audit/output-sf10-validation-umbra). Uses the published parameters for SF1–SF30000. SF1 cards state "not reference-checked" unless a reference is generated.
SNB Interactive (v1)
SF1 → SF3 → SF10 → SF30
23.0M → … → 231.4M → …
Neo4j-produced v1 validation parameters, SF0.1–SF10. v2 publishes no validation set and is the version without audits.
FinBench Transaction
SF1 → SF3 → SF10
6.1M → … → 51.9M
Unconfirmed. The driver has CREATE_VALIDATION mode, and the spec names SF1 as the validation scale, but no published reference file was found. Slice 1 confirms one. Fallback: generate references once with the reference implementation in a container on the host, and pin their digest.
Every dataset archive, parameter set and reference archive is downloaded from datasets.ldbcouncil.org and pinned by SHA-256 in benchmarks/profiles/gdc/*-identity.json.
LDBC publishes no checksum sidecars, so pins are recorded at first download and a later mismatch is a typed failure.
Downloads are cached on the host work root, outside the repo.
LDBC CSV is converted to Parquet by a harness step; Graphalytics uses its published Parquet.
Ordinary product path.
Data loads with gf import-session, and queries run through the public Cypher API or analyst verbs.
Loaded node and edge counts per type, read back from GraphForge after reopen, equal the LDBC-published counts for that rung.
For FinBench, that is the snapshot count excluding the 3% held back as incremental updates.
A mismatch is a typed failure.
Correctness.
Every supported query or algorithm is checked against the reference in the table, using that workload's matching rules.
Refused queries fail closed with the existing typed causes and count against coverage, never as correct.
A wrong answer fails the rung.
Measurement.
BenchExec owns load wall time, whole-run CPU and peak process RSS.
Per-operation latency comes from one declared driver clock around execute plus full result materialization.
That clock is added to docs/development/benchmarking.md as the per-operation latency authority, and gdc_measurement_policy.py enforces it. This amends the 2026-09-19 policy. The old inventory and policy script were removed by ci: cut the PR gate to lint, build, and test (#1699) #1701.
Each query variant gets one warm-up pass, excluded, then a measured pass over the published parameter bindings.
Throughput is single-client and read-only, and is labelled as not the LDBC power/throughput score.
Teardown. After each rung, the work-root inventory is empty: true. The dataset cache may be retained and is declared separately.
Work order
One PR per slice. File each slice as a native sub-issue of this tracker when it is dispatched, within the WIP limit.
Shared harness.
Acquisition and pinning; CSV→Parquet conversion; a GDC rung runner reusing the progressive admission, envelope and teardown code.
Per-operation latency authority and the policy amendment; card renderer and evidence schema.
Confirm the FinBench reference source.
Graphalytics scorecard. Published references and Parquet make this the cheapest end-to-end proof of slice 1.
SNB BI scorecard.
BI1–BI20 have never run beyond BI2 through GraphForge. Map each one or fail it closed.
The Neo4j reference queries use APOC/GDS in BI10, BI15, BI19 and BI20.
SNB Interactive (v1) scorecard.
FinBench Transaction scorecard.
Host runs and publication at one frozen commit.
Standing decisions
Graph500 claim (2026-09-04). Graph500-compliant generated data, clearly labeled. No official Graph500 BFS/SSSP result or TEPS score is claimed. Ingest, persistence, query, interchange and recovery are product lifecycle tests.
Measurement authority (2026-09-19). BenchExec owns process-tree timing and resources, and Divan owns in-process benchmarks. Shared phase timings are diagnostic only. Product semantic counters keep their own named authority. CodSpeed remains permitted.
Fly. The disposable Fly adapter stays as disabled offline tooling (benchmarks/fly/disabled.json), not a close path.
Evidence location. Run output attaches to the issue that produced it. The repository keeps method and content digests only (AGENTS.md § Evidence).
Non-goals
CodSpeed replacement, benchmark-only product hooks, developer-laptop or overlay-VM scale claims, persistent cloud provider infrastructure as the SUT, audited LDBC/GDC results or the LDBC power/throughput scores (they need the update streams GraphForge refuses), implementing currently refused query semantics (separate product issues if wanted), and limits generalized beyond the declared OVHC-AGENCY envelope.
Related
#17 (user-facing LDBC loader, after v0.6.0; may adopt this harness's mapping), #735, #745, #900, #959, #1096, #1208 (claim-to-evidence table consumes the cards), #1387, #1388.
Tracker purpose
Close the benchmark program for v0.6.0: the OVHC-AGENCY Graph500 ladder harness, its parity with the retired legacy orchestration, and published, unaudited scale scorecards for all four GDC suites that prove how many edges GraphForge loads and queries correctly. This is M10's only open issue. It blocks readiness #1096 and is blocked by #900.
Status — 2026-10-01
Two outcomes remain:
Everything else is delivered and verified at
maina5902a01a. That includes the tiny GDC correctness suites the scorecards build on.Verification at
a5902a01a:Closure standard
All native sub-issues are closed or explicitly superseded. build(bench): scaffold an isolated external benchmark workspace #953–docs(bench): inventory GDC SPB without adding an incompatible runner #965, ci(bench): remove unused provider image build from required Rust PR checks #1113 closed; epic(bench): standardize BenchExec and Divan measurement and retire custom benchmark harnesses #959 closed the BenchExec/Divan standardization and bounded parity.
Profiles, controller, BenchExec authority and local-host admission are complete.
benchmarks/profiles/local-linux-cgroups-v2.json; controllerbenchmarks/harness/graphforge_bench/progressive_host_run.py(ladder18,19,20,22,24,25,26, first-failure stop, work-root free-capacity admission with reserve, adjacent-rung projection); BenchExecbenchexec_authority.py; runbookbenchmarks/README.md§ OVHC-AGENCY host ladder.Graph500-compliant input. Vertex scrambling ported from graph500-3.0.0 and checked against the pinned C reference (fix(bench): generate compliant scrambled Graph500 input #1111 / fix(bench): generate compliant Graph500 input #1117,
dd93aec8a), before the current S18–S22 evidence onfa6447cc. Published identity: generator source and executable digests in the host run result (schemas/progressive-host-run-result.json), edge factor 16 and seed13907095936298285200in the profile (runners/certify/src/lib.rs), andscale/live_edgesper rung, withlive_edges == 16 · 2^scaleenforced because one tuple stores one edge.Work-root teardown is implemented and tested.
reclaim_rung_workspace/inventory_work_root, schemahost-work-root-inventory.json;test_progressive_host_run,test_native_ladder_bundle. The final-run proof is part of the ladder item below.Legacy orchestration retired after shadow parity. refactor(bench): retire legacy Graph500 scale orchestration (#959) #1061 retired it; tiny shadow parity and the accepted-difference list are in
benchmarks/scale-parity-index.mdandfixtures/parity/accepted-differences.json.GDC suites pin upstream identities and pass bounded live correctness fixtures (test(bench): add a separate GDC Graphalytics benchmark suite #961–docs(bench): inventory GDC SPB without adding an incompatible runner #965). This is the foundation for the scorecards below.
GDC scorecards published. Every acceptance criterion under GDC scorecards is met, and the four cards are attached here.
Completed-ladder evidence. Once test(scale): complete the S18-S26 lifecycle ladder on OVHC-AGENCY #900 accepts its terminal S26 rung:
make -C benchmarks ingest-ladder-bundle SOURCE=<#900 output dir>validates the native bundle.make -C benchmarks parity-gatereportsfull_ladder_evidence_complete: true. That requires a canonical S18→S26 prefix, every rung compared against legacy asmatchor a declared accepted difference, and a teardown inventory withempty: true.Any difference not already in
accepted-differences.jsonis added there with its reason in the same PR, or the gate stays red. The S20 and S26 resource and admission requirements are test(scale): complete the S18-S26 lifecycle ladder on OVHC-AGENCY #900/test(scale): certify the S26 billion-edge round trip on OVHC-AGENCY #745's acceptance criteria. This tracker consumes that single run and does not restate or rerun them.GDC scorecards (maintainer decision 2026-10-01)
Goal: prove the edges. For each of the four GDC suites, publish an unaudited scorecard showing how many nodes and edges GraphForge loaded through the ordinary public product surface, and how fast and how correctly it answered that workload's queries. Each suite climbs its own scale ladder until the first typed failure, like the Graph500 ladder. The card headlines the largest passing rung and names the failure that stopped the next one. All runs happen on OVHC-AGENCY under
local-linux-cgroups-v2.Card
Graphalytics reports its own metrics instead of p50/p95: per-algorithm load time
Tl, processing timeTp(mean of 3 runs), makespan and EVPS.Ladders
Each rung gets the Graph500 per-rung envelope: four-hour wall, 4 GiB process memory limit, disk admission with the declared reserve, and first-failure stop. A rung that exceeds the memory limit fails typed. Raising the limit is a maintainer decision recorded here, not a harness default. Graphalytics algorithms also use Graphalytics' per-class timeouts (S: 15 min). v0.6.0 floor:
The rows below are LDBC-published node/edge counts. The loaded graph must reconcile to them exactly.
bi-pre-audit/output-sf10-validation-umbra). Uses the published parameters for SF1–SF30000. SF1 cards state "not reference-checked" unless a reference is generated.Sources: SNB datasets, BI pre-generated sets, BI entity counts, FinBench datasets, FinBench entity counts, Graphalytics datasets, fair-use policy.
Acceptance criteria
datasets.ldbcouncil.organd pinned by SHA-256 inbenchmarks/profiles/gdc/*-identity.json.gf import-session, and queries run through the public Cypher API or analyst verbs.benchmarks/; feat(gf-datasets): generic loader + registry infrastructure #17 may adopt it later.docs/development/benchmarking.mdas the per-operation latency authority, andgdc_measurement_policy.pyenforces it. This amends the 2026-09-19 policy. The old inventory and policy script were removed by ci: cut the PR gate to lint, build, and test (#1699) #1701.benchmarks/gdc-suite-index.mddocuments the method and pins.empty: true. The dataset cache may be retained and is declared separately.Work order
One PR per slice. File each slice as a native sub-issue of this tracker when it is dispatched, within the WIP limit.
Standing decisions
benchmarks/fly/disabled.json), not a close path.Non-goals
CodSpeed replacement, benchmark-only product hooks, developer-laptop or overlay-VM scale claims, persistent cloud provider infrastructure as the SUT, audited LDBC/GDC results or the LDBC power/throughput scores (they need the update streams GraphForge refuses), implementing currently refused query semantics (separate product issues if wanted), and limits generalized beyond the declared OVHC-AGENCY envelope.
Related
#17 (user-facing LDBC loader, after v0.6.0; may adopt this harness's mapping), #735, #745, #900, #959, #1096, #1208 (claim-to-evidence table consumes the cards), #1387, #1388.