Skip to content

epic(bench): externalize scale certification and add GDC benchmark suites #952

Description

@DecisionNerd

Tracker purpose

Close the benchmark program for v0.6.0: the OVHC-AGENCY Graph500 ladder harness, its parity with the retired legacy orchestration, and published, unaudited scale scorecards for all four GDC suites that prove how many edges GraphForge loads and queries correctly. This is M10's only open issue. It blocks readiness #1096 and is blocked by #900.

Status — 2026-10-01

Two outcomes remain:

  1. GDC scorecards (see GDC scorecards). Not started, and not blocked by M5. This is the M10 critical path.
  2. Ingesting test(scale): complete the S18-S26 lifecycle ladder on OVHC-AGENCY #900's completed S18–S26 ladder. This waits on the M5 ingest and query redesign (epic(storage): scale complete ingest across cores and reach 1M edges/s #1387, epic(query): bounded queries should cost what their results cost, not what the graph costs #1388), because S25/S26 do not fit the four-hour rung envelope on today's architecture (see test(scale): complete the S18-S26 lifecycle ladder on OVHC-AGENCY #900, 2026-09-17).

Everything else is delivered and verified at main a5902a01a. That includes the tiny GDC correctness suites the scorecards build on.

Verification at a5902a01a:

cd benchmarks && PYTHONPATH=harness uv run --locked python -m unittest \
  tests.test_benchexec_authority tests.test_gdc_contracts tests.test_gdc_finbench_transaction \
  tests.test_gdc_graphalytics tests.test_gdc_measurement_policy tests.test_gdc_snb_bi \
  tests.test_gdc_snb_interactive tests.test_gdc_spb_inventory tests.test_native_ladder_bundle \
  tests.test_parity_gate tests.test_progressive_host_run tests.test_scale_parity
# Ran 143 tests in 387.394s — OK

make -C benchmarks parity-gate
# structural_retirement_ready: true, prefix_parity_ready: true,
# full_ladder_evidence_complete: false (checked-in sample bundle is S18/S19 only)

Closure standard

GDC scorecards (maintainer decision 2026-10-01)

Goal: prove the edges. For each of the four GDC suites, publish an unaudited scorecard showing how many nodes and edges GraphForge loaded through the ordinary public product surface, and how fast and how correctly it answered that workload's queries. Each suite climbs its own scale ladder until the first typed failure, like the Graph500 ladder. The card headlines the largest passing rung and names the failure that stopped the next one. All runs happen on OVHC-AGENCY under local-linux-cgroups-v2.

Card

GraphForge <version> (<commit>) — GDC <suite>, <dataset/SF>, unaudited
These are not LDBC Benchmark Results. <one-line variance from the spec>
Hardware:     OVHC-AGENCY — 16 cores, 125 GiB, NVMe ext4, Ubuntu 26.04
Graph:        <nodes> nodes / <edges> edges loaded (LDBC published: <nodes> / <edges>; exact match)
Load time:    <BenchExec wall for gf import-session>   (CSV→Parquet conversion reported separately)
On disk:      <published project bytes>
Coverage:     <supported>/<total> queries; refused: <ids> (<typed cause>)
Throughput:   <read-only ops/s or queries/hour, single client>
Latency:      p50 <x>  p95 <y>  (per query in the evidence; all-query on the card)
Peak RAM:     <BenchExec peak process RSS>, load and query phases separately
Correctness:  <matched>/<checked> supported results match <reference source>
Next rung:    <dataset/SF> — <passed | typed failure | not admitted (reason)>

Graphalytics reports its own metrics instead of p50/p95: per-algorithm load time Tl, processing time Tp (mean of 3 runs), makespan and EVPS.

Ladders

Each rung gets the Graph500 per-rung envelope: four-hour wall, 4 GiB process memory limit, disk admission with the declared reserve, and first-failure stop. A rung that exceeds the memory limit fails typed. Raising the limit is a maintainer decision recorded here, not a harness default. Graphalytics algorithms also use Graphalytics' per-class timeouts (S: 15 min). v0.6.0 floor:

  • The SF1-class rung passes: SNB BI, SNB Interactive and FinBench at SF1; Graphalytics on wiki-Talk and cit-Patents.
  • The SF10-class rung is attempted: SF10, or graph500-22 for Graphalytics. Its result, pass or typed failure, is on the card.
  • Later rungs run if admitted.

The rows below are LDBC-published node/edge counts. The loaded graph must reconcile to them exactly.

Suite Ladder Edges at each rung Correctness reference
Graphalytics wiki-Talk (2XS) → cit-Patents (XS) → datagen-7_5-fb (S, adds SSSP) → graph500-22 (S) → M-class onward 5.02M → 16.5M → 34.2M → 64.2M → … Reference outputs shipped in each dataset archive. Matching rules: exact (BFS, CDLP), equivalence (WCC), 1e-4 epsilon (PR, LCC, SSSP).
SNB BI SF1 → SF3 → SF10 → SF30 → SF100 17.2M → … → 170.3M → … → 1.70B Umbra validation output, published for SF10 only (bi-pre-audit/output-sf10-validation-umbra). Uses the published parameters for SF1–SF30000. SF1 cards state "not reference-checked" unless a reference is generated.
SNB Interactive (v1) SF1 → SF3 → SF10 → SF30 23.0M → … → 231.4M → … Neo4j-produced v1 validation parameters, SF0.1–SF10. v2 publishes no validation set and is the version without audits.
FinBench Transaction SF1 → SF3 → SF10 6.1M → … → 51.9M Unconfirmed. The driver has CREATE_VALIDATION mode, and the spec names SF1 as the validation scale, but no published reference file was found. Slice 1 confirms one. Fallback: generate references once with the reference implementation in a container on the host, and pin their digest.

Sources: SNB datasets, BI pre-generated sets, BI entity counts, FinBench datasets, FinBench entity counts, Graphalytics datasets, fair-use policy.

Acceptance criteria

  • Acquisition.
    • Every dataset archive, parameter set and reference archive is downloaded from datasets.ldbcouncil.org and pinned by SHA-256 in benchmarks/profiles/gdc/*-identity.json.
    • LDBC publishes no checksum sidecars, so pins are recorded at first download and a later mismatch is a typed failure.
    • Downloads are cached on the host work root, outside the repo.
    • LDBC CSV is converted to Parquet by a harness step; Graphalytics uses its published Parquet.
  • Ordinary product path.
  • Edges proven.
    • Loaded node and edge counts per type, read back from GraphForge after reopen, equal the LDBC-published counts for that rung.
    • For FinBench, that is the snapshot count excluding the 3% held back as incremental updates.
    • A mismatch is a typed failure.
  • Correctness.
    • Every supported query or algorithm is checked against the reference in the table, using that workload's matching rules.
    • Refused queries fail closed with the existing typed causes and count against coverage, never as correct.
    • A wrong answer fails the rung.
  • Measurement.
    • BenchExec owns load wall time, whole-run CPU and peak process RSS.
    • Per-operation latency comes from one declared driver clock around execute plus full result materialization.
    • That clock is added to docs/development/benchmarking.md as the per-operation latency authority, and gdc_measurement_policy.py enforces it. This amends the 2026-09-19 policy. The old inventory and policy script were removed by ci: cut the PR gate to lint, build, and test (#1699) #1701.
    • Each query variant gets one warm-up pass, excluded, then a measured pass over the published parameter bindings.
    • Throughput is single-client and read-only, and is labelled as not the LDBC power/throughput score.
  • Ladders run. Each suite reaches at least the v0.6.0 floor above on OVHC-AGENCY at one frozen merged commit. GDC runs serialize with test(scale): complete the S18-S26 lifecycle ladder on OVHC-AGENCY #900 ladder runs on the host.
  • Cards published.
    • Four cards are attached here as one comment: one per suite, in the format above, with evidence JSON and executable/dataset digests.
    • Labels follow the LDBC fair-use policy: "These are not LDBC Benchmark Results", variances from the spec, and CC-BY 4.0 attribution.
    • No raw output is committed to the repo.
    • benchmarks/gdc-suite-index.md documents the method and pins.
    • docs(release): prepare audience positioning and evidence-backed v0.6.0 launch material #1208 links the cards from its claim-to-evidence table.
  • Teardown. After each rung, the work-root inventory is empty: true. The dataset cache may be retained and is declared separately.

Work order

One PR per slice. File each slice as a native sub-issue of this tracker when it is dispatched, within the WIP limit.

  1. Shared harness.
    • Acquisition and pinning; CSV→Parquet conversion; a GDC rung runner reusing the progressive admission, envelope and teardown code.
    • Per-operation latency authority and the policy amendment; card renderer and evidence schema.
    • Confirm the FinBench reference source.
  2. Graphalytics scorecard. Published references and Parquet make this the cheapest end-to-end proof of slice 1.
  3. SNB BI scorecard.
    • BI1–BI20 have never run beyond BI2 through GraphForge. Map each one or fail it closed.
    • The Neo4j reference queries use APOC/GDS in BI10, BI15, BI19 and BI20.
  4. SNB Interactive (v1) scorecard.
  5. FinBench Transaction scorecard.
  6. Host runs and publication at one frozen commit.

Standing decisions

Non-goals

CodSpeed replacement, benchmark-only product hooks, developer-laptop or overlay-VM scale claims, persistent cloud provider infrastructure as the SUT, audited LDBC/GDC results or the LDBC power/throughput scores (they need the update streams GraphForge refuses), implementing currently refused query semantics (separate product issues if wanted), and limits generalized beyond the declared OVHC-AGENCY envelope.

Related

#17 (user-facing LDBC loader, after v0.6.0; may adopt this harness's mapping), #735, #745, #900, #959, #1096, #1208 (claim-to-evidence table consumes the cards), #1387, #1388.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    testingTest coverage and testing infrastructure

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions