Static site for posttrainbench.com: the leaderboard, the blog and the
trace viewer. There is no build step; GitHub Pages serves the repo as-is (custom domain in CNAME).
python3 -m http.server 8000 # then open http://localhost:8000| Path | What it is |
|---|---|
index.html, script.js, styles.css |
The main page: leaderboard, charts, scoring, setup, observations |
config.js |
Agent display names and metadata, and which agents appear in each chart |
data.js |
Loads a score bundle and computes the leaderboard from it |
scores-v1.2.{js,json} |
Current results (v1.2), generated from data/v1.2/ |
scores.{js,json} |
v1.1 results, generated from data/ |
scores-v1.js |
Archived v1 results, frozen; never regenerated |
generate_data.py |
Builds the score bundles from the CSVs in data/ |
blog/ |
Blog index and posts |
traces/ |
Trace viewer (see below) |
tooltip.js, photo-mode.js |
Leaderboard tooltips; a hidden screenshot mode (type photo or add ?photo) |
paper-plots/ |
Figures and tables for the paper; not used by the site |
The page shows v1.2 by default (data-current-results-version on <html> in index.html);
?version=v1.1 or ?version=v1 shows an older leaderboard.
-
Put the aggregation outputs in
data/v1.2/:aggregated_avg_<Agent>.csvandaggregated_std_<Agent>.csv, one row per base model, scores in 0–1:model,aime2025,arenahardwriting,gpqamain,gsm8k,healthbench,humaneval Qwen3-1.7B-Base,0.022,0.004,0.174,0.509,0.093,0.327
single_metrics_aggregated.csv(agent,avg,std,n): the overall score shown on the leaderboardtime_aggregated.csv(agent,avg_time,std_time,n) andaggregated_time_overview.csvfor runtimes
Baselines (
data/aggregated_baseline*.csv) are shared across versions. Benchmark weights are indata/factors-v1.2.json. -
Regenerate the bundle:
python3 generate_data.py --version v1.2
This writes
scores-v1.2.jsonandscores-v1.2.js. It stops with an error if a published agent is missing scores, standard deviations, an overall score or a runtime. -
Commit the CSVs and both generated files. Never edit the generated files by hand.
generate_data.py: map the agent's names to a key (e.g.opus-5.5-max):CSV_TO_AGENTandSTD_CSV_TO_AGENT: itsaggregated_avg_*/aggregated_std_*filenamesAGGREGATED_NAME_TO_KEY: its name insingle_metrics_aggregated.csvTIME_AGGREGATED_TO_KEY: its name intime_aggregated.csvV12_AGENT_KEYS: add the key to publish it- Single-run agents use
SINGLE_RUN_FINAL_TO_KEYandTIME_OVERVIEW_TO_KEYinstead. - If some of its cells are fallbacks from another agent, add a
CELL_PROVENANCEentry so the leaderboard labels them.
config.js:agentInfo: display name, description, scaffold, optionalreasoningEffort, and flags such asisExternalorprovenanceLabelallAgentKeys: table order before sortingchartAgentKeys(andchartAgentKeysByVersionto hide it from one version's chart): include it in the main charttimeChartAgentKeys: include it in the budget chart
- Regenerate the bundle (above).
Each post is a folder, blog/<slug>/index.html, listed by hand in blog/index.html.
Posts share blog/posttrainbench-1-1/article.css (plus an optional article.css of their own),
blog/nav.js (mobile menu), blog/toc.js (highlights the current section in the contents list)
and the theme toggle from traces/assets/theme.js. Copy an existing post to start a new one.
traces/ is the viewer's code only. Run data is loaded from the Hugging Face dataset
aisa-group/PostTrainBench-Trajectories,
set in traces/config.js. Both are produced by the separate ptb-traces-pipeline repo, whose
deploy.sh uploads the data and copies the viewer code here. Data-only updates need no change to
this repo.