Skip to content

lsv_measure: --only times only the listed parameter cases - #56

Merged
ArjunS07 merged 1 commit into
survey/phase1-template-20261008from
fix/confirm-only-cases
Oct 9, 2026
Merged

ArjunS07 merged 1 commit into
survey/phase1-template-20261008from
fix/confirm-only-cases

Conversation

@ArjunS07

@ArjunS07 ArjunS07 commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

Purpose

On bottleneck#304, --only (from #54) took 46+ minutes for confirm sets of 71 to 96 ids, not the expected ~10 minutes. It kept whole benchmark functions and dropped unlisted ids only after timing. So a parameterized function such as Time2DReductions.time_nanargmin timed all its parameter cases in every round: about 550 cases per side per round. This change makes --only time only the listed parameter cases.

How it works

asv's Benchmarks.benchmark_selection maps each benchmark name to the parameter indices to run (None for a benchmark without parameters). asv.runner.run_benchmark skips every index not in that list and records NaN for it. This is the asv in the task images (/opt/conda/envs/asv_3.12/.../asv/runner.py, the lightspeed fork).

flowchart LR
    A["--only ids<br/>p-3, p-7, x"] --> B["only_cases<br/>{p: {3, 7}}"]
    B --> C["selected.benchmark_selection<br/>p: [3, 7], x: None"]
    C --> D["run_benchmarks, base side"]
    C --> E["run_benchmarks, patched side"]
Loading
  • <name>-<index> ids narrow that benchmark's selection to the listed indices, intersected with its current selection.
  • A plain <name> keeps all cases of that benchmark.
  • The selection is set once on selected. run() uses that one object on both sides in every round, so the base and patched sides time exactly the same cases.
  • filter_out builds a new selection dict, so the full benchmark list is not changed.
  • run() now drops non-finite times (asv's NaN for skipped cases) in the same way as None.

Changes

  1. only_cases(only, names) in lsv_measure.py
    • Effect: returns {name: set of indices} for listed <name>-<index> ids whose name LSV selected, leaving out names that are also listed plain.
    • Before: none.
  2. measure_paired with --only
    • Effect: narrows selected.benchmark_selection before timing. A confirm set of 2 cases out of a 10-case benchmark times 2 cases per side per round.
    • Before: it timed all 10 and dropped 8 afterwards.
  3. run() drops non-finite times.
    • Effect: skipped cases do not appear in base_times / patched_times.
    • Before: only None was dropped. A NaN never came up because nothing was skipped.
  4. tests/docker/test_lsv_confirm_pass.py
    • Effect: the asv stub now follows benchmark_selection like asv's runner and logs every case it runs. New tests:
      • --only {p-3, p-7} on a 10-case benchmark runs exactly p-3, p-7 on each side of each of 4 rounds (16 runs).
      • Without --only, all 12 cases run.
      • Tests for only_cases.

Usage

No change: test.sh calls lsv_measure.py --only confirm_set.json --rounds N --out lsv_measure_confirm.json as before.

Verification

  • tests/docker: 450 passed, 10 skipped. Without the fix, the two new tests fail.
  • Against the real asv in a task image (fc-task/deshaw__versioned-hdf5__332, --network none, no timing): I built a real Benchmarks with a 10-case benchmark, applied the selection, and called asv.runner.run_benchmark with the single-case step replaced by a counter. Result: selection {'p': [3, 7], 'x': None}, ran [('p', 3), ('p', 7), ('x', 0)]. The full benchmark list still selected all 10 cases.
  • lsv_measure.py loads and only_cases runs on Python 3.8.

Notes

Expected time on #304, using the first pass's rate: 841 cases × 2 sides × 3 rounds in about 2670 s, so about 0.53 s per case per side. The confirm pass runs 6 rounds × 2 sides × (number of ids).

confirm ids cases per side per round, before after expected time, after
71 about 550 71 about 7.5 min
86 about 550 86 about 9 min
96 about 550 96 about 10 min

Before, about 550 cases took about 58 min at this rate; 46+ min was observed. That is about 6 to 8 times faster.

A listed <name>-<index> id now narrows asv's benchmark_selection for that benchmark, so the confirmation pass no longer
times every case of a parameterized function. One selection is used for the base and the patched side.
@ArjunS07
ArjunS07 merged commit 90ceb67 into survey/phase1-template-20261008 Oct 9, 2026
0 of 3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant