Skip to content

fix(bench): reject failed processes before scoring output - #18

Open
DivyamTalwar wants to merge 1 commit into
FedericoTs:masterfrom
DivyamTalwar:fix/qp-bench-exit
Open

DivyamTalwar wants to merge 1 commit into
FedericoTs:masterfrom
DivyamTalwar:fix/qp-bench-exit

Conversation

@DivyamTalwar

Copy link
Copy Markdown

Summary

Fixes #17.

The benchmark path parses its captured speed table without checking subprocess returncode. A failed run can therefore be scored or offered as a contribution. Reject nonzero exits before parsing and scoring; preserve successful results and bound the diagnostic tail.

Testing

Base: 252e5193902d466726da9af75047dfffff2ae662. Debian 12 Linux aarch64 in a nonroot disposable container, Python 3.11. The full existing smoke exits 0 with eight explicitly reported pre-existing optional research/hardware skips. It also prints baseline diagnostic skips where simulator/research dependencies are absent; no skipped check is counted as a pass. No Windows run or GPU/real-model benchmark is claimed.

The same final test files fail against unchanged production; the corrected branch gives:

Focused: 13 passed in 0.05s
Full applicable suite: 8 SKIPPED (a skip is not a pass):
all green
python tests/smoke.py
ruff check quantprobe
ruff format --check quantprobe
# Bandit medium-severity checks on changed package modules
git diff --check

All applicable commands above exited 0 locally. The snapshot is a complete upstream checkout plus this branch's exact changed-file bytes; no test is a copied production-function reimplementation. External I/O is mocked where stated. Dependencies and lockfiles are unchanged.

Security And Data Access

No credential, production-data, authentication, read-only guardrail or privileged workflow changes are included. Tests use synthetic inputs and disposable paths. No new benchmark, fitted law, or hardware capability is claimed.

Notes

This validates exit-status handling with a synthetic process boundary, not a performance benchmark. Existing hardware laws, calibration values and scoring formulas are unchanged.

AI-assisted implementation, isolated same-provider source review, and controller regression checks are disclosed. They are not maintainer approval, cross-vendor certification, or hosted CI. One focused, signed-off commit; no generated logs, personal config, model weights or worktree state is included.

Draft pending upstream CI and maintainer review. Companion changes touching the same module/test hook may require rebasing as they land; no combined branch is being submitted.

Address FedericoTs#17 with focused regression coverage.

AI-assisted implementation and isolated source review; exact validation and remaining platform limitations are recorded in the draft PR.

Signed-off-by: Divyam Talwar <divyamtalwar0@gmail.com>
@DivyamTalwar
DivyamTalwar marked this pull request as ready for review September 19, 2026 22:01

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

A failed llama-bench process can be scored after printing a speed table

1 participant