Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
17 changes: 17 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -624,6 +624,23 @@ after its public API and format compatibility policies are established.

### Fixed

- Benchmark report admission refuses identical or conflicting repeated metadata
coordinates with a typed failure naming the coordinate (issue #142). Complete
canonical report admission remains in progress. Ordered metadata keys, exact
headers/catalogs, complete metric widths and canonical unsigned decimals now
refuse malformed reports while admitting the committed historical fixture.
Fixed measurement policies and consistent per-row sample counts are admitted
before publication. Ratios, throughput, percentile order and reused-chunk
bounds now refuse inconsistent evidence with typed expected/observed failures;
arithmetic and numeric parsing refuse without approximation. Byte/count
metrics admit only their portable unsigned 64-bit range; timing and throughput
retain unsigned 128-bit precision. Publication requires an immutable admitted
report; refusal preserves the prior artifact and retained recovery stage.
Direct parser admission enforces the same one-MiB input ceiling as subprocess
capture before decoding, and preserves UTF-8 error sources. A dedicated
I/O-free fuzz facade exercises the production parser with a deterministic
historical seed; seed preparation and campaign discovery include the target.

- The benchmark workload catalog describes current range authentication once
before output, matching its single-pass metric definitions (#71).

Expand Down
9 changes: 8 additions & 1 deletion fuzz/Cargo.toml
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,7 @@ cargo-fuzz = true
[dependencies]
keep = { path = "..", version = "=0.0.0" }
libfuzzer-sys = { version = "=0.4.13", default-features = false, features = ["link_libfuzzer"] }
xtask = { path = "../xtask", version = "=0.0.0", default-features = false, features = ["golden-protocol-fuzz", "repository-json-fuzz"] }
xtask = { path = "../xtask", version = "=0.0.0", default-features = false, features = ["golden-protocol-fuzz", "repository-json-fuzz", "benchmark-report-fuzz"] }

[lints.rust]
warnings = "deny"
Expand Down Expand Up @@ -148,3 +148,10 @@ bench = false
[workspace]
members = ["."]
resolver = "3"

[[bin]]
name = "benchmark_report"
path = "fuzz_targets/benchmark_report.rs"
test = false
doc = false
bench = false
17 changes: 17 additions & 0 deletions fuzz/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -115,3 +115,20 @@ for the finite periods declared in `campaign.env`. A confirmed crash, timeout,
or out-of-memory input must be minimized and promoted into a committed
deterministic regression test; an artifact or cache alone never closes the
defect.

## Benchmark report admission

The `benchmark_report` target calls the production report parser through the
`benchmark-report-fuzz` feature, with task execution and host capture disabled.
Its coordinates are fixed to the committed historical baseline. Seed
preparation materializes that complete canonical report under the derived,
ignored corpus; no host measurement or new performance baseline is implied.
The facade performs no filesystem, process, clock or network operations.

The parser enforces a one-MiB input ceiling and preserves typed encoding,
canonical row, metadata, counter-width and arithmetic refusals. Unit laws
verify that the seed reaches admission and corrupted bytes refuse; these
prevent an always-refusing facade from passing solely because it never panics.
Run this target with the same pinned nightly and resource bounds as the other
targets. Campaign success is bounded exploration evidence, not proof that all
malformed reports have been enumerated.
10 changes: 10 additions & 0 deletions fuzz/fuzz_targets/benchmark_report.rs
Original file line number Diff line number Diff line change
@@ -0,0 +1,10 @@
#![no_main]

//! This target owns bounded canonical benchmark-report admission fuzzing.

use libfuzzer_sys::fuzz_target;
use xtask::admit_benchmark_report;

fuzz_target!(|bytes: &[u8]| {
let _ = admit_benchmark_report(bytes);
});
1 change: 1 addition & 0 deletions xtask/Cargo.toml
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,7 @@ publish = false
[features]
default = ["repository-tasks"]
golden-protocol-fuzz = []
benchmark-report-fuzz = []
repository-json-fuzz = ["dep:serde", "dep:serde_json"]
repository-tasks = [
"dep:blake3",
Expand Down
27 changes: 19 additions & 8 deletions xtask/src/benchmark_baseline/artifact.rs
Original file line number Diff line number Diff line change
Expand Up @@ -3,17 +3,27 @@
use super::BenchmarkBaselineError;
use super::environment::CapturedEnvironment;

pub(super) fn validate(
bytes: &[u8],
/// Immutable report bytes admitted against captured measurement coordinates.
#[must_use]
pub(super) struct AdmittedReport<'a> {
bytes: &'a [u8],
}

impl<'a> AdmittedReport<'a> {
pub(super) const fn bytes(&self) -> &'a [u8] {
self.bytes
}
}

pub(super) fn validate<'a>(
bytes: &'a [u8],
environment: &CapturedEnvironment,
) -> Result<(), BenchmarkBaselineError> {
let report =
std::str::from_utf8(bytes).map_err(|_source| BenchmarkBaselineError::ReportViolation {
reason: "report-is-not-utf8",
})?;
) -> Result<AdmittedReport<'a>, BenchmarkBaselineError> {
let report = super::report_input::decode(bytes)?;
if report.contains('\r') || !report.ends_with('\n') {
return violation("report-line-framing");
}
super::metadata_uniqueness::admit(report)?;
let mut lines = report.lines();
if lines.next() != Some("schema\tkeep.streaming-cas-baseline/v1") {
return violation("report-schema");
Expand Down Expand Up @@ -86,7 +96,8 @@ pub(super) fn validate(
{
return violation("report-profile-count");
}
Ok(())
super::report_grammar::admit(report)?;
Ok(AdmittedReport { bytes })
}

fn require_line(
Expand Down
6 changes: 5 additions & 1 deletion xtask/src/benchmark_baseline/artifact_publication.rs
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,11 @@ const LOCK_NAME: &str = ".streaming-cas-baseline-v1.lock";
const OUTPUT_RELATIVE_PATH: &str = "target/benchmark/streaming-cas-baseline-v1.tsv";
const STAGE_NAME: &str = ".streaming-cas-baseline-v1.tsv.stage";

pub(super) fn persist(repository_root: &Path, bytes: &[u8]) -> Result<(), BenchmarkBaselineError> {
pub(super) fn persist(
repository_root: &Path,
report: &super::artifact::AdmittedReport<'_>,
) -> Result<(), BenchmarkBaselineError> {
let bytes = report.bytes();
let output = repository_root.join(OUTPUT_RELATIVE_PATH);
let parent = output
.parent()
Expand Down
65 changes: 61 additions & 4 deletions xtask/src/benchmark_baseline/artifact_publication_tests.rs
Original file line number Diff line number Diff line change
Expand Up @@ -14,9 +14,11 @@ fn interrupted_baseline_stage_is_recovered_before_publication() -> Result<(), Bo
let stage = parent.join(STAGE_NAME);
fs::write(&stage, b"abandoned")?;

persist(directory.path(), b"exact report\n")?;
let (bytes, environment) = report_fixture()?;
let report = super::super::artifact::validate(bytes.as_bytes(), &environment)?;
persist(directory.path(), &report)?;

assert_eq!(fs::read(&output)?, b"exact report\n");
assert_eq!(fs::read(&output)?, bytes.as_bytes());
assert!(!stage.exists());
directory.close()?;
Ok(())
Expand All @@ -28,8 +30,10 @@ fn failed_baseline_publication_removes_its_stage() -> Result<(), Box<dyn Error>>
let (output, parent) = output_paths(&directory)?;
fs::create_dir(&output)?;

let (bytes, environment) = report_fixture()?;
let report = super::super::artifact::validate(bytes.as_bytes(), &environment)?;
assert!(matches!(
persist(directory.path(), b"exact report\n"),
persist(directory.path(), &report),
Err(BenchmarkBaselineError::Io {
action: "publish report",
..
Expand All @@ -52,8 +56,10 @@ fn concurrent_baseline_publishers_are_refused() -> Result<(), Box<dyn Error>> {
.open(parent.join(LOCK_NAME))?;
lock.try_lock()?;

let (bytes, environment) = report_fixture()?;
let report = super::super::artifact::validate(bytes.as_bytes(), &environment)?;
assert!(matches!(
persist(directory.path(), b"exact report\n"),
persist(directory.path(), &report),
Err(BenchmarkBaselineError::ReportViolation {
reason: "benchmark-publication-already-active"
})
Expand All @@ -71,3 +77,54 @@ fn output_paths(
fs::create_dir_all(&parent)?;
Ok((output, parent))
}

#[test]
fn refused_report_preserves_prior_artifact_and_interrupted_stage() -> Result<(), Box<dyn Error>> {
let directory = TestDirectory::create("benchmark-admission-preservation")?;
let (output, parent) = output_paths(&directory)?;
fs::write(&output, b"previous artifact\n")?;
let stage = parent.join(STAGE_NAME);
fs::write(&stage, b"interrupted stage\n")?;
let (bytes, environment) = report_fixture()?;
let malformed =
format!("{bytes}metadata\tgit-commit\tffffffffffffffffffffffffffffffffffffffff\n");
assert!(
matches!(super::super::artifact::validate(malformed.as_bytes(), &environment)
.and_then(|report| persist(directory.path(), &report)),
Err(BenchmarkBaselineError::DuplicateReportMetadata { coordinate })
if coordinate == "git-commit")
);
assert_eq!(fs::read(&output)?, b"previous artifact\n");
assert_eq!(fs::read(&stage)?, b"interrupted stage\n");
assert!(!parent.join(LOCK_NAME).exists());
directory.close()?;
Ok(())
}

fn report_fixture() -> Result<
(
String,
crate::benchmark_baseline::environment::CapturedEnvironment,
),
Box<dyn Error>,
> {
use crate::benchmark_baseline::environment::CapturedEnvironment;
use crate::benchmark_baseline::host_environment::CapturedHost;
let environment = CapturedEnvironment {
commit: String::from("c529c07f385b5bcd76a4e57c1987001d496f9135"),
tree: "clean",
rustc_version: String::from("rustc 1.96.0 (ac68faa20 2026-05-25)"),
target_triple: String::from("aarch64-apple-darwin"),
host: CapturedHost {
os_description: String::from("Darwin 25.3.0 arm64"),
cpu_model: String::from("Apple M1 Pro"),
logical_cpu_count: std::num::NonZeroUsize::new(10).ok_or("invalid CPU fixture")?,
},
};
Ok((
String::from(include_str!(
"../../../benchmark/baselines/c529c07-aarch64-apple-darwin.tsv"
)),
environment,
))
}
19 changes: 19 additions & 0 deletions xtask/src/benchmark_baseline/captured_environment.rs
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
//! This module owns captured source, compiler and host measurement coordinates.

use std::num::NonZeroUsize;

#[derive(Eq, PartialEq)]
pub(super) struct CapturedEnvironment {
pub(super) commit: String,
pub(super) tree: &'static str,
pub(super) rustc_version: String,
pub(super) target_triple: String,
pub(super) host: CapturedHost,
}

#[derive(Eq, PartialEq)]
pub(super) struct CapturedHost {
pub(super) os_description: String,
pub(super) cpu_model: String,
pub(super) logical_cpu_count: NonZeroUsize,
}
39 changes: 39 additions & 0 deletions xtask/src/benchmark_baseline/catalog_membership_tests.rs
Original file line number Diff line number Diff line change
@@ -0,0 +1,39 @@
//! Model laws for exact closed catalog membership without replacement aliases.

use super::{BenchmarkBaselineError, artifact, environment, report};

#[test]
fn unknown_and_duplicate_catalog_members_refuse_at_the_exact_slot()
-> Result<(), Box<dyn std::error::Error>> {
let environment = environment();
let valid = report(&environment, 13, 5);
for row in valid
.lines()
.filter(|line| line.starts_with("scenario\t") || line.starts_with("profile\t"))
{
let mut fields = row.split('\t');
let kind = fields.next().ok_or("missing catalog kind")?;
let name = fields.next().ok_or("missing catalog name")?;
let width = if kind == "scenario" { 3 } else { 7 };
let expected = row.split('\t').take(width).collect::<Vec<_>>().join("\t");
let unknown = row.replacen(
&format!("{kind}\t{name}\t"),
&format!("{kind}\tunknown\t"),
1,
);
let duplicate = valid
.lines()
.find(|candidate| candidate.starts_with(&format!("{kind}\t")) && *candidate != row)
.ok_or("missing duplicate fixture")?;
for observed in [unknown.as_str(), duplicate] {
let malformed = valid.replacen(&format!("{row}\n"), &format!("{observed}\n"), 1);
assert!(
matches!(artifact::validate(malformed.as_bytes(), &environment),
Err(BenchmarkBaselineError::InvalidReportRow {
expected: actual_expected, observed: actual_observed,
}) if actual_expected == expected && actual_observed == observed)
);
}
}
Ok(())
}
78 changes: 78 additions & 0 deletions xtask/src/benchmark_baseline/completeness_tests.rs
Original file line number Diff line number Diff line change
@@ -0,0 +1,78 @@
//! Model laws for required metadata and complete metric row widths.

use super::{BenchmarkBaselineError, artifact, environment, report};

#[test]
fn missing_and_unknown_metadata_refuse_the_exact_required_coordinate()
-> Result<(), Box<dyn std::error::Error>> {
let environment = environment();
let valid = report(&environment, 13, 5);
for row in valid.lines().filter(|line| line.starts_with("metadata\t")) {
let key = row.split('\t').nth(1).ok_or("missing fixture key")?;
let unknown = row.replacen(&format!("metadata\t{key}\t"), "metadata\tunknown\t", 1);
let successor = valid
.lines()
.skip_while(|line| *line != row)
.nth(1)
.ok_or("missing successor fixture")?;
for (replacement, observed) in [
(String::new(), successor),
(format!("{unknown}\n"), unknown.as_str()),
] {
let malformed = valid.replacen(&format!("{row}\n"), &replacement, 1);
let error = artifact::validate(malformed.as_bytes(), &environment)
.err()
.ok_or("incomplete metadata was admitted")?;
exact_metadata_refusal(error, key, observed);
}
}
Ok(())
}

fn exact_metadata_refusal(error: BenchmarkBaselineError, key: &str, observed: &str) {
if [
"build-profile",
"git-commit",
"git-tree",
"rustc-version",
"target-triple",
"os-description",
"cpu-model",
"logical-cpu-count",
]
.contains(&key)
{
let expected = format!("report-{key}");
assert!(
matches!(error, BenchmarkBaselineError::ReportViolation { reason } if reason == expected)
);
} else {
assert!(matches!(error, BenchmarkBaselineError::InvalidReportRow {
expected, observed: actual,
} if expected == key && actual == observed));
}
}

#[test]
fn every_metric_row_refuses_missing_or_extra_fields_exactly()
-> Result<(), Box<dyn std::error::Error>> {
let environment = environment();
let valid = report(&environment, 13, 5);
for row in valid
.lines()
.filter(|line| line.starts_with("scenario\t") || line.starts_with("profile\t"))
{
let (truncated, _field) = row.rsplit_once('\t').ok_or("missing metric fixture")?;
let extended = format!("{row}\t1");
for observed in [truncated, extended.as_str()] {
let malformed = valid.replacen(&format!("{row}\n"), &format!("{observed}\n"), 1);
assert!(
matches!(artifact::validate(malformed.as_bytes(), &environment),
Err(BenchmarkBaselineError::InvalidReportRow {
expected: "complete metric row", observed: actual,
}) if actual == observed)
);
}
}
Ok(())
}
Loading
Loading