Conversation
|
| Project | Bencher |
| Branch | u/ep/parameters-api/plot-parameters |
| Testbed | intel-v1 |
Click to view all benchmark results
| Benchmark | Latency | Benchmark Result microseconds (碌s) (Result 螖%) | Upper Boundary microseconds (碌s) (Limit %) |
|---|---|---|---|
| Adapter::Json | 馃搱 view plot 馃毞 view threshold | 5.18 碌s(+6.57%)Baseline: 4.86 碌s | 5.90 碌s (87.76%) |
| Adapter::Magic (JSON) | 馃搱 view plot 馃毞 view threshold | 4.96 碌s(+5.30%)Baseline: 4.71 碌s | 5.60 碌s (88.42%) |
| Adapter::Magic (Rust) | 馃搱 view plot 馃毞 view threshold | 27.39 碌s(+3.54%)Baseline: 26.45 碌s | 29.96 碌s (91.40%) |
| Adapter::Rust | 馃搱 view plot 馃毞 view threshold | 4.70 碌s(+20.29%)Baseline: 3.91 碌s | 6.09 碌s (77.16%) |
| Adapter::RustBench | 馃搱 view plot 馃毞 view threshold | 4.67 碌s(+19.52%)Baseline: 3.90 碌s | 6.07 碌s (76.83%) |
284e133 to
9d4a326
Compare
9d4a326 to
cf8ffa3
Compare
cf8ffa3 to
2875eb1
Compare
Review after the rebase onto develRebased onto devel with no source conflicts. Two independent reviews and an adjudication found no major defect introduced by the rebase. Cleanup applied from review: the Follow-up: a pinned plot's filter never reaches the consoleA plot's |
2875eb1 to
69ffb8b
Compare
69ffb8b to
1034559
Compare
b6eebe4 to
26000cd
Compare
A threshold checked one thing: the `value` metric of every variant of its measure. A benchmark now reports as many variants as it has sets of parameters and as many metrics as its harness measured, and neither was addressable. A project that wanted `p99` watched on the one configuration where it matters had nowhere to say so. A threshold gains two optional fields. `metric` is the name it checks, and a threshold that names none checks the conventional `value` name: a threshold always checks exactly one name and a bare one never checks all of them. `parameters` is a filter over variants, a list of partial parameters that is an OR across the list and a subset match within each entry, and a threshold with no filter checks every variant. Both defaults are what every existing threshold already does, so no existing row moves and no project's alert volume changes. Every threshold that matches a metric row runs. There is no winner: not in checking, not in display. A variant that a bare threshold and a filtered threshold both match earns a boundary from each and, on a regression, an alert from each. That is the design and it is pinned by a test, because a row that two people asked to be watched is a row two people hear about. Identity is the three dimensions plus the two new fields, under the null semantics they carry. SQLite treats nulls as distinct in a unique index, so two bare thresholds on one branch, testbed, and measure would no longer collide under a plain unique key over the five columns. The key is declared over the effective values instead, `COALESCE(metric, 'value')` and `COALESCE(parameters, x'')`, and the wire canonicalizes into them: an explicit `value` and an absent name are one threshold, an empty filter and an absent one are one threshold, and a filter has one spelling because its entries sort by their RFC 8785 canonical bytes and duplicates collapse. A threshold checks the sample it names. The historical query behind detection filters on the threshold's metric name, so a threshold on `p99` is tested against `p99` rows and never against the `value` rows beside them, and the per variant separation stays exactly as it was. `boundary` keys on `(metric_id, threshold_id)` rather than on `metric_id` alone, because a metric row may now carry a boundary per threshold that checked it. Both tables are rebuilt with their unique keys built after the copy, which is what keeps the rebuild at the cost of the scan. `JsonAlert` gains `value`, the metric value the alert fired on, and its `metric` triple becomes optional: the triple is a convention over the `value` name, so an alert on any other name has none. Every alert that a threshold could raise before this carries the triple exactly as it did. The checked name is readable at `alert.threshold.metric`. The deprecated singular `threshold`, `boundary`, and `alert` fields carry the bare threshold's boundary and alert and no other threshold's, everywhere they appear. That is precisely what a caller from before named thresholds has always been shown: a row that only a named or filtered threshold checks reports no threshold in them at all. Where a list of boundaries is returned, it is ordered by threshold creation time, oldest first, with the UUID breaking a tie. The in-report `thresholds.models` map is unchanged and still addresses the bare threshold: a map that names a measure and a model says nothing about a name or a variant, so it neither creates nor resets anything narrower.
A report has always carried its thresholds as a map of measure to model, and a map
key names a measure and nothing else. Every threshold a report could declare was
therefore the bare one: the conventional `value` name of every variant. A pipeline
that wanted `p99` watched, or wanted only some variants watched, had to reach for
the thresholds endpoint and then keep it in step with the run by hand.
At BMF version 1 `thresholds.models` is a list. An entry is `parameters`, `measure`,
`metric`, and `model`: the dimensions a threshold hangs off that the report does not
already state, in their canonical order, and the model to check with. `measure` is
the same name, slug, or UUID the map key is today, created if the project has never
seen it. An absent `metric` is the conventional `value` name and an absent
`parameters` checks every variant, so an entry naming only a measure and a model is
the map pair written out longhand. One measure may carry several entries, and each
one creates or updates the threshold with that identity under the report's branch
and testbed, through the same null collapse and canonical filter storage every other
writer goes through.
A threshold the report declares checks the very report that declared it. Nothing
moved to make that true: thresholds are resolved before results are parsed, which is
where they already were. It is worth saying out loud because it is what makes the
list worth having, and it is pinned by a test whose history is five unchecked
reports and whose alerts all belong to thresholds that did not exist when the
request arrived.
The report's BMF version, the one it declares or else the project's default, says
which shape to expect, and the shape is checked rather than guessed at. A list at
version 0 and a map at version 1 are both a 400 naming that version and the shape it
calls for, and the check runs before anything is created for the report.
`reset` reaches as far as the report's version can address. A version 0 map can only
name bare thresholds, so it takes a model away from bare thresholds and nothing
else, which is what it already did. At version 0, `reset` cannot strip a named or
filtered threshold the payload has no way to spell. A version 1 list can name every
identity, so `reset` reaches every threshold on the branch and testbed that the
entries did not name, including a payload that names none at all.
One identity declared twice in one payload is one threshold: the position is
where it was first written and the model is what it was last told, resolved
without an error. Two spellings of one filter are one identity, because a filter
canonicalizes before it is compared.
Two shapes behind one key is a place where error quality quietly dies. The obvious
spelling, an untagged enum, buffers the input, tries each variant, and on failure
says only that nothing matched, so a misspelled model test in a version 0 map would
come back as "data did not match any variant" rather than as the field and the
variants it could have been. The shape is known from the first token, so it is
decided by looking rather than by trying: every error a version 0 client got before
this layer it gets after it, byte for byte, and a malformed version 1 entry is named
by its position and its field. The one message that moves is the one that has to,
where `models` is neither shape and what it expects now names both.
Deleting a variant learns about thresholds. A threshold whose filter names a
variant is a reference to that row, so the variant cannot go out from under it, and
the refusal says which threshold to delete first. Naming is canonical equality and
nothing looser: a filter of `{"a":1}` matches the variant `{"a":1,"b":2}` without
naming it, because a filter names only the keys it cares about and is a predicate
over values rather than a pointer at a row. That variant can be deleted and the
filter still says what it said. The comparison runs in Rust over the project's
filtered thresholds, because canonical equality is what the canonical form defines
and that form is written in Rust. Nothing caps that read: it is small because a
filtered threshold is a rare thing to write and deleting a variant is a rare thing
to ask for, not because a limit says so. The delete and the check share one
transaction with the delete first, so a variant that a report still references is
refused for that reason and the client is sent to the results rather than to a
threshold it would have to delete anyway.
The CLI is unchanged: it declares bare thresholds, which is what the version 0 map
spells, and naming a metric or a variant from the command line is a separate piece
of work.
A pinned plot stores what the perf URL stores. The `plot` table gains a nullable `parameters` column holding the canonical filter list, the same JSONB encoding the threshold column uses, and `JsonPlot`, `JsonNewPlot`, and the plot patch shape gain the field in the canonical dimension order, after `benchmarks` and before `measures`. NULL is match all, so every plot that predates this carries NULL and draws exactly what it drew before.
26000cd to
953768a
Compare
A pinned plot stores what the perf URL stores.
The other dimensions of a plot are entity references, so they live in join tables of UUIDs. The parameters dimension in query space is not a reference: it is a value predicate, a filter, so the plot stores the same canonical list of entries the perf query takes.
A UUID join table was considered and rejected. A pinned variant's full parameters used as a filter also match future supersets of themselves, so UUID pinning cannot faithfully drive the ruled query, and a filter column round-trips "pin the current perf view" exactly.
The field
The
plottable gains a nullableparameterscolumn: the SQLite JSONB encoding of a JSON array of partial parameters, the same encoding thethresholdcolumn uses.JsonPlot,JsonNewPlot, and the plot patch shape gainparametersafterbenchmarksand beforemeasures, the canonical dimension order.NULL is match all, so every plot that predates this migration carries NULL and draws exactly what it drew before. A plot with no filter answers with the field absent, and the migration is a plain
ADD COLUMN.Canonical semantics
The filter means exactly what the perf query's
parametersparam means: OR across the list, subset match within each entry. A variant matches when any entry of the list is a subset of its parameters, so a filter names only the keys it cares about and a variant that pins more keys still matches.Canonicalization is
ParameterFilter's own, unchanged and unduplicated: each entry in its RFC 8785 canonical form, the list sorted by canonical bytes and deduplicated, at most eight entries. A list holding the empty entry and the empty list are both match all, and so is an absent field, so all three are one stored state: no filter at all.On a patch, an absent
parametersleaves the plot's filter alone. An explicitnulland an empty list both clear it back to every variant, which are two spellings of one filter and so land on one stored value.The console will pass the filter through when it builds the perf query. No console change here.