Skip to content

Pin the parameters filter on a plot - #1021

Draft
epompeii wants to merge 3 commits into
u/ep/parameters-api/thresholds-payloadfrom
u/ep/parameters-api/plot-parameters
Draft

epompeii wants to merge 3 commits into
u/ep/parameters-api/thresholds-payloadfrom
u/ep/parameters-api/plot-parameters

Conversation

@epompeii

@epompeii epompeii commented Aug 26, 2026 •

Copy link
Copy Markdown
Member

A pinned plot stores what the perf URL stores.

The other dimensions of a plot are entity references, so they live in join tables of UUIDs. The parameters dimension in query space is not a reference: it is a value predicate, a filter, so the plot stores the same canonical list of entries the perf query takes.

A UUID join table was considered and rejected. A pinned variant's full parameters used as a filter also match future supersets of themselves, so UUID pinning cannot faithfully drive the ruled query, and a filter column round-trips "pin the current perf view" exactly.

The field

The plot table gains a nullable parameters column: the SQLite JSONB encoding of a JSON array of partial parameters, the same encoding the threshold column uses. JsonPlot, JsonNewPlot, and the plot patch shape gain parameters after benchmarks and before measures, the canonical dimension order.

NULL is match all, so every plot that predates this migration carries NULL and draws exactly what it drew before. A plot with no filter answers with the field absent, and the migration is a plain ADD COLUMN.

Canonical semantics

The filter means exactly what the perf query's parameters param means: OR across the list, subset match within each entry. A variant matches when any entry of the list is a subset of its parameters, so a filter names only the keys it cares about and a variant that pins more keys still matches.

Canonicalization is ParameterFilter's own, unchanged and unduplicated: each entry in its RFC 8785 canonical form, the list sorted by canonical bytes and deduplicated, at most eight entries. A list holding the empty entry and the empty list are both match all, and so is an absent field, so all three are one stored state: no filter at all.

On a patch, an absent parameters leaves the plot's filter alone. An explicit null and an empty list both clear it back to every variant, which are two spellings of one filter and so land on one stored value.

The console will pass the filter through when it builds the perf query. No console change here.

@github-actions

github-actions Bot commented Aug 26, 2026 •

Copy link
Copy Markdown
Contributor

馃惏 Bencher Report

ProjectBencher
Branchu/ep/parameters-api/plot-parameters
Testbedintel-v1
Click to view all benchmark results
BenchmarkLatencyBenchmark Result
microseconds (碌s)
(Result 螖%)
Upper Boundary
microseconds (碌s)
(Limit %)
Adapter::Json馃搱 view plot
馃毞 view threshold
5.18 碌s
(+6.57%)Baseline: 4.86 碌s
5.90 碌s
(87.76%)
Adapter::Magic (JSON)馃搱 view plot
馃毞 view threshold
4.96 碌s
(+5.30%)Baseline: 4.71 碌s
5.60 碌s
(88.42%)
Adapter::Magic (Rust)馃搱 view plot
馃毞 view threshold
27.39 碌s
(+3.54%)Baseline: 26.45 碌s
29.96 碌s
(91.40%)
Adapter::Rust馃搱 view plot
馃毞 view threshold
4.70 碌s
(+20.29%)Baseline: 3.91 碌s
6.09 碌s
(77.16%)
Adapter::RustBench馃搱 view plot
馃毞 view threshold
4.67 碌s
(+19.52%)Baseline: 3.90 碌s
6.07 碌s
(76.83%)
馃惏 View full continuous benchmarking report in Bencher

@epompeii
epompeii force-pushed the u/ep/parameters-api/plot-parameters branch 2 times, most recently from 284e133 to 9d4a326 Compare August 27, 2026 04:22
@epompeii
epompeii force-pushed the u/ep/parameters-api/plot-parameters branch from 9d4a326 to cf8ffa3 Compare August 27, 2026 05:11
@epompeii
epompeii force-pushed the u/ep/parameters-api/plot-parameters branch from cf8ffa3 to 2875eb1 Compare September 17, 2026 04:48
@epompeii

Copy link
Copy Markdown
Member Author

Review after the rebase onto devel

Rebased onto devel with no source conflicts. Two independent reviews and an adjudication found no major defect introduced by the rebase.

Cleanup applied from review: the plot_parameters migration now recreates the table so parameters sits after window, rather than ADD COLUMN placing it after modified. schema.rs, QueryPlot, and InsertPlot follow the same order.

Follow-up: a pinned plot's filter never reaches the console

A plot's parameters filter is stored and returned, but PinnedFrame.tsx builds the perf query from branches, testbeds, benchmarks, and measures only, so a plot pinned with a filter through the API or bencher plot create/update draws every variant in the console. specs is missing from the same query already. This is console work for the layer that renders variants: pass the plot's parameters into the pinned frame's perf query.

@epompeii
epompeii force-pushed the u/ep/parameters-api/plot-parameters branch from 2875eb1 to 69ffb8b Compare September 18, 2026 06:06
@epompeii
epompeii force-pushed the u/ep/parameters-api/plot-parameters branch from 69ffb8b to 1034559 Compare September 18, 2026 06:10
@epompeii
epompeii force-pushed the u/ep/parameters-api/plot-parameters branch 3 times, most recently from b6eebe4 to 26000cd Compare September 19, 2026 06:18
A threshold checked one thing: the `value` metric of every variant of its
measure. A benchmark now reports as many variants as it has sets of parameters and
as many metrics as its harness measured, and neither was addressable. A project that
wanted `p99` watched on the one configuration where it matters had nowhere to say
so.

A threshold gains two optional fields. `metric` is the name it checks, and a
threshold that names none checks the conventional `value` name: a threshold
always checks exactly one name and a bare one never checks all of them.
`parameters` is a filter over variants, a list of partial parameters that is an
OR across the list and a subset match within each entry, and a threshold with no
filter checks every variant. Both defaults are what every existing threshold
already does, so no existing row moves and no project's alert volume changes.

Every threshold that matches a metric row runs. There is no winner: not in
checking, not in display. A variant that a bare threshold and a filtered
threshold both match earns a boundary from each and, on a regression, an alert
from each. That is the design and it is pinned by a test, because a row that two
people asked to be watched is a row two people hear about.

Identity is the three dimensions plus the two new fields, under the null
semantics they carry. SQLite treats nulls as distinct in a unique index, so two
bare thresholds on one branch, testbed, and measure would no longer collide under
a plain unique key over the five columns. The key is declared over the effective
values instead, `COALESCE(metric, 'value')` and `COALESCE(parameters, x'')`, and
the wire canonicalizes into them: an explicit `value` and an absent name are one
threshold, an empty filter and an absent one are one threshold, and a filter has
one spelling because its entries sort by their RFC 8785 canonical bytes and
duplicates collapse.

A threshold checks the sample it names. The historical query behind detection
filters on the threshold's metric name, so a threshold on `p99` is tested against
`p99` rows and never against the `value` rows beside them, and the per variant
separation stays exactly as it was.

`boundary` keys on `(metric_id, threshold_id)` rather than on `metric_id` alone,
because a metric row may now carry a boundary per threshold that checked it. Both
tables are rebuilt with their unique keys built after the copy, which is what
keeps the rebuild at the cost of the scan.

`JsonAlert` gains `value`, the metric value the alert fired on, and its `metric`
triple becomes optional: the triple is a convention over the `value` name, so an
alert on any other name has none. Every alert that a threshold could raise before
this carries the triple exactly as it did. The checked name is readable at
`alert.threshold.metric`.

The deprecated singular `threshold`, `boundary`, and `alert` fields carry the
bare threshold's boundary and alert and no other threshold's, everywhere they
appear. That is precisely what a caller from before named thresholds has always
been shown: a row that only a named or filtered threshold checks reports no
threshold in them at all. Where a list of boundaries is returned, it is ordered
by threshold creation time, oldest first, with the UUID breaking a tie.

The in-report `thresholds.models` map is unchanged and still addresses the bare
threshold: a map that names a measure and a model says nothing about a name or a
variant, so it neither creates nor resets anything narrower.
A report has always carried its thresholds as a map of measure to model, and a map
key names a measure and nothing else. Every threshold a report could declare was
therefore the bare one: the conventional `value` name of every variant. A pipeline
that wanted `p99` watched, or wanted only some variants watched, had to reach for
the thresholds endpoint and then keep it in step with the run by hand.

At BMF version 1 `thresholds.models` is a list. An entry is `parameters`, `measure`,
`metric`, and `model`: the dimensions a threshold hangs off that the report does not
already state, in their canonical order, and the model to check with. `measure` is
the same name, slug, or UUID the map key is today, created if the project has never
seen it. An absent `metric` is the conventional `value` name and an absent
`parameters` checks every variant, so an entry naming only a measure and a model is
the map pair written out longhand. One measure may carry several entries, and each
one creates or updates the threshold with that identity under the report's branch
and testbed, through the same null collapse and canonical filter storage every other
writer goes through.

A threshold the report declares checks the very report that declared it. Nothing
moved to make that true: thresholds are resolved before results are parsed, which is
where they already were. It is worth saying out loud because it is what makes the
list worth having, and it is pinned by a test whose history is five unchecked
reports and whose alerts all belong to thresholds that did not exist when the
request arrived.

The report's BMF version, the one it declares or else the project's default, says
which shape to expect, and the shape is checked rather than guessed at. A list at
version 0 and a map at version 1 are both a 400 naming that version and the shape it
calls for, and the check runs before anything is created for the report.

`reset` reaches as far as the report's version can address. A version 0 map can only
name bare thresholds, so it takes a model away from bare thresholds and nothing
else, which is what it already did. At version 0, `reset` cannot strip a named or
filtered threshold the payload has no way to spell. A version 1 list can name every
identity, so `reset` reaches every threshold on the branch and testbed that the
entries did not name, including a payload that names none at all.

One identity declared twice in one payload is one threshold: the position is
where it was first written and the model is what it was last told, resolved
without an error. Two spellings of one filter are one identity, because a filter
canonicalizes before it is compared.

Two shapes behind one key is a place where error quality quietly dies. The obvious
spelling, an untagged enum, buffers the input, tries each variant, and on failure
says only that nothing matched, so a misspelled model test in a version 0 map would
come back as "data did not match any variant" rather than as the field and the
variants it could have been. The shape is known from the first token, so it is
decided by looking rather than by trying: every error a version 0 client got before
this layer it gets after it, byte for byte, and a malformed version 1 entry is named
by its position and its field. The one message that moves is the one that has to,
where `models` is neither shape and what it expects now names both.

Deleting a variant learns about thresholds. A threshold whose filter names a
variant is a reference to that row, so the variant cannot go out from under it, and
the refusal says which threshold to delete first. Naming is canonical equality and
nothing looser: a filter of `{"a":1}` matches the variant `{"a":1,"b":2}` without
naming it, because a filter names only the keys it cares about and is a predicate
over values rather than a pointer at a row. That variant can be deleted and the
filter still says what it said. The comparison runs in Rust over the project's
filtered thresholds, because canonical equality is what the canonical form defines
and that form is written in Rust. Nothing caps that read: it is small because a
filtered threshold is a rare thing to write and deleting a variant is a rare thing
to ask for, not because a limit says so. The delete and the check share one
transaction with the delete first, so a variant that a report still references is
refused for that reason and the client is sent to the results rather than to a
threshold it would have to delete anyway.

The CLI is unchanged: it declares bare thresholds, which is what the version 0 map
spells, and naming a metric or a variant from the command line is a separate piece
of work.
A pinned plot stores what the perf URL stores. The `plot` table gains a
nullable `parameters` column holding the canonical filter list, the same
JSONB encoding the threshold column uses, and `JsonPlot`, `JsonNewPlot`,
and the plot patch shape gain the field in the canonical dimension order,
after `benchmarks` and before `measures`.

NULL is match all, so every plot that predates this carries NULL and
draws exactly what it drew before.
@epompeii
epompeii force-pushed the u/ep/parameters-api/plot-parameters branch from 26000cd to 953768a Compare September 19, 2026 06:27

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant