tpd: re-read only the transports that moved; keep settled days across restarts - #5274
Merged
Merged
Conversation
… restarts The metrics publisher rebuilt today's leaf every minute by reading every transport: ~107k HGETALLs and ~105k GETs per tick on the live deployment, nearly all for figures that had not changed. Bandwidth and latency are written through the store, so it now marks those transports and the tick re-reads only them; other rows carry over and liveness follows the registered set. The first read of a UTC day is still whole. Every restart also re-aggregated the whole 30-day window (~4.9M HGETALLs in 45 s, measured after a deploy). A settled day's leaf is now saved in redis once it settles, and startup loads those instead, reading the window only if a day is missing.
0pcom
added a commit
that referenced
this pull request
Sep 30, 2026
…5275) Keepalive pings move every live transport's byte counters, so the unchanged-shard dedup cannot skip them: both edges of ~80k transports reported every 45 s, ~3k bandwidth and ~1.5k latency scripts a second (~20 redis commands each) on prod01 after #5274 removed the metrics reads — now the bulk of TPD's load. Counters are cumulative, so applying fewer snapshots loses no bytes; the next delta spans the skipped ones. Latency is kept as the latest and throughput as the peak over held reports. A held snapshot is flushed when its window passes, and the first report of a UTC day applies at once. Uptime heartbeats are paced separately, as before.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Measured on prod01 (INFO commandstats every 5 s): the metrics publisher's 60 s tick was the whole steady-state HGETALL load — ~107k HGETALLs and ~105k GETs per tick, one per transport, nearly all for figures that had not changed. Bandwidth and latency are written through the store, so it now marks those transports and the tick re-reads only them; other rows carry over and liveness follows the registered set. The first read of a UTC day is still whole.
Each restart also re-aggregated the full 30-day window: ~4.9M HGETALLs in 45 s after a deploy. A settled day's leaf is now saved in redis when it settles, and startup loads those instead; the window is read only if a day is missing (so the first deploy of this reads it once more).
The redis-gated tests pass against miniredis locally and run in CI's redis-store job.