Skip to content

tpd: re-read only the transports that moved; keep settled days across restarts - #5274

Merged
0pcom merged 1 commit into
skycoin:developfrom
0magnet:tpd-metrics-today-deltas
Sep 30, 2026
Merged

0pcom merged 1 commit into
skycoin:developfrom
0magnet:tpd-metrics-today-deltas

Conversation

@0pcom

@0pcom 0pcom commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator

Measured on prod01 (INFO commandstats every 5 s): the metrics publisher's 60 s tick was the whole steady-state HGETALL load — ~107k HGETALLs and ~105k GETs per tick, one per transport, nearly all for figures that had not changed. Bandwidth and latency are written through the store, so it now marks those transports and the tick re-reads only them; other rows carry over and liveness follows the registered set. The first read of a UTC day is still whole.

Each restart also re-aggregated the full 30-day window: ~4.9M HGETALLs in 45 s after a deploy. A settled day's leaf is now saved in redis when it settles, and startup loads those instead; the window is read only if a day is missing (so the first deploy of this reads it once more).

The redis-gated tests pass against miniredis locally and run in CI's redis-store job.

… restarts

The metrics publisher rebuilt today's leaf every minute by reading every
transport: ~107k HGETALLs and ~105k GETs per tick on the live deployment,
nearly all for figures that had not changed. Bandwidth and latency are
written through the store, so it now marks those transports and the tick
re-reads only them; other rows carry over and liveness follows the
registered set. The first read of a UTC day is still whole.

Every restart also re-aggregated the whole 30-day window (~4.9M
HGETALLs in 45 s, measured after a deploy). A settled day's leaf is now
saved in redis once it settles, and startup loads those instead,
reading the window only if a day is missing.
@0pcom
0pcom merged commit 9d26056 into skycoin:develop Sep 30, 2026
14 of 16 checks passed
0pcom added a commit that referenced this pull request Sep 30, 2026
…5275)

Keepalive pings move every live transport's byte counters, so the
unchanged-shard dedup cannot skip them: both edges of ~80k transports
reported every 45 s, ~3k bandwidth and ~1.5k latency scripts a second
(~20 redis commands each) on prod01 after #5274 removed the metrics
reads — now the bulk of TPD's load.

Counters are cumulative, so applying fewer snapshots loses no bytes; the
next delta spans the skipped ones. Latency is kept as the latest and
throughput as the peak over held reports. A held snapshot is flushed
when its window passes, and the first report of a UTC day applies at
once. Uptime heartbeats are paced separately, as before.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant