Repository navigation
agent 1.6.2: take the push off the poll loop, and two fixes that followed - #15
Merged
Merged
Conversation
…owed Sync from the monorepo (agent/), covering 1.6.0 through 1.6.2. 1.6.0 — the push was on the poll loop's critical path. pollLoop reset its timer only AFTER deliver() returned, and deliver() blocks on an HTTPS round trip, so the real sampling period was `interval + game-server round trips + central's latency`. A venue configured for a 1s in-game rate measured 2.4s in production, and no value of the setting could fix it. Batches now go to a delivery goroutine that owns every push for the process's life and outlives every reconnect — which also means buffered telemetry keeps draining while the agent cannot reach the box, exactly when the buffer matters most. Separately, telemetryInterval() splits how often to LOOK from how often to RECORD. Fast polling while print-server work is in flight is a safety property and stays; it no longer drags the telemetry rate along with it, which had made a nominal 15s idle cadence average 7.9s. 1.6.1 — a queued batch could be discarded without ever being sent. Moving delivery onto its own goroutine let a batch be enqueued while another was in flight; the buffer is bounded and drops oldest-first, so the batch being sent is the one capacity evicts, and the unconditional pop then removed the batch that had taken its place at the head. Nothing logged it and nothing retried it, and it could only happen while the buffer was full enough to be evicting — during an outage, when the queue is the only record there is. The drain now removes a batch only if it is still the one it sent. 1.6.2 — the sample that reports a game has ENDED could be thrown away. The end of a game is the exact moment the recording interval widens, so the poll that discovers it was measured against the new, wider idle interval and rejected for arriving too soon. At 1s in play and 30s idle the state change could sit undelivered for most of a minute. A poll reporting a different server state is now always recorded; the exemption is one-shot, so the idle rate resumes immediately after. Operators: IDLE_POLL_INTERVAL at or above 30s crosses central's default "Agent offline" threshold, and config.go warns about it at startup. Raise that alert rule before raising the interval, or the site flaps offline in between. Gates: gofmt -l . clean, go vet ./... clean, go test ./... green, go build ./... clean, plus -race on the changed packages upstream. Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Pi8roSs8BLnmFXgEgMVU7e
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. Note Currently processing new changes in this PR. This may take a few minutes, please wait... ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Advanced Run ID: 📒 Files selected for processing (10)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Comment |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Sync from the monorepo (
agent/), covering 1.6.0 → 1.6.2.1.6.0 — the push was on the poll loop's critical path.
pollLoopreset its timer only afterdeliver()returned, anddeliver()blocks on an HTTPS round trip, so the real sampling period wasinterval + game-server round trips + central's latency:A venue configured for a 1 s in-game rate measured 2.4 s in production, and no value of the setting could fix it — work time can only lengthen that period, never shorten it. Batches now go to a delivery goroutine that owns every push for the process's life and outlives every reconnect, which also keeps buffered telemetry draining while the agent cannot reach the box — exactly when the buffer matters most.
Separately,
telemetryInterval()splits how often to look from how often to record. Fast polling while print-server work is in flight is a safety property and stays; it no longer drags the telemetry rate with it, which had made a nominal 15 s idle cadence average 7.9 s.1.6.1 — a queued batch could be discarded without ever being sent. Moving delivery onto its own goroutine let a batch be enqueued while another was in flight. The buffer is bounded and drops oldest-first, so the batch being sent is the one capacity evicts, and the unconditional pop then removed the batch that had taken its place at the head. Nothing logged it, nothing retried it, and it could only happen while the buffer was full enough to be evicting — during an outage, when the queue is the only record there is.
Buffer.PopSentnow removes a batch only if it is still the onePeekhanded out.1.6.2 — the sample reporting a game had ENDED could be thrown away. The end of a game is the exact moment the recording interval widens, so the poll that discovers it was measured against the new, wider idle interval and rejected for arriving too soon. At 1 s in play / 30 s idle the state change could sit undelivered for most of a minute. A poll reporting a different server state is now always recorded; the exemption is one-shot, so the idle rate resumes immediately after.
Related issue
n/a — upstream sync.
Type of change
Checklist
gofmt -l .prints nothinggo vet ./...passesgo test ./...passes.env.exampleand the README config table updated (if configuration changed) — no new configuration;IDLE_POLL_INTERVALis pre-existing and already documentedNotes for reviewers
Not one logical change, deliberately. This is a mirror sync, so it carries 1.6.0 and the two fixes that review found in it rather than splitting them. The commit body separates the three.
Deployment note — order matters.
IDLE_POLL_INTERVALat or above 30 s crosses central's defaultAgent offlinethreshold, andconfig.gowarns about it at startup. Raise that alert rule before raising the interval, or the site flaps offline in between. Leaving it unset keeps the 15 s default, which is safe.New tests, each verified non-vacuous by reverting its fix:
TestDrainKeepsABatchEvictedWhileAnotherWasInFlight— reproduces the 1.6.1 loss exactly (central never received {"push_seq":2})TestTheSampleThatEndsAGameIsNeverThrottledandTestAModeChangeBetweenGamesIsRecordedImmediately— the 1.6.2 transitionTestPopSentLeavesAHeadItDidNotSendAlone,TestPopSentStillDiscriminatesAfterALoad— the buffer contract, including that entry ids do not survive the spill fileUpstream also ran
-raceon the changed packages and the full monorepo suite.🤖 Generated with Claude Code
https://claude.ai/code/session_01Pi8roSs8BLnmFXgEgMVU7e
Generated by Claude Code
Summary by CodeRabbit