Skip to content

test(handler): fix TestRelayRoutes flake on /hex/packages and /gem/info - #395

Merged
andrew merged 1 commit into
git-pkgs:mainfrom
wickedOne:relay-test-flake
Oct 1, 2026
Merged

andrew merged 1 commit into
git-pkgs:mainfrom
wickedOne:relay-test-flake

Conversation

@wickedOne

Copy link
Copy Markdown
Contributor

Summary

TestRelayRoutes fails intermittently on CI with write tcp ...: write: connection reset by peer at internal/handler/relay_test.go:86, the upstream test server's final rw.Flush(). main is red on this right now, and it has taken five of the six analytics PRs down with it. The failing subtest is always truncated=false on either /hex/packages/demo or /gem/info/demo.

Cause

Those two routes are the only entries in the table that fan out a second upstream request. With cooldown enabled — and the test sets proxy.Cooldown = &cooldown.Config{Default: "3d"} — HexHandler.fetchPackageAndVersions and GemHandler.fetchIndexAndVersions fetch version timestamps concurrently with the artifact itself:

route under test upstream requests the proxy makes
/hex/packages/demo /packages/demo and /api/packages/demo
/gem/info/demo /info/demo and /api/v1/versions/demo.json
every other route one

The test server has a single handler, so the sidecar request is answered with the same 64 KiB chunked 502 as the artifact request. fetchFilteredVersions returns as soon as it sees a non-200 and its deferred resp.Body.Close() abandons the body unread, so the transport tears that connection down while the test server is still writing to it. The server's next write gets the reset.

Whether the reset arrives before or after the Flush is a scheduling race between two goroutines and the kernel socket buffers, which is why it only shows up on loaded runners and never locally. Reproduced in isolation — a client that reads the headers of a chunked response and then calls Body.Close(), against a hijacking server that keeps writing — the error string is identical once the body outgrows the receive buffer:

64KiB    bodyWrite=<nil> flush=<nil>
256KiB   bodyWrite=<nil> flush=<nil>
1024KiB  bodyWrite=... write: connection reset by peer  flush=... write: connection reset by peer

On a Linux CI container the threshold sits far lower than on a developer macOS loopback, so 64 KiB is enough there.

Fix

The upstream handler cannot tell which of the two connections it is on, so a flush error there says nothing about the response under test — that is asserted on the downstream side, which is where a genuinely broken relay shows up. Report the write, drop the assertion.

Verification

go build ./..., go vet ./..., go test -race ./... and go tool golangci-lint run ./... all clean on main.

The test keeps its teeth. Two mutations of relayResponse are still caught:

  • dropping the relayTrailers call → trailers not relayed: map[X-Checksum:[]]
  • io.Copy → io.CopyN(w, resp.Body, 1024) → body length = 1024, want 65536

Not in scope

Abandoning the sidecar body also costs a reusable upstream connection on every cooldown-filtered hex/gem request that upstream answers with a non-200. Draining it under a byte cap before closing would be a small, separate improvement to the handlers; this PR only stops the test from asserting on a connection it does not own.

Context

Unblocks the pipelines on #388–#393, the six-part split of #381.

🤖 Generated with Claude Code

…ndons

TestRelayRoutes fails intermittently on CI with "write: connection reset
by peer" at the upstream handler's final Flush, on /hex/packages/demo and
/gem/info/demo.

Those two routes are the only ones in the table that fan out a second
upstream request: with cooldown enabled, the hex and gem handlers fetch
version timestamps concurrently with the artifact. The test server answers
that sidecar request with the same 64 KiB chunked 502, but the handler
returns as soon as it sees a non-200 and closes the body unread, so the
proxy drops the connection while the test server is still writing to it.
Whether the reset lands before or after the Flush is a scheduling race,
which is why it only surfaces on loaded runners.

The flush error says nothing about the response under test, which is
asserted on the downstream side, so stop reporting it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@andrew

andrew commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

thanks!

@andrew
andrew merged commit fb1630a into git-pkgs:main Oct 1, 2026
6 checks passed
@wickedOne
wickedOne deleted the relay-test-flake branch October 1, 2026 14:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants