Skip to content

[test] Report deploy queue vs build timings in deploy tests - #99784

Closed
eps1lon wants to merge 1 commit into
canaryfrom
sebbie/deploy-timing-instrumentation
Closed

eps1lon wants to merge 1 commit into
canaryfrom
sebbie/deploy-timing-instrumentation

Conversation

@eps1lon

@eps1lon eps1lon commented Oct 7, 2026 •

Copy link
Copy Markdown
Member

Adds queue-vs-build timing instrumentation to deploy tests, motivated by the -c 8 experiment in #99770 where deploy shard failures were all 240s hook timeouts and the queue-vs-build split had to be inferred from logs (remote builds took 21–31s while deployments took ~5min wall — the rest was Vercel-side queueing).

Per deployment, a [deploy-timing] line is logged, split from the deployment's lifecycle timestamps (createdAt / buildingAt / ready from GET /v13/deployments/:idOrUrl, using the same token and team scope the deploy itself used):

[deploy-timing] dpl_5tjX… state=READY upload=11.2s queued=271.3s build=24.1s cli-overhead=1.8s total=308.4s

Reported on all three paths:

  • Success (deploy(), after the CLI returns)
  • CLI/deploy failure (throwDeploymentError()) — distinguishes queue time from build time on failed builds
  • Test aborted by hook timeout (destroy()): jest runs afterAll even when beforeAll times out, so destroy() now gives an in-flight deployment a bounded 90s grace to settle, then reports timings. If it still hasn't settled, a [deploy-timing] … timings unavailable line is logged instead.

To make the timeout path reachable, nextTestSetup's afterAll now always destroys the instance: previously await next?.destroy() silently skipped teardown when the beforeAll hook never resolved (jest timeout), leaving the running deployment to the module-level leak-detector afterAll (which then also failed the suite with "next instance not destroyed"). The afterAll now tracks the in-flight createNext promise, waits a bounded 30s grace, and falls back to the registered instance — so teardown runs, deployments settle enough to be timed, and the leak detector stays quiet. Both graces are sized to stay well under the 240s deploy hook timeout.

Safety: instrumentation can never fail a test — the API call is wrapped and only ever logs, and it no-ops when there is no token/scope (custom deploy script path) or no identifiable deployment.

Not included (deliberately): a shard-level post-step listing the run's deployments via the API — it would mix in unrelated deployments sharing the same Vercel project.

Co-Authored-By: Claude Code <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Failing CI jobs

Commit: e72eba7 | About building and testing Next.js

// time out.
if (this._startPromise && !this._deployEndedAt) {
await Promise.race([
this._startPromise.catch(() => {}),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Teardown grace Promise.race uses a ref'd timers/promises setTimeout, so the losing timer keeps the Jest worker's event loop alive for the full grace duration (90s / 30s) even after the deployment settles first.

Fix on Vercel

@eps1lon
eps1lon added this pull request to stack #99785 October 7, 2026 09:46
@eps1lon
eps1lon removed this pull request from stack #99785 October 7, 2026 10:02
@eps1lon

eps1lon commented Oct 7, 2026

Copy link
Copy Markdown
Member Author

Closing in favor of a simpler approach (get deployment ID synchronously at deploy start, time from there) to be done separately.

@eps1lon eps1lon closed this Oct 7, 2026
@eps1lon
eps1lon deleted the sebbie/deploy-timing-instrumentation branch October 7, 2026 10:03
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant