Skip to content

QUIC: bound how long a due ACK waits for a packet to carry it - #231

Draft
deadcaf3 wants to merge 1 commit into
apple:mainfrom
deadcaf3:worktree-perf-bound-ack-bundling-wait
Draft

deadcaf3 wants to merge 1 commit into
apple:mainfrom
deadcaf3:worktree-perf-bound-ack-bundling-wait

Conversation

@deadcaf3

@deadcaf3 deadcaf3 commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

Problem

An ACK that the ACK policy wants sent now, but which would be the only frame in its packet, is held back in sendFrames so it can be bundled with outgoing data. The only thing that eventually sends it is the delayed ACK timer, which is armed for the max ACK delay (25 ms).

That hold comes from the ACK bundling work in #5. This change keeps the bundling: an ACK still rides on outgoing data whenever there is some. It only bounds the wait when there is none.

An endpoint that only receives has nothing to bundle with, so it acknowledges once per 25 ms. A sender whose congestion window is waiting for those ACKs then moves one window per 25 ms, whatever the path RTT. In QUICTransfer a 5 MB transfer over the in-process link takes 0.89 s (5.6 MB/s), and the process is idle for most of it.

RFC 9000 Section 13.2.2 recommends an ACK after two ack-eliciting packets and notes that delaying acknowledgments hurts window-based congestion controllers.

Fix

  • A due ACK with no packet to ride on now waits max(smoothed RTT / 2, Recovery.timerGranularity), capped at the max ACK delay, on a timer of its own. If a packet carries the ACK first, the timer does nothing.
  • The delayed ACK timer and the delay for a single unacknowledged packet are unchanged, so ACKs are still bundled whenever there is data to send.
  • Timer.recalculate keeps a wakeup that is already armed for the same deadline. While a deadline inside the 1 ms threshold was pending, every reschedule of another timer re-armed the wakeup. Without this, the new timer cost QUICStreamLoad 3%.

Testing

  • New testQUICOneWayTransferDoesNotWaitForMaxAckDelay: 1 MiB one way with the receiver's max ACK delay stretched to 30 s. It times out without the fix. runQUICTest gains shouldEchoData for it.
  • testDelayedAckTimerCanSendPMTUDProbe relied on the echoed data still being unacknowledged when it fires the timer. That ACK can now leave earlier, so the test owes the ACK explicitly.
  • testQUICEcho1MiBAckBundling passes unchanged.
  • Full test suite passes locally on macOS, debug and release.

Measurements

Setup: MacBook Air (Apple M1, 8 cores, 16 GB), macOS 26.6.2, Swift 6.4, release build, on AC power. Client and server run in one process over the in-process link, so this measures the stack and not a network. Times and rates are the ones each tool prints. Medians of 7 alternating rounds for QUICTransfer, 15 for the other two.

Benchmark Before After Change
QUICTransfer -iterations 10 (5 MB) 0.887 s, 5.6 MB/s 0.099 s, 50.3 MB/s 8.9x faster
QUICTransfer -iterations 300 (150 MB) 3.317 s, 45.2 MB/s 0.953 s, 157.4 MB/s 3.5x faster
QUICTransfer -iterations 125000 -size 1200 (150 MB) 3.338 s, 44.9 MB/s 1.664 s, 90.1 MB/s 2.0x faster
QUICTransfer -iterations 100 -link-delay-ms 1 (50 MB) 2.424 s, 20.6 MB/s 0.567 s, 88.2 MB/s 4.3x faster
QUICTransfer -iterations 100 -link-delay-ms 5 (50 MB) 2.590 s, 19.3 MB/s 1.759 s, 28.4 MB/s 1.5x faster
QUICTransfer -iterations 100 -link-delay-ms 20 (50 MB) 6.562 s, 7.6 MB/s 6.078 s, 8.2 MB/s 1.08x faster
QUICStreamLoad -stream-count 40000 16,478 streams/s 16,382 streams/s -0.6%
QUICHandshake -iterations 2000 741 handshakes/s 737 handshakes/s -0.6%
  • Every QUICTransfer row was faster in 7 of 7 rounds. The gain shrinks as the link delay grows, because the max ACK delay stops being the bottleneck once the RTT approaches it.
  • The last two rows are noise: the change was faster in 6 and 8 of the 15 rounds.
  • Process CPU time for the first three rows, from a separate run: 31% lower, 31% lower, unchanged.
  • Each connection registers one more timer. I did not count allocations.

- An ACK that is due but would be alone in its packet is held back to be
  bundled with data, and only the delayed ACK timer sent it, after the
  max ACK delay. An endpoint that only receives then acknowledges once
  per 25 ms, and a sender whose congestion window waits on those ACKs
  sends one window per 25 ms whatever the RTT
- Such an ACK now waits max(smoothed RTT / 2, timer granularity), capped
  at the max ACK delay, on a timer of its own. A packet that carries the
  ACK first makes that timer a no-op, and the delayed ACK timer is
  unchanged
- Timer.recalculate keeps a wakeup that is already armed for the same
  deadline. With a deadline inside the 1 ms threshold pending, every
  reschedule of another timer re-armed it
- Tests: a one-way transfer completes with the max ACK delay stretched
  to 30 s, and the PMTUD delayed ACK test owes its ACK explicitly
@rpaulo

rpaulo commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

Why isn't Ack::processPending() working for you? That should ACK more aggressively when we need to.

timerNow: now,
in: &eventContext
)
}

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could we run into a situation where we have two timers potentially scheduled here now?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah, I think both can be armed. Whatever fires second should find nothing to do, but it's an extra wakeup. I am all for optimizations. Would it be cleaner to skip the new timer and just pull the delayed ack timer earlier when an ACK is due? Open to thinking things thru cuz there's a tiny design decision here. cc @rpaulo

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One caveat tho, the ACK_FREQUENCY extension where sender asks receiver to ACK less often to save resources exists in QUIC, although there's no handling for it in the repo apart from maybe like a few placeholders, if it's implemented, bundling delay should respect the peer's request than using half the rtt.

@deadcaf3

deadcaf3 commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

@rpaulo

Why isn't Ack::processPending() working for you? That should ACK more aggressively when we need to.

I think it is firing. As far as I can tell, processPending() queues the ACK, but then the isAckOnly guard in sendFrames() holds it for bundling. With nothing to send, it waits for the 25 ms timer. Am I missing a path where it's supposed to go out sooner?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants