Skip to content

Sum checksum16 in 64-byte blocks with an add-with-carry chain - #224

Merged
rnro merged 2 commits into
apple:mainfrom
rnro:perf/checksum16-wide
Oct 6, 2026
Merged

rnro merged 2 commits into
apple:mainfrom
rnro:perf/checksum16-wide

Conversation

@rnro

@rnro rnro commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

checksum16 read one 16-bit word per iteration into a UInt32. At -Osize, which does not vectorize the loop, that is about 2 instructions per byte, and on a link without checksum offload every UDP datagram pays it on send and receive. The accumulator also dropped carries above about 128 KiB.

It now sums 64-byte blocks with a chain of UInt128 additions, which compiles to an add-with-carry chain and needs no SIMD types, with overlapping loads for the tail and no rebinding of the caller's memory. Added tests against a word-by-word sum at every length up to 300 and every start offset up to 15, and for carries in large buffers.

Measured on my Mac at -Osize, instructions retired per call:

85 bytes 184 -> 47
256 bytes 525 -> 81
1,200 bytes 2,413 -> 301
1,500 bytes 3,013 -> 380

At -O the old loop is auto-vectorized, and on a Mac it is faster than this one for buffers above about 576 bytes: about 15 ns against 40 ns at 1,500 bytes.

@rnro rnro added the 🔨 semver/patch No public API change. label Oct 6, 2026

@agnosticdev agnosticdev left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you!

rnro added 2 commits October 6, 2026 16:29
`checksum16` read one 16-bit word per iteration into a `UInt32`. At `-Osize`,
which does not vectorize the loop, that is about 2 instructions per byte, and on
a link without checksum offload every UDP datagram pays it on send and receive.
The accumulator also dropped carries above about 128 KiB.

It now sums 64-byte blocks with a chain of `UInt128` additions, which compiles to
an add-with-carry chain and needs no SIMD types, with overlapping loads for the
tail and no rebinding of the caller's memory. Added tests against a word-by-word
sum at every length up to 300 and every start offset up to 15, and for carries in
large buffers.

Measured on my Mac at `-Osize`, instructions retired per call:

  85 bytes        184 -> 47
  256 bytes       525 -> 81
  1,200 bytes   2,413 -> 301
  1,500 bytes   3,013 -> 380

At `-O` the old loop is auto-vectorized, and on a Mac it is faster than this one
for buffers above about 576 bytes: about 15 ns against 40 ns at 1,500 bytes.
@rnro
rnro force-pushed the perf/checksum16-wide branch from 61d873e to b7b4ba4 Compare October 6, 2026 20:29
@rnro
rnro merged commit d7b6d0e into apple:main Oct 6, 2026
36 of 39 checks passed
@rnro
rnro deleted the perf/checksum16-wide branch October 6, 2026 20:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

🔨 semver/patch No public API change.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants