Skip to content

Add a pure-Rust Zstandard backend behind a zstd-rust feature - #45

Open
polaz wants to merge 9 commits into
SOF3:masterfrom
polaz:zstd-rust-backend
Open

polaz wants to merge 9 commits into
SOF3:masterfrom
polaz:zstd-rust-backend

Conversation

@polaz

@polaz polaz commented Sep 26, 2026 •

Copy link
Copy Markdown

The zstd feature binds the C library through zstd-sys, so every crate that embeds data with zstd needs a C toolchain at build time and links C code. This adds an opt-in zstd-rust feature that provides Zstandard through structured-zstd, a pure-Rust implementation of the format (no FFI, no cmake), as a separate compression method.

Nothing changes for existing users: the default features and the zstd method are untouched.

Behaviour

  • zstd-rust adds CompressionMethod::ZstdRust, selected with with zstd_rust, next to Deflate and Zstd.
  • Each method is gated by its own feature only. zstd and zstd-rust can be enabled together, and enabling one never changes what the other does.
  • zstd-rust alone builds without any C dependency (cargo tree shows neither zstd, zstd-sys nor libflate).
  • Both Zstandard methods produce and read the same frame format, so data embedded with one decodes with the other.
  • CompressionMethod::default() prefers deflate, then zstd, then zstd-rust, among the enabled features.

Changes

  • compress: ZstdRust variants of CompressionMethod, FlateEncoder, FlateDecoder and CompressionError. The ZstdRust encoder collects its input and writes it as one frame on finish, so the frame declares Frame_Content_Size; FlateEncoder::finish is public, so callers of CompressionMethod::encoder can end the stream for every method. The ZstdRust Read decoder is structured-zstd's StreamingDecoder (0.0.59), which decodes every frame of a stream and skips skippable frames; it keeps its state inline, so it is boxed to keep the enum compact (clippy::large_enum_variant).
  • compress: decompress_slice(bytes, method) decodes a whole buffer held in memory, meant for data this crate encoded. For ZstdRust, when every frame declares its size, the stream is decoded in one call into a buffer of exactly their total size (FrameDecoder::decode_all); otherwise it goes through the streaming decoder. Empty input, a declared size larger than a frame can hold (RFC 8878 3.1.1.2), frames holding less than they declare and bytes after the last frame that are not a frame are errors. The other methods go through their streaming decoders.
  • Root crate: decode (behind flate! and IFlate) uses decompress_slice; the feature is passed through, __parse_algo! accepts zstd_rust, and the # Algorithm section of flate! describes both Zstandard methods.
  • codegen: the zstd_rust keyword; with zstd_rust without the feature is a compile error naming it.
  • Tests: every asset test also embeds its file with zstd_rust (as [u8]/str and as IFlate). With both Zstandard features enabled, verify encodes each asset with one method and decodes it with the other, through both decoding entry points. Unit tests in compress cover the same on a synthetic input, the declared content size, concatenated and skippable frames on both entry points, small reads, a source that stays open, and the malformed inputs listed above.
  • CI: matrix entries for --no-default-features --features zstd-rust and --features zstd-rust, plus a check that zstd-rust pulls in no C bindings.

Measurements

Decoding jieba-rs's embedded dictionary (5,071,843 bytes) through include-flate's decode, each method decoding the frame it produced, 300 interleaved rounds, median:

method x86 (Xeon E5-1650 v4) vs zstd Apple M-series vs zstd compressed size
zstd (C, level 0) 10.35 ms 1.00x 6.34 ms 1.00x 2,126,841 B
zstd_rust (CompressionLevel::Default) 9.80 ms 0.95x 5.86 ms 0.92x 2,122,976 B

Test plan

  • cargo nextest run --workspace with default features, --no-default-features --features deflate, --no-default-features --features zstd, --no-default-features --features zstd-rust, and --features zstd-rust
  • cargo clippy --workspace --all-targets -- -D warnings for the same sets
  • cargo fmt --all -- --check
  • cargo tree --no-default-features --features zstd-rust --invert zstd-sys finds nothing

The zstd feature binds the C library through zstd-sys, which needs a C
toolchain at build time and links C code into every consumer. The new
zstd-rust feature provides the same CompressionMethod::Zstd and the same
`with zstd` syntax through structured-zstd, a pure-Rust implementation
of the format. Frames are interchangeable between the two backends; when
both features are enabled, zstd-rust is used.

- compress: backend selected by feature, same public API.
- codegen, root crate: feature passed through.
- Tests run under zstd-rust as well; a new interop test decodes C-zstd
  frames with the Rust backend and the reverse.
- CI: test matrix entries for zstd-rust, and a check that it pulls in
  neither zstd, zstd-sys nor libflate.
structured-zstd keeps the coder state inline, which made the Zstd variants of FlateEncoder and FlateDecoder far larger than the others (clippy::large_enum_variant). Boxing them keeps the enums compact; the C backend already holds its state behind a pointer.
Comment thread compress/src/lib.rs Outdated
type ZstdDecoder<R> = zstd::Decoder<'static, std::io::BufReader<R>>;

#[cfg(feature = "zstd-rust")]
fn zstd_encoder<W: Write>(write: W) -> io::Result<ZstdEncoder<W>> {

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

not a fan of exclusive features. i think they should be exposed as if separate decoding formats rather than replacing one another.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed, replacing one backend with the other made the features non-additive. Reworked in 9a485df: zstd-rust now adds its own CompressionMethod::ZstdRust, selected with with zstd_rust, and Zstd is back to the original code. Each method depends only on its own feature, so both can be enabled together without either one changing the other. The frame format is shared, so data embedded with one still decodes with the other.

Comment thread compress/src/tests.rs
@@ -0,0 +1,53 @@
// include-flate

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

should this have been a unit test instead of an integration test?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, moved it into a unit test of compress (compress/src/tests.rs). With the two methods now separate, it goes through apply_compression / apply_decompression only: encode with Zstd, decode with ZstdRust, and the reverse. On top of that, when both features are enabled, verify in test_util.rs decodes every asset across the two methods, so the interop check covers the real test files as well.

polaz added 2 commits October 6, 2026 11:03
zstd-rust used to replace the C backend behind CompressionMethod::Zstd
whenever both features were enabled. Feature unification made that a
global switch: any crate in the graph enabling zstd-rust changed the
backend of every other crate's `with zstd`, which is not additive.

zstd-rust now adds CompressionMethod::ZstdRust, selected by
`with zstd_rust`, next to Deflate and Zstd. Each method is gated by its
own feature only, and both Zstandard methods can be enabled together.
Frames stay interchangeable between them.

- compress: ZstdRust encoder/decoder variants and ZstdRustError; Zstd is
  back to the upstream code. Default prefers deflate, then zstd, then
  zstd-rust.
- codegen, root crate: zstd_rust keyword and __parse_algo arm.
- Tests: every asset test also embeds with zstd_rust. verify() encodes
  with each Zstandard method and decodes with the other when both are
  enabled; the interop check is now a unit test of the compress crate.
- test_util: unwrap_or instead of a lazy closure (clippy).
- structured-zstd 0.0.56 -> 0.0.58.
0.0.57 and 0.0.58 change only the encoder, dictionary training and the
CLI, and 0.0.58 decodes about 4% slower. include-flate decodes at run
time, so the older release is the better pin.

jieba-rs dictionary (5,071,843 B), one C-encoded frame, 200 interleaved
decodes per decoder, two runs:

  0.0.56: min 7.05 / 7.19 ms, median 9.05 / 7.86 ms
  0.0.58: min 7.33 / 7.39 ms, median 9.90 / 8.18 ms
@SOF3

SOF3 commented Oct 7, 2026

Copy link
Copy Markdown
Owner

Another question. What is the reason for using this instead of ruzstd?

polaz added 5 commits October 7, 2026 18:35
The streaming decoder routes every sequence through its internal ring buffer and hands the output out through Read, which made zstd_rust decoding about 1.5x the C time on x86. The embedded data is always a whole frame in memory, so it can be decoded straight into the final buffer instead.

- The zstd_rust encoder collects its input and writes it as one frame, so the frame declares Frame_Content_Size.
- decompress_slice decodes a frame that declares its size with FrameDecoder::decode_all into an exactly sized Vec, and rejects a frame that holds less than it declares. A frame without the size still goes through the streaming decoder.
- decode uses decompress_slice, so flate! and IFlate take this path.
- FlateEncoder::finish is public, so callers of CompressionMethod::encoder can end the stream; the zstd_rust encoder writes its whole frame there.
- decode_zstd_rust walks the frames first. When each declares Frame_Content_Size it decodes them all in one call into a buffer of their total size, otherwise it uses the library's read_to_end, which continues through following frames; skippable frames are skipped either way. The zstd_rust Read decoder decodes the stream this way on its first read, so a concatenated stream is no longer cut after its first frame.
- A declared size larger than the frame can hold (RFC 8878 3.1.1.2: at least 3 bytes on disk and at most 128 KiB out per block) is an error instead of a capacity-overflow panic or an allocation of that size.
- Regression tests: the public finish path, an impossible declared size, concatenated frames with and without declared sizes, leading and interleaved skippable frames, trailing bytes that are not a frame.
RFC 8878 3: compressed data is one or more frames, so an empty input (a fully truncated embedded frame) is InvalidData rather than an empty successful decode. Regression test: empty_input_is_not_a_zstd_rust_stream.
For ZstdRust the output is allocated at the size the frames declare,
bounded only by what their length could hold, so a crafted input can
request far more memory than it decodes to. decompress_slice is the path
for data this crate encoded (what flate! embeds); untrusted input belongs
on the streaming decoder.
structured-zstd 0.0.59 makes StreamingDecoder::read decode every frame
of a stream (RFC 8878 3), skip skippable frames wherever they stand,
including first, and hand out a finished frame's bytes before reading
the next header. The zstd_rust Read decoder goes back to a plain
StreamingDecoder, so it streams again instead of waiting for the end of
its source and holding the whole input and output in memory, and keeps
its source after a WouldBlock.

decompress_slice keeps the one-call path for frames that declare their
size and falls back to read_to_end over the whole stream otherwise.

Tests: small reads across concatenated and skippable frames, a source
that answers WouldBlock after a whole frame (fails on the previous
read-to-end decoder), skippable frames alone, empty input on both entry
points.
@polaz

polaz commented Oct 8, 2026

Copy link
Copy Markdown
Author

Fair question. Full disclosure first: structured-zstd is a fork of ruzstd that I maintain, and I'm actively working on its performance, so it's naturally the one I reached for. That said, which pure-Rust implementation include-flate uses is entirely your call. The integration is small, and the backend can be swapped without changing anything user-facing (with zstd_rust and the frame format stay the same).

For include-flate, what matters is decode speed at runtime and output size at build time, so I measured the options on the jieba dictionary (5,071,843 bytes), each decoding the frame its own encoder produces at the default level. Medians relative to the existing zstd (C) method:

x86 (Xeon E5-1650 v4) Apple M-series frame size
zstd (C) 1.00x 1.00x 2,126,841 B
structured-zstd, as in this PR 0.95x 0.92x 2,122,976 B
zstandard, into a slice / streaming 1.27x / 1.93x 0.84x / 1.05x 2,123,001 B
ruzstd (Fastest, the only level its encoder implements) ~3.4x¹ 2,712,315 B

¹ Measured on a different x86 machine (Xeon E5-2697 v4), where the C method decodes slower in absolute terms; the ratio is what carries over.

While measuring this I also updated the PR: a zstd_rust frame now declares its content size and is decoded in one call straight into an exactly sized buffer, instead of through the streaming reader. That was where most of the gap to C was (about 1.35x before the change).

If you'd rather not depend on a crate maintained by the contributor, zstandard is the alternative I'd pick over ruzstd on these numbers, and I'm happy to rework the PR onto either one.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants