Add cache-backed numerical model validation - #5
Conversation
## Feature additions - Add cache-backed export that writes a self-contained ONNX model under .onnx-cache and a .webnn graph with Safetensors weights under .webnn-cache. - Add --validate to validate a freshly converted cache and mutually exclusive --validate-cached to compare existing cached WebNN artifacts with native ONNX Runtime using deterministic fixed-shape inputs. - Expose cache and validation library APIs, including free-dimension overrides and validation input/output summaries. - Add Add-model and Tiny RoFormer integration coverage for export, reload, dispatch, and native ORT output comparison. ## Bugfixes - None. ## Refactors - Keep graph outputs ordered during export so cache serialization and output binding are deterministic. ## Behavioral impact and compatibility - Existing conversion behavior is unchanged when ConvertOptions::output_path is None and no cache flags are supplied. - ConvertOptions gains the public output_path field; external struct literals must initialize it or use Default. - Cached validation currently requires fixed shapes or explicit dimension overrides and supports deterministic inputs for float32, float16, int32, int64, bool, and uint8 tensors. ## Validation - cargo fmt --all -- --check passed. - ORT_DYLIB_PATH=/home/fkrall/vscprojects/transformers-convert/tools/onnxruntime/onnxruntime-linux-x64-1.29.0/lib/libonnxruntime.so.1.29.0 PROTOC=/home/fkrall/vscprojects/transformers-convert/tools/protoc/bin/protoc make test passed all 704 tests, including the Add cache round trip and Tiny RoFormer native-ORT comparison.
- Add an opt-in, manifest-driven numerical validation suite with smoke, extended, all, and text-match model selection. - Cache complete Hugging Face ONNX models and external-data sidecars using repository-relative paths, atomic downloads, completion metadata, optional authentication, and explicit refresh control. - Cache converted WebNN artifacts by model configuration and reload each exported .webnn and Safetensors pair before deterministic execution. - Add validation support for pinned inputs and deterministic int8, uint32, and uint64 tensors, with exact integer comparisons and dtype-aware floating-point tolerances. - Add coverage for manifest parsing and selection, safe sidecar paths, stable configuration keys, pinned inputs, accepted integer types, external-data execution, and full-model validation. - Mark Tiny RoFormer as the initial smoke-validation model. - Run the native ORT reference model from its filesystem path so standard external-data sidecars resolve correctly. - Preserve full-width int64 and uint64 outputs during comparison instead of relying solely on lossy floating-point conversion. - Exclude converter-pinned inputs from WebNN dispatch while still supplying their fixed values to native ORT. - Derive cached output keys from remaining WebNN inputs after pinning, preventing false output-name collisions when an ONNX output reuses a pinned input name. - Move manifest parsing and entry labeling into shared test infrastructure used independently by the skeleton and numerical suites. - Replace the hardcoded Tiny RoFormer integration test with the manifest-driven smoke runner while retaining lightweight synthetic cache tests. - Full-model numerical validation remains manual and is skipped unless O2W_MODEL_VALIDATION is set. - The existing skeleton sweep remains independently selectable and continues using the complete manifest. - ValidationSummary now reports pinned inputs separately, and validate_cached_model_with_options is added without removing the existing validation entry points. - Path-based native ORT validation requires RustNN commit 67d8713. - CI uses FelixKrall/rustnn branch fkrall/executable-webnn-reload for the required path-based ORT API; CI does not enable full-model validation. - make fmt passed. - PROTOC=... ORT_DYLIB_PATH=... make test passed all library, integration, generated operator, and documentation tests. - The focused pinned-input/output-name collision regression passed. - O2W_MODEL_VALIDATION=smoke cargo test --test model_validation -- --nocapture passed Tiny RoFormer through download cache, conversion, WebNN/Safetensors export, reload, deterministic CPU dispatch, and native ORT comparison. - An initial make test without ORT_DYLIB_PATH failed because libonnxruntime.so was unavailable; the configured rerun passed. - git diff --cached --check passed.
- Add a manifest-driven `validate-models` CLI with `--weights real|generated`, selection, manifest, and worker controls; real publisher weights remain the default. - Cache complete publisher ONNX models and external-data sidecars with atomic downloads, completion metadata, safe relative paths, optional Hugging Face authentication, and explicit refresh support. - Generate deterministic bounded replacement weights from the source revision, model fingerprint, tensor identity, role, and generator version. - Preserve small and shape/control tensors exactly, including range-reading external data; generate positive quantization scales and valid zero points; reject ambiguous or unsupported initializer roles. - Record retained/generated tensor roles, reasons, sizes, and digests in versioned cache metadata. - Add focused coverage for deterministic generation, constrained quantization values, exact external control tensors, ambiguous initializers, cache keys, selection, and generated-weight WebNN round trips. - Document the local `protoc`, RustNN, and ONNX Runtime prerequisites. - None. - Move manifest parsing, full-model caching, and skeleton handling from test-only helpers into the library so the validator CLI and skeleton integration test can share them. - Promote `ureq` to a runtime dependency and add `sha2` for deterministic generation and artifact fingerprints. - Replace the environment-gated numerical integration-test entry point with the reusable validation runner and CLI command. - Numerical model validation is now invoked with `validate-models`; callers using `O2W_MODEL_VALIDATION` or `O2W_MODEL_VALIDATION_JOBS` must migrate to `--selection` and `--jobs`. - `--weights real` downloads and validates publisher weights, while `--weights generated` avoids full-weight downloads by retaining structural constants and replacing learned parameters deterministically. - Every selected case is reconverted, exported to a mode/version-specific WebNN cache entry, reloaded, dispatched, and compared with native ORT using the same ONNX fixture. - The existing `O2W_MODELS=hub|dir=…|strip=…` skeleton sweep remains unchanged. - Generated validation is deliberately fail-closed: ambiguous tensor roles are rejected, and invalid accidental replacements of control metadata such as reshape/rescaling axes produce false-negative failures rather than being interpreted as coverage. Both execution paths consume the same generated ONNX fixture; generated-mode success is scoped to that fixture and does not certify publisher-weight behavior. - `cargo test --all-targets --locked --quiet` passed the library, cache, contribution, dialect, skeleton, and generated ONNX operator suites. - Tiny RoFormer passed end-to-end with `validate-models --selection smoke --weights generated --jobs 1`: 1/1 models, three inputs, and one compared output. - Tiny RoFormer passed end-to-end with `validate-models --selection smoke --weights real --jobs 1`: 1/1 models, three inputs, and one compared output. - `git diff --check` currently reports trailing whitespace in `tests/common/mod.rs`.
## Feature additions - None. ## Bugfixes - Register ONNX graph inputs with explicit empty shapes as rank-0 scalar operands instead of skipping them during conversion. - Add direct ONNX-to-RustNN execution coverage for a float32 scalar input and scalar `Identity` output, compared against native ONNX Runtime. ## Refactors - None. ## Behavioral impact and compatibility - ONNX models with runtime scalar inputs can now convert, build, and dispatch successfully. - Missing shape metadata and unresolved dynamic dimensions retain their existing validation behavior. - No public API, serialization format, dependency, or feature-flag changes. ## Validation - `cargo fmt --all -- --check` passed. - `cargo test --lib --quiet` passed all 204 tests. - `cargo test --test scalar_io --quiet` passed the scalar input/output ORT comparison. - The CI-configured release model-skeleton sweep passed 52 models with 0 failures and skipped 11 heavy entries. - `git diff --check` passed.
## Feature additions
- Add an explicit `untriaged` validation tier for newly generated and previously untriaged manifest cases.
- Add focused Python coverage for preserving branch-specific validation metadata and initializing new or null-valued cases.
## Bugfixes
- Preserve validation metadata for exact retained cases, including distinct prefill and decode entries that share a file and pinned inputs.
- Prevent `generate_manifest.py --apply` from silently dropping validation triage metadata.
## Refactors
- Replace the key-only manual-field lookup with branch-aware metadata selection shared with existing manifest baseline handling.
- Include `validation` in the canonical generated manifest field order.
## Behavioral impact and compatibility
- Generated manifest entries now include a `validation`` object.
- Exact retained cases keep their existing validation tier and blocker reason.
- New entries and retained entries without validation metadata receive `{"tier":"untriaged"}`.
- Validation metadata is not inherited across dtype variants because each new export requires independent triage.
- `untriaged` cases are excluded from `smoke` and `extended` selection, while remaining available through `all` and explicit matching.
- Existing `smoke`, `extended`, and `blocked` entries remain compatible.
## Validation
- `python3 -m py_compile scripts/generate_manifest.py scripts/test_generate_manifest.py` passed.
- `PYTHONPATH=scripts python3 -m unittest -v scripts/test_generate_manifest.py` passed 2 tests.
- `cargo test --all-targets --locked --quiet` passed all 727 tests with the configured `PROTOC` and `ORT_DYLIB_PATH`.
- `cargo fmt --all -- --check` passed.
- `git diff --check` passed.
## Feature additions - Add cache round-trip coverage for semantic token and mask inputs, plus zero-element symbolic inputs. - Add a dedicated real-weight validation failure report grouped by first blocker and repair area. ## Bugfixes - Generate zero-valued `token_type_ids` and one-valued `attention_mask` while preserving pinned-input precedence. - Preserve zero-element tensors instead of allocating one element, while continuing to treat `[]` as a one-element scalar. - Skip WebNN tensor reads and writes for empty buffers and reject missing non-empty cached inputs. - Detect tensor element-count overflow explicitly. - Report the cached graph’s actual input count in validation summaries. ## Refactors - Centralize checked element counting and typed deterministic input generation. - Separate semantic input policies from the generic deterministic patterns. ## Behavioral impact and compatibility - Exact `token_type_ids` and `attention_mask` inputs now receive model-valid deterministic values. - Zero-sized KV-cache tensors remain empty throughout native ORT and cached WebNN execution. - No CLI, manifest schema, or cache-format changes are introduced. - The real-weight sweep improves from 28 to 32 passing cases out of 52 and eliminates I1, I2, and I3 as first blockers, exposing subsequent V1 and N1 failures. - RustNN is unchanged. ## Validation - `cargo fmt -- --check` - `cargo check --all-targets` - `cargo test --all-targets`: all library and integration tests passed, including 483 ONNX operator tests and 8 cache-validation tests. - Generated Tiny RoFormer validation: 2/2 cases passed. - Complete real-weight sweep: 32/52 cases passed in 5m08.8. - `cargo clippy --all-targets` completed with four unrelated pre-existing warnings; `-D warnings` remains blocked only by those warnings. - `git diff --check` passed.
## Feature additions - Add cache-validation coverage for pinned `If` branches, pruned inputs, empty outputs, invalid extra descriptors, non-empty output omissions, and sanitized-name collisions. - Document the separate prefill/decode artifact pairs, fixed KV-cache shape limitation, and possible RAM/VRAM duplication when both graphs are resident. ## Bugfixes - Drive cached WebNN dispatch from the reloaded graph’s actual input interface, mapping retained descriptors back to unpinned source ONNX inputs instead of requiring inputs removed by branch specialization. - Reconstruct cached output mappings from the actual cached interface and compare native ORT outputs by name. - Permit a source output omitted by specialization only when native ORT proves it is empty; reject omitted non-empty outputs, unmatched cached descriptors, duplicate mappings, and sanitized input-name collisions. ## Refactors - Replace source-interface iteration and positional output comparison with explicit source-to-cache input and output maps. - Refresh the current validation ledger and remove the resolved V1 cached-interface failure section. ## Behavioral impact and compatibility - Pinned prefill and decode branches with intentionally pruned interfaces now validate successfully against their original merged ONNX model. - Invalid or ambiguous cached interfaces now fail with specific mapping diagnostics rather than being silently accepted or miscompared. - The complete real-weight sweep improves from 32 to 42 passing cases out of 52, clearing all ten V1 failures without introducing a new blocker family. - The `If` documentation now classifies constant or pinned branch selection as `Folded/static`; general runtime `If` remains unavailable in WebNN. - No CLI, manifest schema, cache format, conversion behavior, or RustNN changes are introduced. ## Validation - `cargo fmt -- --check` passed. - `cargo check --all-targets` passed. - `ORT_DYLIB_PATH=../rustnn/target/onnxruntime/onnxruntime-linux-x64-1.29.0/lib/libonnxruntime.so.1.29.0 cargo test --test cache_validation` passed all 13 tests. - `ORT_DYLIB_PATH=../rustnn/target/onnxruntime/onnxruntime-linux-x64-1.29.0/lib/libonnxruntime.so.1.29.0 cargo test --all-targets` passed, including 207 library tests, 13 cache-validation tests, 17 contrib-op tests, 15 dialect-op tests, 483 generated ONNX-op tests, and scalar-I/O coverage. - The initial cache-validation invocation without `ORT_DYLIB_PATH` failed during ONNX Runtime dynamic-library discovery before exercising the implementation. - Complete real-weight validation passed 42/52 manifest cases in 5m09.9s with no downloads or skipped cases. - `git diff --check` passed.
## Feature additions - Add an end-to-end packed Uint4 `MatMulNBits` fixture covering conversion, `.webnn` and Safetensors export, reload, execution, and comparison with native ORT. - Record invalid Chronos publisher artifacts and Voxtral’s unsupported `accuracy_level=4` execution mode as explicit, selectable manifest blockers. - Document packed-weight storage, reconstructed runtime dtypes, comparison tolerances, and the latest complete validation ledger. ## Bugfixes - Apply a measured `1e-3 + 1e-3 * abs(reference)` Float32 tolerance automatically to source graphs containing `MatMulNBits`, accounting for fused-ORT versus decomposed-WebNN accumulation differences without weakening ordinary Float32 comparisons. - Preserve strict classification of Voxtral rather than hiding its Int8-activation execution difference behind a broader tolerance. ## Refactors - None. ## Behavioral impact and compatibility - Depends on RustNN commit `<65e76e672aa64f59315b002099326b25f240802d>` (`Support packed 4-bit WebNN cache round trips`) for versioned packed Int4/Uint4 Safetensors export and reload. - MatMulNBits tolerance selection is deterministic and source-graph-driven; it does not depend on model names or manifest entries. - SmolLM2 prefill/decode and Janus now pass full numerical validation. Voxtral remains blocked because WebNN’s Float32 dequantize-plus-matmul lowering cannot reproduce ORT’s `accuracy_level=4` Int8 activation mode. - The latest real-weight ledger records 45 passes and 7 failures across all 52 manifest cases. ## Validation - `ORT_DYLIB_PATH=... make test` passed: 209 library tests, 14 cache-validation tests, 17 contrib-op tests, 15 dialect-op tests, 2 skeleton tests, 483 generated ONNX-op tests, and the scalar-I/O test. - `make check` and `git diff --cached --check` passed. - The complete cache-warm real-weight sweep passed 45 of 52 cases; the remaining failures are four numerical disagreements, Voxtral’s unsupported execution mode, and two invalid Chronos publisher artifacts.
## Feature additions - Report aggregate numerical-comparison diagnostics, including failure count, maximum and mean absolute error, RMSE, and maximum normalized error. - Derive separate q4 and q8 `MatMulNBits` comparison profiles from source-node attributes and report nonzero `accuracy_level` modes as explicitly unsupported. - Add cache round-trip coverage for numeric-to-Boolean Cast and positive-stride Slice. ## Bugfixes - Normalize numeric-to-Boolean Cast through a typed zero comparison so all nonzero values, including NaN, become true instead of retaining their numeric byte values. - Preserve positive ONNX Slice steps in `MLSliceOptions.strides` and pass input extents rather than output element counts as WebNN sizes. - Apply the measured q8 tolerance to Qwen instead of incorrectly evaluating it with the q4 envelope. ## Refactors - Add a shared Slice builder helper that records explicit strides. - Replace the Boolean `uses_matmul_nbits` comparison switch with attribute-derived comparison profiles. - Update the validation ledger, failure diagnosis, operator scope documentation, and provenance for the latest measured state. ## Behavioral impact and compatibility - DETR, Donut encoder, and Qwen now pass real-weight numerical validation, increasing measured coverage from 45/52 to 48/52. - Positive-stride Slice conversion now emits different, correct `.webnn` operation arguments. Previously converted affected caches must be regenerated to receive the fix. - FastVLM prefill is marked blocked because its fused GroupQueryAttention path materially differs from the decomposed WebNN path; its decode specialization remains passing. Blocked entries remain attempted by `all` and `match`. - Voxtral `MatMulNBits accuracy_level=4` now fails with an explicit unsupported-mode diagnostic rather than a generic numerical mismatch. - No CLI, manifest schema, cache format, or public library API changes are introduced. ## Validation - `cargo check --all-targets` passed. - `ORT_DYLIB_PATH=... cargo test --all-targets` passed all 743 onnx2webnn tests, including the unchanged skeleton sweep. - Focused Boolean Cast, positive-stride Slice, and q4/q8 comparison-profile regressions passed. - Targeted real-weight validation passed for DETR, all three Donut cases, and both Qwen cases. - The complete real-weight sweep finished in 7m 13.1s with 48/52 passing. Remaining failures were the documented FastVLM prefill, Voxtral `accuracy_level=4`, and two invalid upstream Chronos artifacts. - `cargo fmt --all -- --check` and `git diff --check` passed.
## Feature additions - Add a curated CI validation manifest with immutable Hugging Face revisions and SHA-256 verification. - Run curated real-weight numerical validation in non-Windows CI jobs, while providing a manually dispatched, diagnostic full-manifest workflow for a suitably provisioned self-hosted runner. - Make real and generated model acquisition revision-aware, including external-data sidecars, cache identities, and generated-model source selection. - Record revision and digest information in cache metadata and verify primary ONNX files after download and cache reuse. ## Bugfixes - Invalidate stale model caches when their format, source revision, or expected digest changes. - Include revision and digest information in source/cache keys so different publisher revisions cannot incorrectly share artifacts. ## Refactors - Remove `smoke`, `extended`, `blocked`, and `untriaged` metadata from the generated transformers.js manifest and move numerical status tracking exclusively into the validation documentation. - Restore the broad manifest to the generator-owned upstream population of 51 cases, moving the FP32 Tiny RoFormer case into the curated CI manifest. - Restrict manifest selection to explicit `all` or `match=<text>` modes and add coverage rejecting removed validation metadata. - Update validation ledgers and README guidance for the curated CI gate, diagnostic full sweep, cache requirements, and current 51-case baseline. ## Behavioral impact and compatibility - `validate-models` now requires `--selection`; callers must use `--selection all` or `--selection match=<text>`. The former `smoke` and `extended` values are no longer accepted. - Manifest entries containing the removed `validation` field are rejected as unknown input. Regenerate or remove that metadata and keep blocker classifications in the validation documentation. - The public manifest schema removes `ValidationConfig`, `ValidationTier`, and `Entry::validation`, while adding optional `revision` and `sha256` fields. - Real-model cache metadata advances to format version 2, so older completion records are invalidated and regenerated. - The broad manifest changes from 52 to 51 cases by removing one passing duplicate-model variant. Comparable coverage remains 51/51 skeleton, 20/51 generated, and 47/51 real, with the same four documented real-weight blockers. - Full-manifest model failures remain diagnostic in the manual workflow; setup, build, and runner-contract failures remain fatal. ## Validation - `cargo test --lib` with repository-local Protobuf and ONNX Runtime: 213 passed. - `cargo test --test cache_validation` with repository-local Protobuf and ONNX Runtime: 16 passed. - `cargo check --all-targets`: passed. - `python3 -m unittest discover -s scripts -p 'test_generate_manifest.py'`: passed. - `cargo fmt --all -- --check` and `git diff --check`: passed. - Verified the broad manifest matches `upstream/main`, contains 51 entries and no validation fields, and the current validation ledger contains 51 rows. - Inspected `validate-models --help` to confirm the required `all`/`match=<text>` selection interface.
| # .webnn reload and path-based native ORT execution for external-data models. | ||
| # Switch back to rustnn/rustnn only after both prerequisite rustnn PRs land. | ||
| RUSTNN_REPO: FelixKrall/rustnn | ||
| RUSTNN_REF: fkrall/executable-webnn-reload |
There was a problem hiding this comment.
I assume that needs to change in the final pr. you are just testing the CI
| let target = cache_root().join(&relative_cache_path); | ||
| let metadata_path = metadata_path(&target); | ||
| let refresh = std::env::var_os("O2W_MODEL_CACHE_REFRESH").is_some(); | ||
| if !refresh && complete_cache(&metadata_path, &target, revision, entry.sha256.as_deref()) { |
There was a problem hiding this comment.
do they just offer sha256 for verification? blake3 would be faster since parallelizable (so can use multi-threading or GPU accel)
There was a problem hiding this comment.
SHA-256 is used only by the numerical validation downloader for manually added CI models. The expected digest is manually recorded in ci-validation.json. After downloading the primary ONNX file and again before reusing it from cache. the validator computes its SHA-256 and compares it with that manifest entry. A mismatch prevents that model from proceeding to conversion and numerical comparison. The digest could be sourced from the Hugging Face LFS identity, but the current implementation does not retrieve it automatically. This verification is test infrastructure and does not affect normal onnx2webnn conversion or production artifacts.
## Feature additions - None. ## Bugfixes - None. ## Refactors - Remove the generated-weight model builder, cache preparation path, generator versioning, and associated tests. - Narrow `WeightMode` and manifest validation to publisher weights while retaining the existing `--weights` interface and `-real` WebNN cache-key suffix. - Simplify the manual full-model workflow to always run real weights and emit real-weight logs. - Remove generated-weight commands, status columns, failure families, timing, and cache claims from the README and validation ledger while preserving real-weight history. - Document why publisher-model verification continues using manifest-provided SHA-256 identities. ## Behavioral impact and compatibility - `--weights` still defaults to `real`, and `--weights real` remains supported. - `--weights generated` is now rejected during argument parsing; generated-model creation and generated cache reuse are no longer available on this branch. - Real-weight downloads, SHA-256 verification, conversion, artifact naming, reload, execution, and comparison remain unchanged. - The manually dispatched validation workflow no longer exposes a weight-mode selector and uploads `full-model-validation-real`. ## Validation - `cargo test --all-targets --locked -q` with `ORT_DYLIB_PATH` configured: 741 tests passed. - Complete skeleton sweep: 51/51 models passed. - Complete real-weight coverage: 47/51 models passed with the same FastVLM U1, Voxtral Q1, and two Chronos O1 blockers; the final case was resumed after the all-model command reached its wrapper timeout. - Confirmed `--weights generated` is rejected during CLI parsing. - `cargo check --all-targets --locked -q`, `cargo check --no-default-features --locked -q`, `cargo fmt --all -- --check`, and `git diff --check` passed. - Clippy remains blocked by four pre-existing warnings in unrelated ONNX conversion code.
## Feature additions - Add deterministic end-of-run reporting with complete succeeded and failed model lists, failure details, pass counts, and overall percentage. - Preserve manifest order in summaries so parallel validation results remain easy to interpret in CI terminal output. - Add focused coverage for summary formatting and percentage calculation. ## Bugfixes - None. ## Refactors - Replace separate atomic pass counting and failure-string collection with structured per-model results. - Expose `ModelFailure` and summary helpers for formatting and failure detection. ## Behavioral impact and compatibility - `validate-models` now prints its complete summary before returning a nonzero exit status when any selected model fails, making diagnostic CI logs unambiguous even when the workflow tolerates failures. - `run_manifest_validation` now returns a `RunSummary` containing model failures instead of returning an error for validation failures; setup and manifest errors remain errors. - Successful validation behavior and model execution are unchanged. ## Validation - `cargo test --lib model_validation::runner::tests --locked` — 3 focused tests passed. - `cargo check --all-targets --locked` — passed. - `cargo fmt --all -- --check` — passed. - `git diff --check` — passed. - Cached Tiny RoFormer real-weight validation completed successfully and printed `1/1 passed (100.0%)`.
## Feature additions
- Add centralized ONNX and WebNN cache resolution using the operating system cache directory under `onnx2webnn/{onnx,webnn}`.
- Add `O2W_CACHE_DIR` for relocating both caches while preserving the existing `O2W_ONNX_CACHE` and `O2W_WEBNN_CACHE` per-cache overrides.
- Add focused coverage for override precedence, shared cache roots, OS defaults, and unavailable cache directories.
## Bugfixes
- Stop installed binaries from deriving runtime cache paths from the build-time `CARGO_MANIFEST_DIR` checkout.
## Refactors
- Share cache resolution between conversion output, cached validation, full-model downloads, and manifest validation.
- Resolve the full-model cache root once per operation so model files, sidecars, and completion metadata use one consistent location.
- Add `dirs` as a direct dependency and document the new cache layout and overrides.
## Behavioral impact and compatibility
- Default artifacts move from checkout-local `.onnx-cache` and `.webnn-cache` directories to the platform cache directory.
- Existing checkout-local caches are not migrated automatically; set `O2W_ONNX_CACHE` and `O2W_WEBNN_CACHE` to reuse them.
- Resolution precedence is the per-cache override, then `O2W_CACHE_DIR`, then the operating-system default.
- Systems without a discoverable cache directory now receive an actionable error instead of silently using a build-machine path.
- `RunOptions::new` and `model_validation::full_model::cache_root` now return `Result`, requiring direct callers to handle cache-resolution failure.
## Validation
- `cargo test --all-targets --locked` — 743 tests passed.
- `cargo check --all-targets --offline` — passed.
- `cargo fmt --all -- --check` — passed.
- `git diff --check` — passed.
- CLI smoke testing with an explicit `XDG_CACHE_HOME` resolved ONNX and WebNN artifacts beneath `onnx2webnn/onnx` and `onnx2webnn/webnn` as intended.
## Feature additions - Split integer cache validation into independent `int8`, `uint32`, and `uint64` cases backed by a shared exact-round-trip fixture. - Add a strict CoreML XFAIL for unsupported `int8` Identity graph boundaries; unrelated failures and unexpected success remain test failures. - Mark the Voxtral merged decoder unsupported on CoreML because native MLProgram compilation overflows its worker-thread stack. ## Bugfixes - None. ## Refactors - Refresh the WebNN Graph branch lock from `c6473ef` to `067a86d`, matching the parent `rustnn` dependency. - Consolidate integer cache conversion, reload, execution, and validation into a reusable helper. ## Behavioral impact and compatibility - Production conversion and public APIs are unchanged. - Non-CoreML backends continue requiring exact integer round trips. - CoreML still exercises the incompatible `int8` path as an expected failure; `uint32` and `uint64` remain required to pass. - The CoreML skeleton sweep now reports Voxtral as explicitly unsupported instead of terminating with `SIGABRT`. ## Validation - CoreML cache-validation suite: 18 passed before rebase. - CoreML non-heavy skeleton sweep: 43 passed before rebase. - Heavy Qwen, SmolLM, and OpenAI Privacy Filter skeletons passed; Voxtral reproduced the annotated native stack overflow. - Curated Tiny model validation: 1 passed. - `git diff --check`: passed after rebase. - Tests were not rerun after the clean rebase.
| relative_path: &Path, | ||
| expected_sha256: Option<&str>, | ||
| ) -> Result<CachedFile, String> { | ||
| if let Some(parent) = target.parent() { |
There was a problem hiding this comment.
it would probably be better to let this download logic be handled by an official hugging face crate https://github.com/huggingface/huggingface_hub_rust Then, it would also be able to reuse caches with python hugging face and Rust hugging face.
There was a problem hiding this comment.
There was a problem hiding this comment.
Like it would also probably be able to resume downloads after being interrupted and you would need to start from zero when you have to interrupt at 31GiB/32GiB
There was a problem hiding this comment.
I agree with this. I didnt know about this crate. Changed to your suggested huggingface crate implementation. ONNX Models are now safed to the HF default cache dir by the validation executable and get all the HF download functionality.
## Feature additions - Download full validation models and external-data sidecars through the official Hugging Face Rust client with blocking and Xet support. - Reuse the standard Hugging Face `blobs` and `snapshots` cache while retaining onnx2webnn completion records, manifest SHA-256 verification, revision selection, and forced refreshes. - Support cache precedence through `O2W_ONNX_CACHE`, `O2W_CACHE_DIR`, `HF_HUB_CACHE`, `HUGGINGFACE_HUB_CACHE`, `HF_HOME`, and the platform cache directory. - Add unit coverage for cache precedence, completion-record identity, and cache-format validation. ## Bugfixes - Force-refresh a cached primary model once when its manifest SHA-256 does not match, then reject it if the refreshed artifact still differs. - Resolve every completed primary model and external-data sidecar locally before accepting a warm cache entry. ## Refactors - Replace the custom `ureq` full-model downloader, retry loop, partial files, and flat cache layout with Hugging Face repository downloads. - Introduce completion-metadata format 3, keyed by source identity and stored under `.onnx2webnn-validation`. - Record repository-relative paths and lengths while leaving downloaded model data in the standard Hugging Face cache. - Update validation documentation with the new cache contract and latest full-sweep evidence. ## Behavioral impact and compatibility - Full-model validation now defaults to the standard Hugging Face cache instead of `onnx2webnn/onnx`; existing flat-cache completion records are not reused and models may be downloaded once into the new layout. - `O2W_ONNX_CACHE` and `O2W_CACHE_DIR/onnx` remain supported but are interpreted as Hugging Face Hub cache roots. - `O2W_MODEL_CACHE_REFRESH=1`, manifest revision handling, primary ONNX SHA-256 verification, external-data discovery, and WebNN cache behavior remain supported. - The curated manifest’s recorded SHA-256 values are unchanged. - Ordinary `convert --output` caching remains unchanged. ## Validation - `cargo fmt --all -- --check` passed. - `cargo check --all-targets --locked` passed. - `cargo test --all-targets --locked` passed. - Strict all-target Clippy passed when allowing four pre-existing warnings in untouched converter files; unmodified strict Clippy remains blocked by those warnings. - The complete skeleton sweep passed all 51 cases. - The complete cold-cache real-weight sweep passed 47/51 cases in 10m43.6s with the same four documented FastVLM, Voxtral, and Chronos blockers and no download failures or skipped cases. - Tiny RoFormer passed after download, forced refresh, and offline cache reuse; an intentionally incorrect manifest digest was rejected. - Tiny RoFormer and RMBG both passed with an unreachable Hugging Face endpoint, confirming warm cache-only resolution; Tiny RoFormer also revalidated its manifest SHA-256.
## Summary RustNN changes to enable full roundtrip numeric validation of onnx2webnn conversion. Tied to [onnx2webnn #5](rustnn/onnx2webnn#5) and dependend on [webnn-graph #19](rustnn/webnn-graph#19). The reference in this PR needs to updated after the webnn-graph was merged. Make serialized WebNN graphs independently reloadable and executable while unifying graph recording and shape inference between `MLGraphBuilder` and the `.webnn` loader. Completed graphs now use an unambiguous shape model: `[]` always means a known rank-zero scalar, bounded dynamic dimensions remain explicit, and unresolved descriptors exist only as temporary internal inference state. ## Feature additions - Add `MLContext::rustnn_build_graph` as a direct compilation entry point for deserialized `GraphInfo`. This avoids constructing a second, unused `GraphRecorder` when compiling a graph reconstructed by the `.webnn` loader. - Add `run_onnx_path_with_inputs` so native ONNX Runtime execution can resolve external-data sidecars relative to the model file. This supports reference validation of filesystem-backed models whose weights are not embedded in the ONNX protobuf. - Serialize complete operation arguments, including reshape targets, Slice parameters, concat axes, permutations, and operand-valued options referenced by stable names. - Reload serialized graphs through the same recording and inference path used by `MLGraphBuilder`. - Add a versioned packed-4-bit Safetensors extension: - Logical `Int4` and `Uint4` dtype and shape remain in `.webnn`. - Packed low-nibble-first bytes are stored as Safetensors `U8`. - Archive metadata identifies the RustNN extension. - Reload validates marker, dtype, shape, byte length, and tensor-name resolution. - Support mixed ordinary and packed-4-bit constants in one Safetensors archive. ## Bugfixes - Infer and record every output of multi-output operations rather than only the first. - Prevent unresolved operands from silently becoming scalar descriptors. - Treat scalar GRU hidden states as rank-zero values and reject them through normal GRU rank validation. - Preserve scalar inputs, constants, intermediates, outputs, quantize/dequantize operands, and converter shape-map entries. - Always serialize required scalar `Reshape` and `Expand` targets as `newShape: []`. - Reject missing or malformed required shape-valued arguments explicitly. - Serialize operand-valued options by stable name instead of unstable numeric operand IDs. - Validate packed 4-bit logical element counts and exact packed storage lengths. - Correct Slice lowering for non-unit strides by deriving backend end indices from `start + extent`; this fixes both the ONNX and LiteRT lowering paths. ## Refactors - Introduce a private shared `GraphRecorder` used by both graph construction and JSON loading. - Centralize descriptor inference, operation insertion, dependency tracking, and graph-output marking. - Make operation insertion atomic: failed inference or validation no longer leaves partially recorded graph state. - Remove the historical empty-vector unknown-shape heuristic and obsolete temporary shape-table workaround. - Keep unresolved shapes represented through `Option` or missing internal entries rather than public `OperandDescriptor` values. - Update setup documentation with concise optional prerequisites for feature-specific backends. ## Behavioral impact and compatibility - `GraphRecorder` remains crate-private and is not a new public API. - `MLContext::rustnn_build_graph` and `run_onnx_path_with_inputs` are additive public entry points. - Every `OperandDescriptor` in completed `GraphInfo` now has a known shape: - `[]` is scalar. - Nonempty shapes may contain bounded `Dimension::Dynamic` entries. - Unknown shapes cannot escape graph construction. - The experimental `.webnn` format has intentional compatibility breaks: - Required shapes may no longer be omitted. - Scalar shape arguments must be serialized explicitly. - Direct input or constant graph outputs are rejected; an explicit operation such as Identity is required. - Ordinary Safetensors remain unchanged. Packed 4-bit tensors require the RustNN metadata marker and are rejected if presented as an unmarked or malformed extension. - Browser WebNN still does not natively execute 4-bit tensors. The ORT backend reconstructs an executable graph using its supported representations. ## Validation - Latest `cargo test --lib`: 379 tests passed. - Default, no-default-feature, `dynamic-inputs`, ONNX Runtime, and available backend-mock checks passed across the refactor. - Focused round trips passed for: - Scalar and bounded-dynamic descriptors. - Reshape, Expand, Slice, Concat, and Gemm bias. - Multi-output inference. - Even- and odd-sized `Int4` and `Uint4` constants. - Mixed native and packed-4-bit Safetensors. - Packed-4-bit graph compilation and ORT execution. - Malformed packed archive rejection. - Downstream onnx2webnn validation passed 743 tests and completed the 52-case skeleton and real-weight sweeps with the expected documented blockers. - Formatting and diff checks passed. - LiteRT regression coverage was added, but the LiteRT feature test could not run locally because `flatc` was unavailable. - Clippy remained blocked by pre-existing `-D warnings` failures outside this change.
## Feature additions - None. ## Bugfixes - None. ## Refactors - Update regular CI and full-model validation to consume merged `rustnn/rustnn:main` instead of the prerequisite fork branch. - Resolve `webnn-graph` from merged upstream and refresh RustNN’s transitive documentation dependencies in `Cargo.lock`. - Move the machine-readable model exclusion list to `docs/transformersjs_excluded_models.md` and update the manifest generator’s default path. - Document how the broad Transformers.js manifest is selected, when its current population was generated, and how repository/component exclusions are parsed. - Record an exclusions-disabled skeleton audit and clarify that NLLB decode builds while prefill remains blocked. ## Behavioral impact and compatibility - CI now follows upstream RustNN `main`; the merged upstream tree matched the former feature branch when switched. - The generated manifest remains unchanged at 51 cases across 28 repositories and 44 unique ONNX files. - Repository tooling now expects the exclusion list at `docs/transformersjs_excluded_models.md`. Explicit `--unsupported-file` callers remain supported, while external references to the former `tests/models/UNSUPPORTED_OPS.md` path must be updated. - Operator capability remains documented separately in `docs/operator-conversion-status.md`. ## Validation - `cargo check --all-targets --locked` passed. - `PYTHONPATH=scripts .venv/bin/python -m unittest scripts/test_generate_manifest.py` passed. - Confirmed the generator parses all 17 exclusions from the renamed document. - Ran an exclusions-disabled 26-case skeleton audit: 13 passed and 13 retained conversion/operator blockers; no excluded component became fully supported. - Verified all relative Markdown links resolve. - `git diff --check` passed.
| if path.exists() { | ||
| fs::remove_file(path).map_err(|e| format!("replace {}: {e}", path.display()))?; | ||
| } | ||
| fs::rename(&part, path) |
There was a problem hiding this comment.
add TOCTOU workaround like already one in webnn-graph for unlikely race condition that path is recreated between remove-file and rename.
| name = "webnn-graph" | ||
| version = "0.3.0" | ||
| source = "git+https://github.com/FelixKrall/webnn-graph?branch=fkrall%2Fpacked4-external-weights#067a86d93dfdd185ecfb8aaab855b112bb0064a3" | ||
| source = "git+https://github.com/rustnn/webnn-graph?branch=main#0bcd517d30cf740f55bf19386be254e37010ab2b" |
There was a problem hiding this comment.
should we really request this specific commit? this way we won't be able to make use of new versions without updating this repository.
There was a problem hiding this comment.
I see the point, it is in here because CI can run build with --locked to get reproducable results from onnx2webnn changes.
Summary
Accompanying RustNN PR #238 and webnn-graph PR #19. The refs in the cargo lock files of this PR need to be changed back to the respective main branches after these were merged
Add cache-backed numerical validation for converted models. Models can now be converted, exported as
.webnnplus Safetensors, reloaded independently, executed through RustNN’s CPU ORT backend, and compared with native ONNX Runtime using identical deterministic inputs.The manifest-wide skeleton sweep remains independent and unchanged in purpose. Full-model validation uses explicit
allormatch=<text>selection and supports real or deterministically generated weights.Feature additions
--output,--validate, and mutually exclusive--validate-cachedconversion options..onnx-cacheand exported artifacts under.webnn-cache.validate-modelscommand with:--selection all|match=<text>--weights real|generatedHF_TOKEN, SHA-256 verification, refresh support, and safe repository-relative paths..webnnand Safetensors artifacts without conversion-time state, dispatch deterministic inputs, and compare outputs with native ORT.attention_maskandtoken_type_idsvalues, and all converter-supported validation dtypes.Bugfixes
accuracy_levelvalues explicitly.Refactors
Behavioral impact and compatibility
.webnnartifacts before reloading them.--validate-cachedprovides a reload-only path.Ifmay require multiple independent.webnn/Safetensors pairs, potentially duplicating RAM or VRAM when loaded concurrently.FelixKrall/rustnn:fkrall/executable-webnn-reload; it should return torustnn/rustnnafter the prerequisite RustNN PR lands.Validation
cargo test --all-targets: 743 tests passed after the final conversion fixes.cargo test --lib: 213 tests passed.cargo test --test cache_validation: 16 tests passed.cargo check --all-targetspassed.