Skip to content

CoreML backend fails on larger hub demos despite small graph coverage passing #19

Description

@tarekziade

Summary

CoreML is available and works for small/static pywebnn graphs, but larger Hugging Face Hub demos expose CoreML converter/runtime coverage gaps.

This came up while adding explicit backend coverage. The default test suite now has CPU/ONNX and CoreML smoke coverage, and the supported CoreML cases pass. The larger demos still fail when run with an explicit CoreML backend.

Passing baseline

Small CoreML graphs compile and execute:

.venv/bin/python tests/test_coreml_basic.py --cleanup
.venv/bin/python -m pytest tests/test_performance.py -k coreml -v
.venv/bin/python -m pytest tests/test_backend_coverage.py -v

Observed focused coverage result:

6 passed, 4 xfailed

Supported CoreML cases covered there include:

  • add + relu execution
  • comparison execution with WebNN uint8 boolean outputs
  • quantizeLinear with a constant zero point

Repro 1: MobileNetV2 CoreML demo

ORT_DYLIB_PATH="$(.venv/bin/python tools/resolve_ort_dylib.py)" \
  .venv/bin/python examples/mobilenetv2_from_hub.py examples/images/test.jpg --backend coreml

The graph loads and the CoreML context is created:

Graph loaded successfully
Operand count: 261
Operation count: 154
Inputs: [input]
Outputs: [output]
Context created (accelerated=True)

Failure occurs when compiling/running the graph:

RuntimeError: Failed to build graph: coreml runtime failed: in-memory model load failed
compiler error: Encountered an error while compiling a model: in operation of type conv:
Param x has incorrect type for operator ios17.conv.
Expected { tensor<fp32, [?, ?, ...]>, tensor<fp16, [?, ?, ...]> }; got tensor<fp32, [1]>.

Likely direction: investigate loaded graph metadata/shape materialization on the CoreML conversion path. A convolution input appears to collapse to shape [1].

Repro 2: SmolLM CoreML demo

ORT_DYLIB_PATH="$(.venv/bin/python tools/resolve_ort_dylib.py)" \
  .venv/bin/python examples/smollm_from_hub.py --backend coreml --max-new-tokens 1

The graph loads and the CoreML context is created:

Operands: 5140
Operations: 2814
Inputs: 63
Outputs: 61
Context created (accelerated=True)
Backend requested: coreml

Failure occurs during graph conversion/build for dispatch:

RuntimeError: Failed to build graph: graph conversion failed for coreml_mlprogram: Unsupported operation: shape

Likely direction: add CoreML MLProgram lowering for WebNN shape, or reject/fallback earlier with a clearer unsupported-op diagnostic.

Other CoreML gaps now tracked in tests

The new strict xfails in tests/test_backend_coverage.py document additional CoreML limitations:

  • logical ops with WebNN numeric truthiness inputs
  • logical ops chained from comparison outputs currently emit a duplicate bool temporary name
  • where / CoreML select with integer conditions
  • quantize/dequantize with dynamic zero_point

Expected behavior

The demos should either run successfully on --backend coreml, or fail during capability/graph validation with a clear unsupported operation or unsupported pattern message before CoreML model compilation.

Environment

  • macOS arm64
  • Python 3.11
  • pywebnn built with onnx-runtime,coreml-runtime
  • rustnn dependency from main

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions