Skip to content

Milestone 6: Integrate NAT into browser/PWA with ONNX runtime #43

Description

@vm75

Objective

Make the Tamil NAT correction model usable in browsers and PWAs while keeping the UI responsive, startup cost low, and deterministic transliteration fully functional without NAT.

Runtime direction

Use ONNX as the model interchange/deployment format unless Milestone 5 provides strong evidence for a better option.

Browser runtime should support the conceptual fallback chain:

WebGPU-capable runtime
        ↓
WASM fallback
        ↓
deterministic inditrans-only fallback

Use a Web Worker (or equivalent background execution mechanism) so model loading and inference do not block the UI.

Requirements

  1. Convert the validated Milestone 5 model to ONNX if not already exported.
  2. Validate numerical/output equivalence against the reference model.
  3. Quantize only after establishing a floating-point reference.
  4. Benchmark quantized vs non-quantized accuracy and latency.
  5. Package the model as an independently versioned asset.
  6. Do not bundle a large model into the initial JS application payload if lazy loading can avoid it.
  7. Load the model only when NAT is requested/needed.
  8. Cache model assets using an appropriate browser mechanism.
  9. Use a Worker for model initialization and inference.
  10. Batch/debounce requests where appropriate; do not run expensive inference on every keystroke by default.
  11. Provide deterministic fallback if:
    • model download fails
    • model initialization fails
    • WebGPU is unavailable
    • WASM runtime fails
    • inference fails
  12. Preserve the existing synchronous deterministic transliteration API.
  13. Expose NAT through an explicitly optional async API.
  14. Ensure repeated calls do not repeatedly initialize the model.
  15. Add cancellation/lifecycle handling where practical.

Performance measurements

Measure on representative browsers/devices:

  • initial model download size
  • cached startup time
  • cold initialization time
  • warm inference latency
  • paragraph latency
  • peak memory where measurable
  • WebGPU vs WASM behavior
  • fallback behavior

Do not make performance claims without measurements.

Acceptance criteria

  • Production-capable ONNX artifact exists.
  • Browser inference works offline after model caching.
  • Inference runs outside the main UI thread.
  • Model loading is lazy.
  • Model is cached.
  • WebGPU path works where supported.
  • WASM fallback works where WebGPU is unavailable.
  • Deterministic fallback works when NAT is unavailable.
  • Existing non-NAT browser behavior remains unchanged.
  • Quantized model is validated against the reference.
  • Performance measurements are documented.
  • Model versioning and compatibility are explicit.

Out of scope

  • Flutter/mobile runtime
  • Node runtime
  • Bengali/Hindi
  • automatic model downloading from an uncontrolled third-party server

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    milestoneUmbrella milestone tracking issueroadmapPart of the long-term project roadmap

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions