Objective
Make the Tamil NAT correction model usable in browsers and PWAs while keeping the UI responsive, startup cost low, and deterministic transliteration fully functional without NAT.
Runtime direction
Use ONNX as the model interchange/deployment format unless Milestone 5 provides strong evidence for a better option.
Browser runtime should support the conceptual fallback chain:
WebGPU-capable runtime
↓
WASM fallback
↓
deterministic inditrans-only fallback
Use a Web Worker (or equivalent background execution mechanism) so model loading and inference do not block the UI.
Requirements
- Convert the validated Milestone 5 model to ONNX if not already exported.
- Validate numerical/output equivalence against the reference model.
- Quantize only after establishing a floating-point reference.
- Benchmark quantized vs non-quantized accuracy and latency.
- Package the model as an independently versioned asset.
- Do not bundle a large model into the initial JS application payload if lazy loading can avoid it.
- Load the model only when NAT is requested/needed.
- Cache model assets using an appropriate browser mechanism.
- Use a Worker for model initialization and inference.
- Batch/debounce requests where appropriate; do not run expensive inference on every keystroke by default.
- Provide deterministic fallback if:
- model download fails
- model initialization fails
- WebGPU is unavailable
- WASM runtime fails
- inference fails
- Preserve the existing synchronous deterministic transliteration API.
- Expose NAT through an explicitly optional async API.
- Ensure repeated calls do not repeatedly initialize the model.
- Add cancellation/lifecycle handling where practical.
Performance measurements
Measure on representative browsers/devices:
- initial model download size
- cached startup time
- cold initialization time
- warm inference latency
- paragraph latency
- peak memory where measurable
- WebGPU vs WASM behavior
- fallback behavior
Do not make performance claims without measurements.
Acceptance criteria
Out of scope
- Flutter/mobile runtime
- Node runtime
- Bengali/Hindi
- automatic model downloading from an uncontrolled third-party server
Objective
Make the Tamil NAT correction model usable in browsers and PWAs while keeping the UI responsive, startup cost low, and deterministic transliteration fully functional without NAT.
Runtime direction
Use ONNX as the model interchange/deployment format unless Milestone 5 provides strong evidence for a better option.
Browser runtime should support the conceptual fallback chain:
Use a Web Worker (or equivalent background execution mechanism) so model loading and inference do not block the UI.
Requirements
Performance measurements
Measure on representative browsers/devices:
Do not make performance claims without measurements.
Acceptance criteria
Out of scope