Objective
Build the first project-specific neural model whose job is to resolve ambiguous/correctable cases produced by the deterministic inditrans pipeline, rather than replacing the entire transliteration engine.
The model should be small enough to be a realistic browser/PWA/mobile component.
Model philosophy
Prefer:
Tamil input
↓
inditrans baseline
↓
candidate/ambiguity extraction
↓
small neural correction model
↓
corrected pronunciation/transliteration
over:
Tamil input → large neural model → complete transliteration
The correction model should solve the smallest problem that produces measurable improvement.
Requirements
- Use Milestone 4 data to identify high-value ambiguity categories.
- Construct training examples from baseline-vs-gold differences.
- Keep correct baseline examples in the dataset so the model learns when not to change output.
- Support contextual input sufficient to resolve the ambiguity.
- Prefer span/candidate-level prediction over full-document generation when feasible.
- Keep model architecture simple and well documented.
- Establish a deterministic fallback for uncertain/low-confidence predictions.
- Define a confidence/acceptance threshold from validation data rather than guessing.
- Evaluate both improvement and regression:
- baseline accuracy
- hybrid accuracy
- false-correction rate
- coverage: percentage of inputs/cases actually changed
- latency
- model size
- memory
- Keep model artifacts separate from application source code.
- Record model version independently from the library/package version.
- Document training data provenance and licensing.
Target constraints
These are engineering targets, not guaranteed requirements; benchmark before locking them:
- model download ideally below ~20 MB
- low tens of MB or less runtime memory where practical
- short interactive inference latency
- usable for paragraph-sized input
- no mandatory server inference
If these targets conflict with accuracy, record the trade-off and optimize based on measurements.
Packaging
Use a model manifest containing at least:
- model name
- model version
- language
- input/output contract
- tokenizer/preprocessing version
- expected runtime
- model file(s)
- checksum if applicable
- license/provenance
- minimum compatible NAT API version
Acceptance criteria
Out of scope
- Web Worker integration
- WebGPU
- Flutter/mobile runtime integration
- support for multiple languages
- replacing the deterministic engine
Objective
Build the first project-specific neural model whose job is to resolve ambiguous/correctable cases produced by the deterministic
inditranspipeline, rather than replacing the entire transliteration engine.The model should be small enough to be a realistic browser/PWA/mobile component.
Model philosophy
Prefer:
over:
The correction model should solve the smallest problem that produces measurable improvement.
Requirements
Target constraints
These are engineering targets, not guaranteed requirements; benchmark before locking them:
If these targets conflict with accuracy, record the trade-off and optimize based on measurements.
Packaging
Use a model manifest containing at least:
Acceptance criteria
Out of scope