Objective
Extend the proven Tamil NAT architecture to additional languages only after Tamil is production-stable, while keeping language-specific behavior isolated and the shared NAT API stable.
Initial candidates from the roadmap are Bengali and Hindi.
Principles
Do not generalize prematurely. Reuse abstractions that were proven necessary for Tamil, but avoid creating a large framework merely for hypothetical future languages.
Each language should provide:
- language-specific ambiguity analysis
- language-specific candidate/pronunciation logic
- model asset(s)
- model manifest
- evaluation corpus
- backend-independent model contract
Bengali focus areas
Evaluate and document cases such as:
- inherent vowel / a vs ô behavior
- ব pronunciation variation (including b/w-like contexts)
- contextual pronunciation
- Sanskrit-derived vocabulary
- conjuncts and clusters
- proper names
Do not encode Bengali behavior as Tamil-specific rules.
Hindi focus areas
Evaluate and document:
- schwa deletion
- contextual pronunciation
- consonant clusters
- Sanskrit-derived vocabulary
- word/phrase context
- proper names
Again, keep these rules/model behavior language-specific.
Requirements
- Reuse the shared NAT API.
- Reuse model manifest/versioning.
- Reuse evaluation tooling.
- Add one language at a time.
- Build a language-specific gold corpus before training/integration.
- Establish deterministic baseline measurements.
- Measure neural-only vs hybrid performance.
- Train or adapt a compact model only after identifying measurable baseline gaps.
- Keep each model independently versioned.
- Add cross-language regression tests without coupling language behavior.
- Do not change Tamil behavior unless a genuine shared architectural bug is discovered.
- Keep model downloads lazy and language-specific.
Acceptance criteria
Out of scope
- adding every Indic language at once
- creating a generalized ML platform
- changing the deterministic core solely to accommodate hypothetical languages
Objective
Extend the proven Tamil NAT architecture to additional languages only after Tamil is production-stable, while keeping language-specific behavior isolated and the shared NAT API stable.
Initial candidates from the roadmap are Bengali and Hindi.
Principles
Do not generalize prematurely. Reuse abstractions that were proven necessary for Tamil, but avoid creating a large framework merely for hypothetical future languages.
Each language should provide:
Bengali focus areas
Evaluate and document cases such as:
Do not encode Bengali behavior as Tamil-specific rules.
Hindi focus areas
Evaluate and document:
Again, keep these rules/model behavior language-specific.
Requirements
Acceptance criteria
Out of scope