Skip to content

Milestone 9: Expand NAT architecture to additional Indic languages #46

Description

@vm75

Objective

Extend the proven Tamil NAT architecture to additional languages only after Tamil is production-stable, while keeping language-specific behavior isolated and the shared NAT API stable.

Initial candidates from the roadmap are Bengali and Hindi.

Principles

Do not generalize prematurely. Reuse abstractions that were proven necessary for Tamil, but avoid creating a large framework merely for hypothetical future languages.

Each language should provide:

  • language-specific ambiguity analysis
  • language-specific candidate/pronunciation logic
  • model asset(s)
  • model manifest
  • evaluation corpus
  • backend-independent model contract

Bengali focus areas

Evaluate and document cases such as:

  • inherent vowel / a vs ô behavior
  • ব pronunciation variation (including b/w-like contexts)
  • contextual pronunciation
  • Sanskrit-derived vocabulary
  • conjuncts and clusters
  • proper names

Do not encode Bengali behavior as Tamil-specific rules.

Hindi focus areas

Evaluate and document:

  • schwa deletion
  • contextual pronunciation
  • consonant clusters
  • Sanskrit-derived vocabulary
  • word/phrase context
  • proper names

Again, keep these rules/model behavior language-specific.

Requirements

  1. Reuse the shared NAT API.
  2. Reuse model manifest/versioning.
  3. Reuse evaluation tooling.
  4. Add one language at a time.
  5. Build a language-specific gold corpus before training/integration.
  6. Establish deterministic baseline measurements.
  7. Measure neural-only vs hybrid performance.
  8. Train or adapt a compact model only after identifying measurable baseline gaps.
  9. Keep each model independently versioned.
  10. Add cross-language regression tests without coupling language behavior.
  11. Do not change Tamil behavior unless a genuine shared architectural bug is discovered.
  12. Keep model downloads lazy and language-specific.

Acceptance criteria

  • At least one additional language has a complete evaluation pipeline.
  • The language has documented ambiguity/pronunciation categories.
  • A gold corpus exists.
  • Deterministic baseline is measured.
  • Neural/hybrid improvement is demonstrated before production integration.
  • Model is independently versioned and packaged.
  • Existing Tamil behavior remains unchanged.
  • Shared NAT API remains stable.
  • Language-specific code/model assets are isolated.
  • Documentation explains how another language can be added.

Out of scope

  • adding every Indic language at once
  • creating a generalized ML platform
  • changing the deterministic core solely to accommodate hypothetical languages

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    milestoneUmbrella milestone tracking issueroadmapPart of the long-term project roadmap

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions