Skip to content

Milestone 5: Train and integrate a compact Tamil correction model #42

Description

@vm75

Objective

Build the first project-specific neural model whose job is to resolve ambiguous/correctable cases produced by the deterministic inditrans pipeline, rather than replacing the entire transliteration engine.

The model should be small enough to be a realistic browser/PWA/mobile component.

Model philosophy

Prefer:

Tamil input
   ↓
inditrans baseline
   ↓
candidate/ambiguity extraction
   ↓
small neural correction model
   ↓
corrected pronunciation/transliteration

over:

Tamil input → large neural model → complete transliteration

The correction model should solve the smallest problem that produces measurable improvement.

Requirements

  1. Use Milestone 4 data to identify high-value ambiguity categories.
  2. Construct training examples from baseline-vs-gold differences.
  3. Keep correct baseline examples in the dataset so the model learns when not to change output.
  4. Support contextual input sufficient to resolve the ambiguity.
  5. Prefer span/candidate-level prediction over full-document generation when feasible.
  6. Keep model architecture simple and well documented.
  7. Establish a deterministic fallback for uncertain/low-confidence predictions.
  8. Define a confidence/acceptance threshold from validation data rather than guessing.
  9. Evaluate both improvement and regression:
    • baseline accuracy
    • hybrid accuracy
    • false-correction rate
    • coverage: percentage of inputs/cases actually changed
    • latency
    • model size
    • memory
  10. Keep model artifacts separate from application source code.
  11. Record model version independently from the library/package version.
  12. Document training data provenance and licensing.

Target constraints

These are engineering targets, not guaranteed requirements; benchmark before locking them:

  • model download ideally below ~20 MB
  • low tens of MB or less runtime memory where practical
  • short interactive inference latency
  • usable for paragraph-sized input
  • no mandatory server inference

If these targets conflict with accuracy, record the trade-off and optimize based on measurements.

Packaging

Use a model manifest containing at least:

  • model name
  • model version
  • language
  • input/output contract
  • tokenizer/preprocessing version
  • expected runtime
  • model file(s)
  • checksum if applicable
  • license/provenance
  • minimum compatible NAT API version

Acceptance criteria

  • A project-specific Tamil correction model exists.
  • Training/evaluation is reproducible.
  • Dataset generation from baseline-vs-gold differences is documented.
  • Correct baseline cases are represented.
  • Hybrid accuracy improves on the agreed target set without unacceptable regressions.
  • False-correction behavior is explicitly measured.
  • Confidence/fallback behavior is defined and tested.
  • Model size and latency are measured.
  • Model metadata/versioning is documented.
  • No model dependency is added to the deterministic core.

Out of scope

  • Web Worker integration
  • WebGPU
  • Flutter/mobile runtime integration
  • support for multiple languages
  • replacing the deterministic engine

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    milestoneUmbrella milestone tracking issueroadmapPart of the long-term project roadmap

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions