Objective
Build a research/prototype evaluation pipeline for Tamil NAT and determine where neural assistance provides measurable value over the existing deterministic engine.
This milestone is about evidence and experimentation, not production integration.
Required comparisons
Measure at least:
- existing
inditrans alone
- neural model alone
inditrans + neural hybrid
The hybrid approach is the primary architectural direction.
Candidate baseline
Evaluate an existing Tamil-capable transliteration model such as AI4Bharat IndicXlit as a baseline/teacher candidate.
Do not assume its benchmark numbers apply to this project. Measure on the project's own Tamil corpus, especially literature, scripture, bhajans, Sanskrit-derived words, compounds, and pronunciation-sensitive cases.
The exact model may change if licensing, runtime, quality, or size makes another model more appropriate. Record the decision and evidence.
Requirements
- Build a reproducible evaluation script/tool.
- Use the gold corpus from Milestone 3.
- Record per-example outputs for:
- deterministic baseline
- neural-only
- hybrid
- Measure:
- exact accuracy where appropriate
- phoneme/pronunciation accuracy where applicable
- error category
- false corrections
- unchanged/correct baseline cases
- latency
- model size
- memory footprint where measurable
- Identify which ambiguity categories are improved by neural assistance.
- Identify categories where neural assistance makes the result worse.
- Measure paragraph/sentence behavior, not only isolated words.
- Preserve reproducibility: document model version, tokenizer, runtime, preprocessing, dataset version, and evaluation command.
- Keep external models/downloads out of the core package.
- Do not prematurely optimize for a model that has not demonstrated value.
Important evaluation principle
A neural system must not be allowed to silently replace correct deterministic output simply because it produces a different answer.
The hybrid evaluator should make it possible to answer:
- Was the baseline already correct?
- Did the neural system change it?
- Was the change correct?
- Which linguistic category caused the change?
- What is the false-correction rate?
Deliverables
- reproducible evaluation tool/script
- baseline results
- per-category error analysis
- model/runtime metadata
- recommendation for the smallest useful correction-model target
- documented go/no-go criteria for proceeding to Milestone 5
Acceptance criteria
Out of scope
- production browser integration
- production mobile integration
- committing large model binaries
- changing the deterministic engine based solely on anecdotal examples
Objective
Build a research/prototype evaluation pipeline for Tamil NAT and determine where neural assistance provides measurable value over the existing deterministic engine.
This milestone is about evidence and experimentation, not production integration.
Required comparisons
Measure at least:
inditransaloneinditrans + neuralhybridThe hybrid approach is the primary architectural direction.
Candidate baseline
Evaluate an existing Tamil-capable transliteration model such as AI4Bharat IndicXlit as a baseline/teacher candidate.
Do not assume its benchmark numbers apply to this project. Measure on the project's own Tamil corpus, especially literature, scripture, bhajans, Sanskrit-derived words, compounds, and pronunciation-sensitive cases.
The exact model may change if licensing, runtime, quality, or size makes another model more appropriate. Record the decision and evidence.
Requirements
Important evaluation principle
A neural system must not be allowed to silently replace correct deterministic output simply because it produces a different answer.
The hybrid evaluator should make it possible to answer:
Deliverables
Acceptance criteria
Out of scope