Skip to content

Milestone 4: Prototype neural model and establish evaluation baseline #41

Description

@vm75

Objective

Build a research/prototype evaluation pipeline for Tamil NAT and determine where neural assistance provides measurable value over the existing deterministic engine.

This milestone is about evidence and experimentation, not production integration.

Required comparisons

Measure at least:

  1. existing inditrans alone
  2. neural model alone
  3. inditrans + neural hybrid

The hybrid approach is the primary architectural direction.

Candidate baseline

Evaluate an existing Tamil-capable transliteration model such as AI4Bharat IndicXlit as a baseline/teacher candidate.

Do not assume its benchmark numbers apply to this project. Measure on the project's own Tamil corpus, especially literature, scripture, bhajans, Sanskrit-derived words, compounds, and pronunciation-sensitive cases.

The exact model may change if licensing, runtime, quality, or size makes another model more appropriate. Record the decision and evidence.

Requirements

  1. Build a reproducible evaluation script/tool.
  2. Use the gold corpus from Milestone 3.
  3. Record per-example outputs for:
    • deterministic baseline
    • neural-only
    • hybrid
  4. Measure:
    • exact accuracy where appropriate
    • phoneme/pronunciation accuracy where applicable
    • error category
    • false corrections
    • unchanged/correct baseline cases
    • latency
    • model size
    • memory footprint where measurable
  5. Identify which ambiguity categories are improved by neural assistance.
  6. Identify categories where neural assistance makes the result worse.
  7. Measure paragraph/sentence behavior, not only isolated words.
  8. Preserve reproducibility: document model version, tokenizer, runtime, preprocessing, dataset version, and evaluation command.
  9. Keep external models/downloads out of the core package.
  10. Do not prematurely optimize for a model that has not demonstrated value.

Important evaluation principle

A neural system must not be allowed to silently replace correct deterministic output simply because it produces a different answer.

The hybrid evaluator should make it possible to answer:

  • Was the baseline already correct?
  • Did the neural system change it?
  • Was the change correct?
  • Which linguistic category caused the change?
  • What is the false-correction rate?

Deliverables

  • reproducible evaluation tool/script
  • baseline results
  • per-category error analysis
  • model/runtime metadata
  • recommendation for the smallest useful correction-model target
  • documented go/no-go criteria for proceeding to Milestone 5

Acceptance criteria

  • Existing deterministic baseline is measured.
  • At least one Tamil neural baseline is evaluated.
  • Neural-only and hybrid outputs are compared.
  • Results are reproducible.
  • False corrections are explicitly measured.
  • Results are broken down by ambiguity category.
  • Model size and inference latency are recorded.
  • Licensing/provenance of the evaluated model is documented.
  • The evaluation does not alter production behavior.
  • A concrete specification for Milestone 5 is produced from measured results.

Out of scope

  • production browser integration
  • production mobile integration
  • committing large model binaries
  • changing the deterministic engine based solely on anecdotal examples

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    milestoneUmbrella milestone tracking issueroadmapPart of the long-term project roadmap

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions