This repo contains a prototype machine translation tool for IATI documentation.
It uses a combination of deterministic logic and LLM calls to take an IATI docs repo to a state of having full translations to a consistent standard.
This software has been written by @robredpath with extensive use of Claude Code. It has been through ~7 rapid iterations via several dozen prompts.
This has been tested and refined using the IATI Publisher documentation. It has been run and refined repeatedly, but "works on my machine" caveats may apply.
It hasn't been tested on any other IATI documentation. It's likely that it will need further refinement as we start to work with a wider range of documentation.
The scripts have three fundamental sets of functionality that are combined in various ways to create useful tooling:
- Using an LLM to translate English text into other languages
- Using an LLM to review translation in the ways that LLMs are better at than deterministic code
- Using deterministic checks to review text in the ways that deterministic code is better at than LLMs
Clone the repo; create a venv; install the requirements Sign up for a Mistral API key. @robredpath can provide you with one if you'd like. It needs to be a paid account.
All scripts are run with the format `MISTRAL_API_KEY=XXXXX ./script.py /path/to/an/IATI/docs/repo
glossary.csv should always match the latest version of the IATI Glossary.
Each run also proposes recurring terms it finds on the site into scripts/project_glossary.csv in the target docs repo, with how they are currently translated. Set a row's status to accepted to enforce it on every future run, or rejected to stop it being proposed; commit that file with the docs.
translation_config.json contains few-shot examples and "translation notes" that are sent with every prompt to LLMs
The glossary is loaded from glossary.csv. Using the format provided by YI, a ui_terms.xlsx translation spreadsheet will be included with the glossary if provided in the target docs repo.
Reviews the English text to catch spelling, grammar or broken formatting issues before they propagate to the translation
Does whatever is needed to get an IATI docs repo to the state of being translated.
Exit status: 0 when every automated check passed; 1 when the run failed; 2 when it finished but some strings were left for manual attention (they are marked fuzzy and render in English until fixed — the summary says how).
Gives an overview of the current translation status of the repo
Read-only tool that reviews existing translations and reports issues found by the LLM reviewer. Site-wide by default; --per-file runs the per-file reviewer instead. Does not modify files — use translate.py to actually update translations.
There is a unit-test suite covering the deterministic checks (formatting/italics, list prefixes, URL language codes and integrity, role names and identifiers, leaked model output, glossary preservation, length ratios, revision safety), the PO utilities (obsolete stripping, lock/review markers), the LLM JSON parsing and structured-output fallback, the review gates (evidence, fragments, parse failures), and the orchestration (what gets stamped REVIEWED, what the summary reports). These don't call the API and run in under a second.
pip install -r requirements-dev.txt
pytest