Skip to content

Repository files navigation

Suffeffix

An explainable lexical knowledge graph that jointly represents morphology, semantic decomposition, etymology and affix alignment — for Telugu, Hindi, English, French and Tamil: the T.H.E.F.T. framework.

Website: suffeffix.com · Technical report: v0.2 · Dataset: downloads · Version: 0.2.0 · Licences: code Apache-2.0, data CC BY-SA 4.0

Suffeffix takes words apart (stem, affix, and the function each affix performs), lines them up across five languages, decomposes their meaning into a small set of semantic atoms, and traces where each word came from. Every etymology carries a source and a confidence, and every explanation is generated by named rules with a trace back to the facts it used. There is no language model, embedding or AI inference anywhere in the system.


Contents

  1. What Suffeffix is — and is not
  2. The T.H.E.F.T. framework
  3. The five languages
  4. How a word is modelled — a worked example
  5. The data model
  6. Research-integrity rules (enforced by the validator)
  7. Explanations that show their working
  8. The dataset in numbers
  9. Repository layout
  10. Running it locally
  11. The API
  12. The website
  13. Build, test and deploy
  14. Contributing
  15. Known limitations
  16. The name
  17. Citing and licences

1. What Suffeffix is — and is not

It is:

  • a morphology explorer: every word is segmented into a stem and an ordered list of affixes, and each affix use names the job it does there (abstract state, agent, privative, causative…);
  • an affix-alignment framework: affixes that do the same job in different languages are grouped into equivalence classes, so English -ness, French -té, Hindi -आई, Telugu -తనం and Tamil -மை line up as functional equivalents;
  • a semantic decomposition system: each meaning is expressed as a small typed structure over at most fifty semantic atoms (STATE_OF(GOOD));
  • an etymology graph: typed, sourced, confidence-scored edges between words, historical forms and reconstructed roots;
  • explainable by construction: explanations are assembled from graph facts by deterministic template rules, each sentence carrying a trace.

It is not a translator, a chatbot, a language model, a dictionary clone, or a claim about how humans think. The semantic atoms are an engineering interlingua, not a theory of cognition.


2. The T.H.E.F.T. framework

Telugu · Hindi · English · French · Tamil. Five languages, two families — and between the families, only theft.

The acronym is a mnemonic with a thesis attached. Inside a language family, words are inherited. Across the family line, they can only be taken. Linguists say borrowing, but nothing is ever given back. That is the framework's one hard rule, and the validator enforces it: no inheritance or cognate edge may cross from Indo-European to Dravidian or back.

The five languages were chosen for contrast, not coverage. Each grouping isolates one relation a lexical graph has to get right:

Grouping What it isolates Example in the graph
Telugu · Tamil One family (Dravidian), two answers to a prestige donor ‘library’: Telugu గ్రంథాలయం (Sanskritic) against Tamil நூலகம் (a native coinage from நூல் ‘book’)
Telugu · Tamil Close cognacy within Dravidian మూడు / மூன்று ‘three’; కన్ను / கண் ‘eye’; నీరు / நீர் ‘water’
English · French Two branches (Germanic, Romance) joined by heavy borrowing Old French joie → English joy; beauty, merchant likewise
English · French · Hindi Distant cognacy across three Indo-European branches mother · mère · माता; name · nom · नाम; three · trois · तीन
Sanskrit → Hindi, Telugu, Tamil One donor crossing the family line in several directions pustaka → पुस्तक, పుస్తకం, புத்தகம்

Two classical donors. Sanskrit is to Hindi and Telugu roughly what Latin is to French and English: a prestige source of learned words. Indian grammatical tradition distinguishes tatsama words (borrowed unchanged, ‘same as that’) from tadbhava words (worn down by inheritance, ‘born of that’). French shows the same split in its learned and popular doublets — humanité with learned -ité beside bonté with inherited -té, both from Latin -itātem. Suffeffix records both kinds of layering in a single register field (S for the Sanskritic layer, L for the learned Latin and Greek layer).

In this dataset, 46% of Telugu words are Sanskritic against 18% of Tamil words, and 30% of English words against 28% of French words are learned. These figures describe the meanings that were selected, not the languages as a whole.

The choice of five languages is a design decision (ENGINEERING), not a linguistic claim. See docs/DECISIONS.md entry 13.


3. The five languages

Code Language Family Branch Script Words Affixes
en English Indo-European Germanic Latin 114 28
fr French Indo-European Romance Latin 113 34
hi Hindi Indo-European Indo-Aryan Devanagari 113 29
te Telugu Dravidian South-Central Telugu 111 28
ta Tamil Dravidian South Tamil 113 32

Everywhere words appear on the website, the lane order is fixed — English | French | Hindi ‖ Telugu | Tamil — with a dashed gutter marking the family boundary. The order is a claim, not a convenience.


4. How a word is modelled — a worked example

The meaning the state of being good in all five languages:

English French Hindi Telugu Tamil
Word goodness bonté अच्छाई (acchāī) మంచితనం (mañcitanaṁ) நன்மை (naṉmai)
Stem good bon अच्छा మంచి நல்
Affix -ness -té -आई -తనం -மை
Function fn:ST abstract state fn:ST fn:ST fn:ST fn:ST
Register native native native native native

All five words express one concept, concept:GOODNESS, whose meaning is decomposed as STATE_OF(GOOD) over the atom atom:GOOD. The five suffixes are members of the equivalence class eq:ST. The words are aligned to one another by relation: here all are TRANSLATION, while mère and mother would be COGNATE, and Telugu పుస్తకం and Tamil புத்தகம் are SHARED_LOAN.

A word can carry several affixes in order — French soigneusement is soin + -eux (POS) + -ment (ADV), like English care-ful-ly — and prefixes are first-class (in-, un-, निर्-, சிறு-).


5. The data model

The graph is stored as normalised JSON tables in data/, loaded into typed Pydantic v2 models (packages/core/suffeffix_core/schema/models.py) and indexed in memory. Every reference is by string identifier and checked by the validator. JSON Schemas for every model are exported to docs/schema/.

Node types

Type File What it is ID pattern
LexicalEntry entries/{lang}.json A word in one language: lemma (native script + transliteration), part of speech, register, concept, morphology, etymology edges, alignments lex:te:mancitanam
Affix affixes/{lang}.json An affix: form, kind, the functions it can perform, register, productivity, allomorphs, examples, cross-lingual equivalents affix:te:-tanam
AffixFunction affix_functions.json A cross-lingual job an affix can do fn:ST
EquivalenceClass equivalence_classes.json The affixes, across languages, that perform one function eq:ST
Concept concepts.json A meaning shared across languages, with its atom structure concept:GOODNESS
SemanticAtom atoms.json One of ≤ 50 semantic primitives, with an exponent in each language atom:GOOD
Root roots.json A reconstructed root (Proto-Indo-European or Proto-Dravidian) root:pie:mehter
EtymologyEdge etymology_edges.json A typed, sourced link between two historical forms edge:mother-pie-la
Source sources.json A bibliography entry; every source_ref must resolve here oed

IDs use ASCII transliterations so URLs stay portable; display forms keep full script and diacritics.

Enumerations

Register — which layer of the vocabulary a word or affix belongs to:

Value Meaning
N Native / inherited (Indian tadbhava; French mots populaires; the English Germanic stratum)
S Sanskritic (tatsama)
P Perso-Arabic
L Learned Latin or Greek stratum (English and French mots savants)
E Loan from English
mixed A hybrid of layers (e.g. a Sanskrit base with a native suffix)

Epistemic status — on every piece of linguistic content:

Value Meaning
ESTABLISHED Standard linguistic knowledge, citable from reference works
ENGINEERING An abstraction chosen for implementation convenience — not a linguistic claim
HYPOTHESIS Plausible and testable, not yet demonstrated
FUTURE Out of scope for this release

Review status (separate from epistemic status): draft → reviewed → published. Every record in v0.2 is draft.

Etymology edge types: INHERITED, COGNATE, BORROWED, CALQUE, DERIVED, COMPOUNDED, RECONSTRUCTED, REBORROWED. Each edge also has a status (accepted or contested), a confidence in [0, 1], one or more sources, and a list of semantic drift labels (NONE, NARROWING, WIDENING, METAPHOR, METONYMY, PEJORATION, AMELIORATION, BLEACHING, SPECIALISATION).

Alignment relations between entries of one concept: TRANSLATION, COGNATE, SHARED_LOAN, CALQUE.

Affix kinds: suffix, prefix, circumfix, compound_element. Productivity: high, mid, low, dead.

Atom structure operators: STATE_OF, CAUSE, BECOME, NEG, HAVE, AGENT_OF, EVENT_OF, PLACE_OF, DEGREE — stored as nested JSON ({"op": "STATE_OF", "args": [{"atom": "atom:GOOD"}]}), never as strings. Atom categories: FOUNDATIONAL, RELATIONAL, ACTION, STATE, EMOTIONAL, SOCIAL.

The twenty affix functions

ID Function ID Function
fn:ST abstract state fn:COLL collective / place
fn:AG agent fn:PEJ approximative / pejorative
fn:EV event / result noun fn:DOCT doctrine / -ism
fn:POS possessive / characterised-by fn:FEM feminine
fn:PRIV privative fn:FIELD field of study
fn:ABL ability / passive potential fn:INF verbal noun / infinitive
fn:ADJ relational adjective fn:COMP comparative
fn:ADV adverbialiser fn:HAB habitual agent
fn:CAUS causative fn:ORD ordinal
fn:DIM diminutive fn:NEGV negative verbaliser

Members of an equivalence class are functional equivalents: they differ in register, productivity and the bases they attach to, and swapping one for another usually produces nonsense. Every class says so in its note.

More detail: docs/DATASET.md (dataset design), docs/SEMANTICS.md (atoms and structures), docs/ARCHITECTURE.md (graph model).


6. Research-integrity rules (enforced by the validator)

scripts/validate.py runs suffeffix_core.validate and fails the build on any violation:

  1. The family boundary. No COGNATE or INHERITED edge may connect two top-level families. Reconstructed Proto-Indo-European counts as Indo-European and Proto-Dravidian as Dravidian. Telugu అమ్మ and Tamil அம்மா are cognates with each other; neither is linked to mother or माता, however alike they sound.
  2. Every edge has a source that exists in sources.json.
  3. Confidence caps. An edge supported only by Wiktionary is capped at 0.6. Reference works cited without an entry number (CDIAL, DEDR, the Tamil Lexicon) are held at 0.7 or below by policy — entry numbers are never invented.
  4. Contested means unresolved. Where scholarship disagrees (e.g. Hindi कुत्ता / Telugu కుక్క), the edge is stored contested with low confidence and badged wherever it appears.
  5. Every reference resolves — affixes, functions, concepts, atoms, roots, edges, aligned entries, examples, equivalents — and every ID matches its documented pattern.
  6. Affix uses are honest. An entry may only use an affix for a function the affix declares, and only an affix of its own language.
  7. Coverage is declared. Every concept is lexicalised in all five languages or flagged partial_coverage.
  8. Bounds. 300–850 entries, 30–50 atoms, 20–30 affix functions; no language outside the five.

7. Explanations that show their working

The explanation for a word is assembled by a deterministic template engine (explain.py) from the entry's morphology, its concept's atom structure, its etymology edges and its alignments. Each sentence comes from a named rule and carries a trace of the affixes, atoms, edges, rules and sources it used:

Rule Produces
rule:morphology-chain the stem and affix chain, with each affix's function
rule:simplex a statement that the word has no affixes
rule:atom-structure the meaning's atom structure (phrased “can be approximated as” when the structure is a hypothesis)
rule:etymology-hop each etymological step, with type, source and confidence; contested edges are flagged in the text
rule:alignment the aligned forms in the other languages and how they relate

The generated explanation of Telugu మంచితనం:

మంచితనం (mañcitanaṁ) is a Telugu noun formed from the stem మంచి with the suffix -తనం (-tanaṁ) (abstract state). Its meaning ('the state of being good') is represented as STATE_OF(GOOD) in the Suffeffix atom interlingua. Aligned forms: English goodness (translation equivalent); French bonté (translation equivalent); Hindi अच्छाई (acchāī) (translation equivalent); Tamil நன்மை (naṉmai) (translation equivalent).

The same input always gives the same output, and golden tests pin the output for ten entries. Free-text annotator notes are stored and shown separately, so a reader can always tell what the engine derived from what a person wrote.


8. The dataset in numbers

Count
Words (entries) 564
Meanings (concepts) 115 — 6 flagged as partially covered
Affixes 151
Affix functions / equivalence classes 20 / 13
Semantic atoms 50 (NSM primes plus engineering additions)
Etymology edges 127 — 73 inherited, 29 borrowed, 22 cognate, 3 derived; 6 contested
Reconstructed roots 18
Cited sources 13

Register profile of the selection:

Native Sanskritic Perso-Arabic Learned (L) English loan Mixed
English 78 — — 34 — 2
French 80 — — 32 1 —
Hindi 43 45 24 — — 1
Telugu 47 51 3 — — 10
Tamil 84 20 — — 2 7

How the meanings were chosen. The 115 meanings showcase derivation and cross-family contrast — abstract states, agents, privatives, possessives, ability adjectives, adverbs, causatives, doctrines, places and fields — plus a basic-vocabulary set (mother, name, three, water, sugar, king…) for etymology demonstrations. They are a selection, not a sample: statistics over this dataset describe the selection, not the languages.

Sources. The OED; the Trésor de la langue française informatisé (TLFi); Turner's Comparative Dictionary of the Indo-Aryan Languages (CDIAL); Burrow and Emeneau's Dravidian Etymological Dictionary (DEDR); the University of Madras Tamil Lexicon; Monier-Williams; Platts; Brown; McGregor; Krishnamurti, The Dravidian Languages; Watkins, American Heritage Dictionary of Indo-European Roots; Natural Semantic Metalanguage (Goddard and Wierzbicka); and Wiktionary (confidence-capped). Full citations are in data/sources.json.

Downloads. The release on suffeffix.com/data contains CSV tables (entries, affixes, etymology edges, atoms, concepts, sources; UTF-8 with a byte-order mark so spreadsheets open Indic scripts correctly), the full dataset as one JSON file, the JSON Schemas, a SHA256SUMS.txt and a zip of everything. The release is only produced if the dataset passes the full validator.


9. Repository layout

suffeffix/
├── data/                      the dataset — JSON, CC BY-SA 4.0 (data/LICENSE)
│   ├── atoms.json  affix_functions.json  equivalence_classes.json
│   ├── concepts.json  roots.json  etymology_edges.json  sources.json
│   ├── affixes/{en,fr,hi,te,ta}.json
│   └── entries/{en,fr,hi,te,ta}.json
├── packages/core/             pure Python domain library (no web framework)
│   ├── suffeffix_core/
│   │   ├── schema/            Pydantic models and enumerations
│   │   ├── load.py            JSON → typed Dataset
│   │   ├── index.py           in-memory graph indexes
│   │   ├── validate.py        the research-integrity rules
│   │   ├── search.py          script-agnostic search with transliteration folding
│   │   └── explain.py         the deterministic explanation engine
│   └── tests/                 validator, search, explanation goldens
├── apps/api/                  FastAPI read-only API over packages/core (+ smoke tests)
├── apps/web/                  Next.js 15 static site (App Router, TypeScript strict, Tailwind)
│   ├── app/                   routes (see §12)
│   ├── components/            triptych, concordance, atlas, figures, lane-row, …
│   └── lib/                   lang.ts (the five lanes), data access, census, citation
├── scripts/                   validate, export_site, export_downloads, export_schema,
│                              build_stats, check_links, check_brand
├── deploy/                    nginx config, server setup, deploy script, local preview server
├── brand/                     the ff symbol and wordmark
├── docs/                      ARCHITECTURE, DATASET, SEMANTICS, DECISIONS, API, FRONTEND,
│                              DEPLOY, ROADMAP, LAUNCH_CHECKLIST, schema/*.json
└── .github/workflows/         ci.yml, deploy-vps.yml

Dependencies point inward only: data/ → packages/core → apps/api → apps/web.


10. Running it locally

Requirements: Python 3.12, Node 20+ and pnpm, and GNU Make (or run the commands in the Makefile by hand).

# Python dependencies — from the repo root
pip install "pydantic>=2.7,<3" fastapi uvicorn pytest httpx

# Website dependencies
cd apps/web && pnpm install && cd ../..
Command What it does
make validate Validate the dataset; exits non-zero on any rule violation
make test Run the core and API test suites (30 tests)
make api Start the FastAPI server on http://localhost:8000 (OpenAPI at /openapi.json)
make site-data Export the API's responses for the static site, and build the dataset release into apps/web/public/data/
make web Start the Next.js dev server on http://localhost:3000 (run make site-data once first)
make site The full production build: site data, brand check, next build, then a check of every internal link
make preview Serve the built static site on http://localhost:3000 with nginx-like URL behaviour
make schema Re-export the JSON Schemas into docs/schema/
make stats Print dataset counts
make deploy HOST=deploy@your-server Build and deploy to a VPS (see §13)

The Makefile sets PYTHONPATH=packages/core;apps/api. If you run scripts by hand, set it yourself.


11. The API

A read-only REST API versioned under /v0. Responses are JSON; errors follow RFC 7807 (application/problem+json). There are no write endpoints.

Endpoint Returns
GET /v0/search?q=&lang= Ranked entries: exact form > exact transliteration > normalised transliteration > form prefix > gloss substring. Searching manchitanam finds మంచితనం; nanmai finds நன்மை.
GET /v0/entries?lang=&function=&affix=&atom=&register=&page= Filtered, paginated entries (50 per page)
GET /v0/entries/{id} A full entry: resolved affixes, concept, atoms, etymology edges, aligned forms, and the generated explanation with its trace
GET /v0/affixes?lang=&function= · GET /v0/affixes/{id} Affixes; the detail includes example words and cross-lingual equivalents
GET /v0/affix-functions All twenty functions
GET /v0/equivalence-classes Classes with resolved member affixes
GET /v0/atoms · GET /v0/atoms/{id} Atoms; the detail includes exponents, related atoms and the entries that use the atom
GET /v0/etymology/{entryId} The lineage subgraph: nodes with families, edges with type, drift, sources, confidence and status
GET /v0/concepts/{id} A concept with all its words and their alignment relations
GET /v0/meta Counts and epistemic/review-status breakdowns
GET /health Status, dataset counts and validation status

See docs/API.md.


12. The website

The public site is a fully static export. scripts/export_site.py replays the FastAPI app and saves its exact responses, and Next.js builds every page from them. The website, the API and the downloadable files therefore cannot disagree. The production site needs no Python and no database.

Route Page
/ Home: a live one-meaning-in-five-languages demonstration, the dataset in numbers, the THEFT framework, featured research, four ways into the graph
/research/suffeffix-v0-2/ The technical report: abstract, the THEFT framework (§2), schema, design constraints, interactive figures, explanations, etymology, validation, limitations, data statement, citations. /research/suffeffix-v0-1/ is kept as a superseded notice so old citations resolve.
/data/ Dataset release: every file with size, row count and checksum, the data card, and BibTeX
/lexicon/ Concordance: all 115 meanings × five lanes, filterable by text in any script, by affix function, or to rows derived in all five languages
/lexicon/{lang}/{slug}/ One word: its morphology, meaning, the same meaning in all five languages, etymology and its generated explanation with trace
/concepts/{slug}/ One meaning: the five-lane triptych, where affixes sit on the row of their function so equivalents align horizontally
/affixes/ · /affixes/{lang}/{slug}/ Affix Atlas: every affix filed by function, with register and productivity; each affix page lists example words and equivalents placed in their lanes
/atoms/ · /atoms/{name}/ The fifty semantic atoms, their exponents in all five languages, and the meanings built from each
/etymology/{lang}/{slug}/ Lineage graphs, coloured by family, so every change of colour is a borrowing
/about/ What Suffeffix is, the name, the T.H.E.F.T. framework, the two families, epistemic statuses, the data statement
/docs/ The project documentation, rendered

The site has a global search palette (⌘K / Ctrl+K) that accepts any of the four scripts or Latin transliteration, light and dark themes, a static OG image, a sitemap, and figures that each have an accessible table view. Type is the Anek family across Latin, Devanagari, Telugu and Tamil scripts. Design notes: docs/FRONTEND.md.


13. Build, test and deploy

Continuous integration (.github/workflows/ci.yml) runs on every push:

  • validates the dataset;
  • runs the Python tests;
  • checks the brand spelling;
  • typechecks and builds the site;
  • checks every internal link;
  • verifies the release checksums.

What make site verifies:

  1. The dataset passes the validator (the release export refuses to run otherwise).
  2. The wordmark is spelled suffeffix everywhere (scripts/check_brand.py).
  3. next build succeeds with strict TypeScript.
  4. Every internal link in the built site resolves (scripts/check_links.py; currently ~51,000 links across 984 pages).

Tests (make test, 30 in total) cover:

  • the validator, including deliberately malformed edges that must be rejected (a cross-family cognate, an over-confident Wiktionary-only edge);
  • search normalisation;
  • determinism, traces and golden outputs of the explanation engine;
  • every API route.

Deployment. The site is served by nginx on a VPS. The deploy-vps.yml workflow runs on every push to main (or manually):

  • It builds the site and rsyncs it into a fresh timestamped release directory.
  • It then flips a current symlink in one step, so visitors never see a half-uploaded site.
  • The last five releases are kept, so rolling back is instant.

The workflow needs the repository secrets VPS_HOST and VPS_SSH_KEY (and optionally VPS_USER). One-time server setup (DNS, deploy/setup-server.sh, Let's Encrypt) is documented step by step in docs/DEPLOY.md.


14. Contributing

The most valuable contribution is review: checking a draft record against a reference work and citing it.

Correcting or reviewing an entry

  1. Edit the JSON in data/, following docs/schema/ and docs/DATASET.md.
  2. Every etymology edge needs a source_ref that exists in data/sources.json. Wiktionary-only edges cap at confidence 0.6, and you must never invent CDIAL, DEDR or Tamil Lexicon entry numbers.
  3. Run make validate && make test.
  4. In the pull request, cite the reference work and page or entry, so a reviewer can promote review_status from draft to reviewed. Contested facts stay contested.

Adding a language. This is how French and Tamil were added in v0.2:

  1. Extend the Lang literal and AtomExponents in schema/models.py.
  2. Update LANGS in load.py and validate.py.
  3. Add data/affixes/xx.json and data/entries/xx.json.
  4. Add an exponent to every atom, and add the new affixes to the equivalence classes.
  5. On the website, add a lane to LANES in apps/web/lib/lang.ts; every grid renders from it.

Mind the family rule: a new language joins exactly one family.

See CONTRIBUTING.md. Report errors at github.com/VSSK007/suffeffix/issues.


15. Known limitations

  • Everything is a draft. Records were authored against reference works but have not been reviewed item by item. Treat them as hypotheses with citations. The French and Tamil records are newest.
  • Reference entry numbers are absent. CDIAL, DEDR and the Tamil Lexicon are cited without entry numbers until someone checks them against the printed volumes.
  • The selection is not a sample. 115 meanings chosen to show derivation say little about overall vocabulary, productivity or frequency.
  • Some analyses are judgement calls. Treating these as affixes is an engineering convenience, tagged ENGINEERING:
    • Telugu -వాడు and Tamil -அவர், rather than analysing them as bound pronouns;
    • Telugu -ఐన and Tamil -ஆன, rather than analysing them as participles.
  • Atom structures are coarse. Emotions and abstractions decompose only approximately into fifty atoms; many structures are marked HYPOTHESIS.
  • No sandhi and no frequency. Sandhi alternations are recorded only as allomorphs, and there is no corpus frequency.
  • Explanations are in English only, about all five languages.

What comes next — and what is deliberately left out — is in docs/ROADMAP.md. Every assumption that was challenged along the way is logged in docs/DECISIONS.md.


16. The name

Suffeffix means suffix — and every effing affix.

Open suffix between its stem and its ending (suff·ix) and drop eff into the middle: suff·eff·ix, a suffix with an infix inside it. It is a nod to English expletive infixation (abso-bloody-lutely, fan-effing-tastic; McCarthy 1982), but a playful cousin of it rather than an instance. The first half of the name says suffix; the second half is the real scope. Of the 151 affixes in v0.2, 23 are prefixes or compound elements that the name, taken literally, leaves out.

The wordmark is su·ff·e·ff·ix — two fs, twice — and the build checks the spelling.


17. Citing and licences

@techreport{suffeffix2026,
  title       = {Suffeffix: an explainable lexical knowledge graph for morphology,
                 semantic decomposition, etymology and affix alignment},
  author      = {{Suffeffix contributors}},
  institution = {Suffeffix},
  type        = {Technical report},
  number      = {v0.2},
  year        = {2026},
  month       = oct,
  url         = {https://suffeffix.com/research/suffeffix-v0-2/}
}

Corrections are welcome, ideally as an issue citing the reference work.

About

An explainable lexical knowledge graph: morphology, semantic decomposition, etymology, and affix alignment for English, Telugu, and Hindi.

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages