Skip to content

Inflection-208 Add multisource and DMLex dictionary parsing support - #212

Open
nciric wants to merge 1 commit into
unicode-org:mainfrom
nciric:Inflection-208-dmlex-multisource-parser
Open

nciric wants to merge 1 commit into
unicode-org:mainfrom
nciric:Inflection-208-dmlex-multisource-parser

Conversation

@nciric

@nciric nciric commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

Fixes #208

Summary

Adds support for parsing OASIS DMLex v1.0 JSON dictionaries alongside (or instead of) Wikidata lexeme dumps in tools/dictionary-parser, with configurable multi-source conflict resolution.

Key Changes

  • Explicit Multi-Source CLI Flags (ParserOptions):
    • --wikidata <file1> [<file2> ...] and --dmlex <file1> [<file2> ...] (supporting multiple space- or comma-separated files as well as --wikidata=<files> / --dmlex=<files>, including .bz2 compressed inputs).
    • --merge-conflicts[={true|false}] (default false): controls whether same-part-of-speech entries across multiple input files are overridden by later files (false) or merged (true).
    • Replaced static in-place mutation of Grammar.TYPEMAP with per-instance overrides (customTypeMap and ignoredGrammemes) on ParserOptions.
  • OASIS DMLex v1.0 Deserialization & Streaming (DMLexEntry, DMLexInflectedForm, DMLexStringListDeserializer, ParseWikidata):
    • Streaming Jackson parser (parseDMLexStream) validating the top-level DMLex JSON object (langCode and entries) and converting entries into the shared Lemma / Inflection pipeline.
    • Supports both flat string tags and nested DMLex tag/label/pronunciation objects, synthesizing a default form from headword when inflectedForms is omitted.
    • Extended Grammar lookup (getMappedGrammemes) to normalize camelCase, snake_case, and hyphenated DMLex/Universal Dependencies grammatical tags in addition to Wikidata Q-IDs.
  • Two-Tier Cross-Source Conflict Resolution (ParseWikidata, DictionaryEntry):
    • Tier 1 (Lemma/Paradigm level): When --merge-conflicts=false (default) and multiple source files are provided, stages lemmas and preempts any earlier (lemma, PartOfSpeech) paradigm superseded by a later source file before generating inflection patterns (inflectional.xml) or surface forms (dictionary.lst), preventing orphaned surface forms and inflated pattern counts.
    • Tier 2 (Surface-form level): Tracks per-PartOfSpeech provenance (sourceIndex) inside DictionaryEntry so same-POS surface-form collisions from a later source replace earlier ones while preserving non-conflicting parts of speech (e.g., keeping interjection from an earlier source when a later source adds noun) and intra-file homographs.
  • Testing (ParseWikidataTest):
    • Added unit tests and test fixtures covering standalone DMLex parsing, multi-source Wikidata + DMLex override (--merge-conflicts=false), cross-source union merging (--merge-conflicts=true), partial multi-POS overrides, space/comma-separated file lists, CLI option compatibility (--map-grammeme, --expand-grammemes, --ignore-property, --ignore-entries-with-grammemes, --exclude-language), and schema validation errors.

@nciric
nciric requested a review from grhoten October 9, 2026 23:00

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add support for lexical dictionary format for custom vocabulary

1 participant