Repository navigation
Conversation
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #208
Summary
Adds support for parsing OASIS DMLex v1.0 JSON dictionaries alongside (or instead of) Wikidata lexeme dumps in
tools/dictionary-parser, with configurable multi-source conflict resolution.Key Changes
ParserOptions):--wikidata <file1> [<file2> ...]and--dmlex <file1> [<file2> ...](supporting multiple space- or comma-separated files as well as--wikidata=<files>/--dmlex=<files>, including.bz2compressed inputs).--merge-conflicts[={true|false}](defaultfalse): controls whether same-part-of-speech entries across multiple input files are overridden by later files (false) or merged (true).Grammar.TYPEMAPwith per-instance overrides (customTypeMapandignoredGrammemes) onParserOptions.DMLexEntry,DMLexInflectedForm,DMLexStringListDeserializer,ParseWikidata):parseDMLexStream) validating the top-level DMLex JSON object (langCodeandentries) and converting entries into the sharedLemma/Inflectionpipeline.headwordwheninflectedFormsis omitted.Grammarlookup (getMappedGrammemes) to normalize camelCase, snake_case, and hyphenated DMLex/Universal Dependencies grammatical tags in addition to Wikidata Q-IDs.ParseWikidata,DictionaryEntry):--merge-conflicts=false(default) and multiple source files are provided, stages lemmas and preempts any earlier(lemma, PartOfSpeech)paradigm superseded by a later source file before generating inflection patterns (inflectional.xml) or surface forms (dictionary.lst), preventing orphaned surface forms and inflated pattern counts.PartOfSpeechprovenance (sourceIndex) insideDictionaryEntryso same-POS surface-form collisions from a later source replace earlier ones while preserving non-conflicting parts of speech (e.g., keepinginterjectionfrom an earlier source when a later source addsnoun) and intra-file homographs.ParseWikidataTest):--merge-conflicts=false), cross-source union merging (--merge-conflicts=true), partial multi-POS overrides, space/comma-separated file lists, CLI option compatibility (--map-grammeme,--expand-grammemes,--ignore-property,--ignore-entries-with-grammemes,--exclude-language), and schema validation errors.