Skip to content

Repository files navigation

esTS

esTS

¿Cómo esTáS, texto?

Spanish Texts Statistics - a library for statistics extraction from texts in Spanish

Demo · Documentation · PyPI · Español

Version Supported Python versions Build Ruff License Demo on Hugging Face Spaces DOI


esTS computes statistics of Spanish texts: basic statistics, readability, lexical diversity, lexical sophistication, style, phonostatistics, morphology, syntax and cohesion - by published formulas with the coefficients and the scales of their authors, and by the parts of speech and the features of Universal Dependencies.

The library works both with raw strings and with Doc objects of spaCy; most statistics need no trained model (see Installation). Try it without installing in the demo on Hugging Face Spaces: paste a text and get its readability, the metrics, the plots and the highlighting of its fragments.

  • Object extraction - configurable sentence, word and character N-gram tokenizers that know the inverted marks, the dialogue dash and the abbreviations of Spanish
  • Syllables and stress - rule-based syllabification and the stressed syllable derived from the spelling, with no dictionary
  • Basic statistics - counts of sentences, words, letters, syllables and punctuation marks by type, with distributions and normalized shares
  • Readability metrics - Fernández Huerta, Szigriszt-Pazos with the INFLESZ scale, Gutiérrez de Polini, Crawford, Legibilidad µ, SOL, LIX and RIX, with a consensus grade, the school stages of Spain and reading time
  • Lexical diversity metrics - TTR and its variations, MATTR, MSTTR, MTLD, HD-D, Simpson's and Yule's indices, entropy, Zipf's and Heaps' laws
  • Morphological statistics - parts of speech and fifteen grammatical features of Universal Dependencies, with the markers of Spanish: the moods, the non-finite forms, ser against estar, the adverbs in -mente
  • Corpus measures - keywords against a reference corpus, collocations, the dispersion of a word over the parts of a text, a KWIC concordance and the stylometry of authorship: Burrows's Delta with its variants, Zeta, the Mendenhall curve, the profile of the function words; the comparison of two corpora by 132 features of a text with effect sizes
  • Visualizers - Zipf's law, literature fingerprinting, a word tree, lexical dispersion and keywords, a network of collocations, a dendrogram, PCA and MDS by Delta, vocabulary growth, sentence lengths, and the highlighting of the fragments the statistics count, such as long sentences and passives
  • Datasets - Spanish-language literature in the public domain: 150 works by 33 authors in four genres; 4,259 sonnets of the 15th-20th centuries with the metrical pattern and the rhyme of every line; a frequency dictionary of 83,785 lemmas by Google Books Ngram
  • spaCy components - every statistics class as a component of a pipeline, the statistics attached to the Doc in one pass
  • Cohesion statistics - the overlap of nouns, arguments and content words between sentences, givenness and temporal cohesion in the manner of Coh-Metrix, with the density of 255 Spanish discourse markers
  • Lexical sophistication statistics - how rare the words of a text are in the language: the frequency, range and dispersion of the lemmas by a dictionary of Google Books Ngram, the frequency bands top-1000 to 10000, surprisal, perplexity and lexical density
  • Style metrics - the SEO indicators of Advego and Text.ru (nausea, water content, spam score, naturalness by Zipf's law, keyword density) and the markers of the officialese style by the Spanish guides to plain language: verbal nouns, compound prepositions, parenthetical expressions and clichés
  • Phonostatistics - the shares of the classes of sounds, consonant clusters, hiatuses, open syllables, hardness and the indices of alliteration and assonance, over the sounds of a rule-based transcription
  • Verse statistics - the scansion of Spanish verse by its syllabic meter: the metrical syllables, the meter of a poem, the stress profile and the types of the endecasílabo, the rhyme in full and by assonance with its scheme, the strophes and the form of a poem
  • Syntactic statistics - the dependency tree by distances, depth, clauses and coordination, with the constructions of the administrative style: the passive with ser and with se, the participial and the gerund clauses, the chains of de, the split predicates

Installation

Requires Python 3.11 or newer. The language-independent part - the extractors, the basic statistics and the common readability formulas, the lexical diversity metrics, the corpus measures and the plots - comes from the anyTS core, installed with the package.

pip install pyests

Or with uv:

uv add pyests

The distribution on PyPI is pyests, the package it installs is ests. The basic statistics, readability, lexical diversity, the style metrics except the verbal nouns, phonostatistics and verse statistics need no spaCy model. The morphological, syntactic, cohesion and lexical sophistication statistics of a string need one, and so do the verbal nouns, the profile of the function words, the features of a text, the comparison of corpora and parsing a Doc yourself:

python -m spacy download es_core_news_sm

The statistics by the frequency dictionary need it downloaded once with FreqDict().download(). The datasets go to the directory ests_data next to the installed package, or to the one passed in data_dir or set in ESTS_DATA_DIR before the package is imported.

Quick start

>>> from ests import BasicStats, DiversityStats, ReadabilityStats

>>> text = "Hay tres clases de mentiras: mentiras, malditas mentiras y estadísticas"

>>> BasicStats(text).get_stats()
{'c_letters': {1: 1, 2: 1, 3: 1, 4: 1, 6: 1, 8: 4, 12: 1},
 'c_syllables': {1: 4, 2: 1, 3: 4, 5: 1},
 'n_sents': 1,
 'n_words': 10,
 'n_unique_words': 8,
 'n_long_words': 5,
 'n_complex_words': 5,
 'n_simple_words': 5,
 'n_monosyllable_words': 4,
 'n_polysyllable_words': 6,
 'n_chars': 71,
 'n_letters': 60,
 'n_spaces': 9,
 'n_syllables': 23,
 'n_punctuations': 2,
 'c_punctuations': {'comma': 1, 'period': 0, 'question': 0, 'exclamation': 0,
                    'ellipsis': 0, 'colon': 1, 'semicolon': 0, 'dash': 0,
                    'hyphen': 0, 'angle_quotes': 0, 'straight_quotes': 0,
                    'parentheses': 0, 'other': 0}}

>>> ReadabilityStats(text).flesch_reading_easy
53.545000000000016

>>> DiversityStats(text).ttr
0.8

Features

Object extraction

Configurable tools extract sentences, words and character N-grams from a text for the statistics. The sentence splitter knows the inverted marks, the dialogue dash and the abbreviations of Spanish; the word tokenizer keeps clitics, ordinals and Spanish numbers whole, and lemmas come from simplemma, which needs no model.

>>> from ests import CharNgramsExtractor, SentsExtractor, WordsExtractor

>>> SentsExtractor().extract("—¿Vienes? —preguntó María. ¡Claro que sí!")
('—¿Vienes? —preguntó María.', '¡Claro que sí!')

>>> we = WordsExtractor(use_lexemes=True, stopwords=["de", "y"], filter_nums=True, ngram_range=(1, 2))
>>> we.extract("Hay 3 clases de mentiras y estadísticas")
('haber', 'clase', 'mentira', 'estadística', 'haber_clase', 'clase_mentira', 'mentira_estadística')

>>> CharNgramsExtractor(n=3, lowercase=True).extract("estadísticas")[:5]
('est', 'sta', 'tad', 'adí', 'dís')

More in the documentation.

Syllables and stress

The syllables and the stress follow from the Spanish spelling, with no dictionary: the written accent, or the ending of the word when there is none, gives the stressed syllable. Adverbs in -mente and hyphenated compounds carry two stresses.

>>> from ests.syllables import stress_type, syllabify, word_stress, word_stresses

>>> syllabify("murciélago")
['mur', 'cié', 'la', 'go']

>>> syllabify("averiguáis")
['a', 've', 'ri', 'guáis']

>>> word_stress("construir"), stress_type("construir")
(1, 'aguda')

>>> word_stresses("fácilmente")
[0, 2]

More in the documentation.

Basic statistics

The library allows extracting the following statistics from a text:

  • the number of sentences
  • the number of words
  • the number of unique words
  • the number of long words
  • the number of complex words
  • the number of simple words
  • the number of monosyllable words
  • the number of polysyllable words
  • the number of characters
  • the number of letters
  • the number of spaces
  • the number of syllables
  • the number of punctuation marks and their distribution by type
  • the distribution of words by the number of letters
  • the distribution of words by the number of syllables

A complex word has three or more syllables and a long word seven or more letters. Any statistic can be printed in a readable form:

>>> from ests import BasicStats

>>> text = "Hay tres clases de mentiras: mentiras, malditas mentiras y estadísticas"
>>> BasicStats(text).print_stats()
     Statistic      |  Value
------------------------------
Sentences           |    1
Words               |    10
Unique words        |    8
Long words          |    5
Complex words       |    5
Simple words        |    5
Monosyllabic words  |    4
Polysyllabic words  |    6
Characters          |    71
Letters             |    60
Spaces              |    9
Syllables           |    23
Punctuation marks   |    2

More in the documentation.

Readability metrics

The library allows counting the following readability metrics:

  • Flesch reading ease with the coefficients of Szigriszt-Pazos or of Fernández Huerta
  • Gutiérrez de Polini comprehensibility formula
  • Crawford grade
  • Legibilidad µ
  • SOL grade, the SMOG index converted to Spanish
  • LIX readability measure
  • RIX readability measure

On top of the formulas: the band of a scale for the reading ease and for Legibilidad µ, a consensus grade (the median of the grade formulas), the school stage and the reader age in Spain, and reading time.

The coefficients of the Flesch reading ease are selected by preset: by default the fórmula de perspicuidad of Szigriszt-Pazos with the INFLESZ scale (general), or Fernández Huerta with the bands of its author (classic).

>>> from pprint import pprint
>>> from ests import ReadabilityStats

>>> text = "Hay tres clases de mentiras: mentiras, malditas mentiras y estadísticas"
>>> rs = ReadabilityStats(text)

>>> pprint(rs.get_stats(), sort_dicts=False)
{'flesch_reading_easy': 53.545000000000016,
 'gutierrez_polini_index': 33.5,
 'crawford_grade': 5.812999999999999,
 'mu_index': 50.943396226415096,
 'sol_grade': 9.258359866374562,
 'lix': 60.0,
 'rix': 5.0,
 'consensus_grade': 9.0,
 'reading_time': 0.03597122302158273}

>>> rs.print_stats()
                   Metric                    |  Value
-------------------------------------------------------
Flesch reading ease (Szigriszt-Pazos)        |  53.55
Gutiérrez de Polini comprehensibility        |  33.50
Crawford grade                               |   5.81
Legibilidad µ                                |  50.94
SOL grade (SMOG for Spanish)                 |   9.26
LIX readability index                        |  60.00
RIX readability index                        |   5.00
Consensus grade                              |   9.00
Reading time (min)                           |   0.04

>>> rs.describe_level()
'algo difícil'

>>> rs.describe_grade()
'ESO (12-16 years)'

More in the documentation.

Lexical diversity metrics

The library allows counting 32 lexical diversity metrics, among them:

  • Type-Token Ratio and its variations: RTTR, CTTR, Herdan, Summer, Maas, Dugast
  • Moving Average Type-Token Ratio and Mean Segmental Type-Token Ratio
  • Measure of Textual Lexical Diversity and its moving-window variants MA-MTLD and MTLD-W
  • Hypergeometric Distribution D
  • Simpson's index, its reciprocal and the Gini-Simpson index
  • the hapax index (Honoré's R), the measures of Yule, Herdan, Sichel, Michéa, Brunet, Dugast and Baayen
  • Shannon entropy, evenness and perplexity
  • the slope of Zipf's law, the Zipf-Mandelbrot fit and the exponent of Heaps' law

Any metric can be computed over windows of equal length, to compare texts of different lengths, with a confidence interval of the mean.

>>> from ests import DiversityStats

>>> text = "Hay tres clases de mentiras: mentiras, malditas mentiras y estadísticas"
>>> ds = DiversityStats(text)

>>> ds.ttr, ds.mtld, ds.yule_k
(0.8, 14.000000000000004, 600.0)

>>> ds.frequency_spectrum
{1: 7, 3: 1}

>>> DiversityStats("La legibilidad de un texto depende de la longitud de sus oraciones y de sus "
...                "palabras. Las fórmulas clásicas miden esas dos magnitudes y las combinan en "
...                "un solo número. Ninguna de ellas mide la comprensión: miden la superficie "
...                "del texto.").windowed("ttr", window_len=10)
WindowStats(mean=0.875, std=0.1258305739211792, lower=0.6747754774664074, upper=1.0752245225335926, n_windows=4)

More in the documentation.

Morphological statistics

The parts of speech and the grammatical features of Universal Dependencies, as the Spanish models of spaCy give them:

  • the part of speech and fifteen features: case, definiteness, degree, gender, mood, numeral type, number, person, polarity, politeness, possessive, pronoun type, reflexive, tense and verb form
  • the distribution of the words by the values of any feature, and the parse of the text word by word
  • the markers of Spanish: the moods among the finite forms, the non-finite forms, ser against estar, the adverbs in -mente
>>> from ests import MorphStats

>>> ms = MorphStats("Si tuviera tiempo, leería el libro que me recomendaste ayer")

>>> ms.get_stats("mood", "tense", filter_none=True)
{'mood': {'Sub': 1, 'Cnd': 1, 'Ind': 1}, 'tense': {'Imp': 1, 'Pres': 1}}

>>> ms.tags[1]
'Mood=Sub|Number=Sing|Person=3|Tense=Imp|VerbForm=Fin'

>>> ms.explain_text("pos", "mood", filter_none=True)[1]
('tuviera', {'pos': 'VERB', 'mood': 'Sub'})

>>> MorphStats("Ella es alta pero hoy está cansada y habla lentamente").get_markers()["p_ser"]
0.5

A text is parsed with es_core_news_sm; another pipeline can be passed in nlp.

More in the documentation.

Syntactic statistics

The dependency tree of Universal Dependencies and the constructions that the Spanish guides to clear language warn about:

  • the complexity of the tree: dependency distances, depth, leaves and subtrees, valency of the finite verbs, coordination chains, clauses and subordinate clauses, modifiers per noun
  • the constructions: the passive with ser and with se, the participial and the gerund clauses, the chains of de, the split predicates, the impersonal se, the words of negation, the ratio of nouns to verbs
>>> from ests import SyntaxStats

>>> text = ("La revisión de las cuentas fue realizada por el comité. "
...         "Se llevó a cabo la reforma sin que nadie hiciera mención de los problemas.")
>>> ss = SyntaxStats(text)

>>> ss.n_sents, ss.n_words
(2, 24)

>>> ss.p_passive, ss.noun_verb_ratio
(0.6666666666666666, 2.3333333333333335)

>>> ss.split_predicates
('llevó cabo', 'hiciera mención')

>>> round(ss.mean_dependency_distance, 2)
2.05

A text is parsed with es_core_news_sm; another pipeline with a parser can be passed in nlp.

More in the documentation.

Cohesion statistics

Referential cohesion in the manner of Coh-Metrix and its Spanish adaptation Coh-Metrix-Esp:

  • the overlap of nouns, of arguments and of content words between adjacent sentences and between all pairs of sentences, binary and proportional
  • givenness: pronouns, demonstratives and the content words whose lemma was already used
  • temporal cohesion: the repetition of the tense and of the mood of the verbs of adjacent sentences
  • the density of 255 Spanish discourse markers by class - causal, adversative, concessive, temporal, additive, conditional, reformulative - and by kind
>>> from ests import CohesionStats

>>> text = ("El informe fue aprobado por la comisión. Sin embargo, el informe no resuelve el problema. "
...         "Por lo tanto, la comisión aplazó la decisión.")
>>> cs = CohesionStats(text)

>>> cs.n_sents, cs.n_words
(3, 23)

>>> cs.noun_overlap_adjacent, round(cs.p_given, 3)
(0.5, 0.167)

>>> cs.c_connectors
{'por lo tanto': 1, 'sin embargo': 1}

>>> round(cs.connectors, 2), round(cs.connectors_causal, 2)
(86.96, 43.48)

A text is parsed with es_core_news_sm; another pipeline can be passed in nlp.

More in the documentation.

Lexical sophistication statistics

How rare the words of a text are in the language, in the manner of TAALES: the mean frequency, range and dispersion of the lemmas by the frequency dictionary of Google Books Ngram, the shares of the frequency bands top-1000, 2000, 5000 and 10000, surprisal and perplexity, the lexical density. The statistics by the dictionary need it downloaded once.

>>> from ests import LexicalStats
>>> from ests.datasets import FreqDict

>>> FreqDict().download()
>>> ls = LexicalStats("El felinólogo examinaba al minino con parsimonia")
>>> ls.coverage, ls.p_top1000, round(ls.surprisal, 2)
(0.8571428571428571, 0.42857142857142855, 13.71)

More in the documentation.

Style metrics

The SEO indicators of Advego and Text.ru - nausea, water content, spam score, naturalness by Zipf's law, keyword density - and the markers of the officialese style that the Spanish guides to plain language warn about: the nouns derived from a verb, the compound prepositions of the administrative style, the parenthetical expressions and the clichés.

>>> from ests import StyleStats

>>> ss = StyleStats("Se procedió a la revisión del expediente en el marco del plan a la mayor brevedad.")
>>> ss.compound_prepositions, ss.cliches, ss.verbal_nouns
(6.25, 12.5, 20.0)

More in the documentation.

Phonostatistics

The shares of the classes of sounds, the consonant clusters, the hiatuses, the open syllables, the hardness and the indices of alliteration and assonance, counted over the sounds of a rule-based transcription, not over the letters: h is silent, ll and rr are one sound.

>>> from ests import PhonStats

>>> ps = PhonStats("Los suspiros se escapan de su boca de fresa")
>>> round(ps.p_voiceless, 3), round(ps.hardness, 3), ps.sounds[-1]
(0.371, 0.684, ('f', 'r', 'e', 's', 'a'))

More in the documentation.

Verse statistics

The scansion of Spanish verse by its syllabic meter: the metrical syllables of a line with the synalepha and the law of the final stress, fitted to the meter of the poem; the meter, the stress profile and the types of the endecasílabo; the rhyme in full (consonante) and by assonance (asonante) with its scheme, the strophes and the form of a poem. No model and no dictionary are needed.

>>> from ests import VerseStats

>>> text = """Cuando me paro a contemplar mi estado
... y a ver los pasos por do me han traído,
... hallo, según por do anduve perdido,
... que a mayor mal pudiera haber llegado."""
>>> vs = VerseStats(text)
>>> vs.meter, vs.patterns[0], vs.c_rhythms
('endecasílabo', '---+---+-+-', {'4-8-10': 2, '4-7-10': 1, '3-6-10': 1})
>>> vs.syllables[0]
('cuan', 'do', 'me', 'pa', 'ro‿a', 'con', 'tem', 'plar', 'mi‿es', 'ta', 'do')
>>> vs.rhyme_schemes, vs.strophes
(('ABBA',), ('cuarteto',))

More in the documentation.

spaCy components

Every statistics class is also a component of a pipeline, so a text is annotated and measured in one pass and the statistics travel with the Doc:

>>> import ests
>>> import spacy

>>> nlp = spacy.load("es_core_news_sm")
>>> for factory in ("basic", "morph", "syntax"):
...     _ = nlp.add_pipe(f"ests_{factory}", name=factory, last=True)

>>> nlp.pipe_names[-3:]
['basic', 'morph', 'syntax']

>>> doc = nlp("El gato duerme en la ventana. Los niños juegan en el parque.")
>>> doc._.basic.n_words, doc._.morph.pos[:2], doc._.syntax.tree_depth
(12, ('DET', 'NOUN'), 2.0)

The factories are ests_basic, ests_readability, ests_diversity, ests_morph, ests_syntax, ests_cohesion, ests_lexical, ests_style, ests_phon and ests_verse; the name of the pipe is the name of the extension.

More in the documentation.

Corpus measures

The measures of corpus linguistics that compare corpora and describe the use of a word:

  • keywords of a target corpus against a reference one or against the frequency dictionary of Google Books Ngram: the log-likelihood with its p-value, Log Ratio, chi-square, %DIFF, BIC, ELL and the odds ratio
  • collocations by logDice, MI, MI³, t-score, Dice, log-likelihood, NPMI and minimum sensitivity
  • the dispersion of a word over the parts of a text: DP of Gries, normalized DP, Juilland's D, Carroll's D2, Rosengren's S and the Kullback-Leibler divergence
  • a KWIC concordance by word form or by lemma
  • stylometry: the distances between texts by Burrows's Delta and its variants, with the attribution of a text to reference authors, the markers of preferred and avoided words by Zeta, Kilgarriff's chi-square, the Mendenhall curve and the profile of the function words
  • the comparison of two corpora by 132 features of a text over windows of about the same size: Cliff's delta, Cohen's d and the AUC of every feature, the Mann-Whitney test with Holm's correction and a bootstrap interval of the difference of the medians
>>> from ests import WordsExtractor
>>> from ests.corpus import collocations, keyness, kwic, zeta

>>> we = WordsExtractor(use_lexemes=True, lowercase=True)
>>> target = we.extract("El gato estaba en la ventana y miraba a los pájaros. Los pájaros se fueron y el gato "
...                     "se durmió en la ventana. Mañana el gato volverá a estar en la ventana y mirará a los pájaros.")
>>> reference = we.extract("El perro estaba en el suelo y dormía. Después el perro comió y volvió a dormir. "
...                        "Mañana el perro saldrá a pasear.")

>>> [(k.word, round(k.g2, 2)) for k in keyness(target, reference, top_n=2)]
[('gato', 2.8), ('pájaro', 2.8)]

>>> [(c.left, c.right, round(c.score, 1)) for c in collocations(target, window=2, top_n=2)]
[('en', 'ventana', 13.0), ('estar', 'en', 12.7)]

>>> [line.keyword for line in kwic("Los gatos juegan y el gato duerme.", "gato", by_lemma=True)]
['gatos', 'gato']

>>> scores = zeta(target, reference, segment_size=10)
>>> [z.word for z in scores[:2]], [z.word for z in scores[-2:]]
(['gato', 'ventana'], ['dormir', 'perro'])

Words are compared as they are, so case, lemmas and stop words are chosen at the extraction.

More in the documentation.

Datasets
  • spanish_literature - Spanish-language literature in the public domain: 150 works by 33 authors from Spain, Latin America and the Philippines, from Cervantes to the 1920s, in prose, poems, drama and publicism
  • spanish_sonnets - Spanish sonnets of the Diachronic Spanish Sonnet Corpus (DISCO) in the public domain: 4,259 sonnets by 1,167 authors of the 15th-20th centuries, every line with its metrical pattern and the label of its rhyme, the automatic annotation of DISCO (CC BY 4.0)
  • freq_dict - a frequency dictionary of 83,785 Spanish lemmas from the books of Google Books Ngram of 1980-2019: ipm, range and Juilland's D over the years, the number of books and the part of speech (CC BY 3.0)

The texts are cut to the text of the author and come with the genre, the author, the title, the years of the first publication and the country; the records can be filtered by any of them and by the length of the text.

>>> from ests.datasets import SpanishLiterature

>>> sl = SpanishLiterature()
>>> sl.download()
>>> for record in sl.get_records(author="galdos", year_from=1880, year_to=1884):
...     print(record["title"], record["year_from"], len(record["text"]))
La desheredada 1881 820454
El amigo Manso 1882 521156
La de Bringas 1884 413730
Tormento 1884 477286

The archive (19 MB) is downloaded once by download() into the data directory; before that get_texts() and get_records() raise DatasetNotFoundError.

More in the documentation.

Visualizers

The matplotlib plots take the axes ax and return Axes; the network of collocations and the word tree are rendered by the executables of Graphviz. Cosine Delta separates three novels by Galdós from three by Unamuno:

import matplotlib.pyplot as plt
from ests import WordsExtractor
from ests.corpus import delta
from ests.datasets import SpanishLiterature
from ests.visualizers import dendrogram_plot, pca_plot

# six novels by Galdós and Unamuno from the corpus of literature
titles = (
    "Marianela",
    "Misericordia",
    "Torquemada en la hoguera",
    "Niebla",
    "Abel Sánchez",
    "La tía Tula",
)
sl = SpanishLiterature()
sl.download()
novels = {
    record["title"]: record["text"]
    for author in ("galdos", "unamuno")
    for record in sl.get_records(author=author)
    if record["title"] in titles
}
we = WordsExtractor(lowercase=True)
corpus = {title: we.extract(novels[title]) for title in titles}
distances = delta(corpus, n_mfw=100, variant="cosine")

fig, (left, right) = plt.subplots(1, 2, figsize=(13, 5))
dendrogram_plot(distances, ax=left)
pca_plot(corpus, n_mfw=100, ax=right)

Stylometric plots

More in the documentation.

Development

The project uses uv for dependency management and ruff for linting and formatting.

git clone https://github.com/SergeyShk/esTS.git
cd esTS

make deps        # create the environment and install dependencies
make test        # run the tests and docstring examples (doctest)
make lint        # ruff + mypy

Run make help for the full list of commands.

The documentation is bilingual: English pages are docs/*.md, Spanish ones are docs/*.es.md next to them (mkdocs-static-i18n); when editing a page, update both versions.

The installed version is ests.__version__. All exceptions inherit ests.EstsError and a built-in class (ValueError, TypeError, KeyError, OSError or RuntimeError), so except ValueError keeps working. The library prints nothing on its own: its messages go to the ests logger and are silent by default. More in the documentation.

Before submitting changes, install the hooks that run the linters on commit and the tests on push:

uv run pre-commit install

Contributing

Bug reports, ideas and pull requests are welcome in the issues. The workflow and the checks to run before a pull request are in CONTRIBUTING.md; the rules of conduct are in the code of conduct.

Project structure
  • demo - the demo on Hugging Face Spaces (Gradio)
  • docs - project documentation
  • ests:
    • basic_stats.py - basic text statistics of the anyTS core with the Spanish syllables
    • cohesion_stats.py - cohesion statistics
    • lexical_stats.py - lexical sophistication statistics
    • components.py - components of a spaCy pipeline
    • corpus - measures of corpus linguistics: keywords, collocations, dispersion, concordance, stylometry, comparison of corpora
    • datasets - datasets: Spanish-language literature, Spanish sonnets, the frequency dictionary
    • constants.py - constants of the Spanish language and of the metrics
    • diversity_stats.py - lexical diversity metrics of the anyTS core over a string or a Doc
    • exceptions.py - library exceptions
    • extractors.py - the extractors of the anyTS core with the Spanish tokenizers
    • morph_stats.py - morphological statistics
    • readability_stats.py - readability metrics: the Spanish formulas over those of the anyTS core
    • style_stats.py - style metrics
    • phon_stats.py - phonostatistics
    • syntax_stats.py - syntactic statistics
    • verse_stats.py - verse statistics
    • syllables.py - syllabification and stress
    • utils.py - helper tools
    • visualizers - plots: Zipf's law, fingerprinting, word tree, corpus and stylometric plots, vocabulary growth, sentence lengths, text highlighting
  • scripts - scripts that build the archives of the datasets and fetch the anyTS pages of the documentation
  • tests - tests mirroring the package structure

Authors

License

MIT

Citation

Please use the following BibTeX entry for citing esTS if you use it in your research or software. The same metadata is in CITATION.cff ("Cite this repository" on GitHub). The Concept DOI 10.5281/zenodo.22924655 points at every version of the library; the DOI of a single version is on the page of its release.

@software{esTS,
  author = {Sergey Shkarin},
  title = {{esTS, a library for statistics extraction from texts in Spanish}},
  year = 2026,
  doi = {10.5281/zenodo.22924655},
  url = {https://github.com/SergeyShk/esTS}
}

About

Spanish Texts Statistics: readability, lexical diversity, morphology, syntax, cohesion, stylometry, metre and rhyme for Spanish text

Topics

Resources

Code of conduct

Contributing

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages