A COMPACT EXTENSION OF HEAPS' LAW: INCORPORATING PHONOLOGICAL STRUCTURE AND LEXICAL REPETITION
DOI:
https://doi.org/10.51891/rease.v12i9.30297Palavras-chave:
Heaps' Law. Quantitative Linguistics. Zipf's Law of Abbreviation.Resumo
The classical formulation of Heaps' Law establishes a bivariate relationship between the length of a text and the growth of its vocabulary, disregarding morphological, phonological, and stylistic factors inherent to natural language. This work proposes and validates a compact, generalized extension of Heaps' Law, incorporating the proportion of monosyllabic words (p₁), theoretically grounded in Zipf's Law of Abbreviation, and a variable indicating high lexical repetition (Drep). The simulation-based validation adopts a mechanistic data-generating process (DGP) based on a Zipf-Mandelbrot urn, in which vocabulary and the proportion of monosyllables emerge from a token sampling process, without imposing the functional form of either compared model on the generating process. The results tables report the average estimated parameters across 200 Monte Carlo replications, ensuring that the estimates do not depend on a single sample realization. Empirical validation on a large English-language dataset, the Movie Reviews Corpus, corroborates the gains of the proposed model in adjusted R², AIC, BIC, and nested-model tests (ANOVA).
Downloads
Downloads
Publicado
Como Citar
Edição
Seção
Categorias
Licença
Atribuição CC BY