Sensitivity of dispersion measures to distributional patterns and corpus design
Bibliographic Data
| ID | 21701255 |
|---|---|
| Authors | Lukas Sönning (0000-0002-2705-395X, University of Bamberg), Jamie Egbert (0000-0002-3751-2865, Northern Arizona University) |
| Year | 2026 |
| Publication date | 2026-07-03 |
| Peer Reviewed | Yes |
| Open Access | Yes |
| Type | ARTICLE |
| Venue | International Journal of Corpus Linguistics (JOURNAL) |
| Journal identifiers | ISSN: 1384-6655 • E-ISSN: 1569-9811 |
| Publisher | John Benjamins Publishing Company (PUBLISHER • NL) |
| DOI | 10.1075/ijcl.25008.son |
| OpenAlex | W4399600906 |
| Language | EN |
| References cited | 24 |
Recent work has shown that dispersion measures respond to multiple features in the data: Juilland’s D varies systematically with the number of corpus parts, and all commonly used indices are affected by the frequency of an item. This study uses a simulation approach to provide further insights into the sensitivity of dispersion measures to differences in corpus design (number of texts, average text length, distribution of text lengths) and distributional milieu (frequency and evenness of distribution). Our results suggest that, within the settings covered by our analysis, the factors frequency and evenness of distribution have roughly the same impact, though there is some variation among measures. The average text length emerges as another feature that leaves its mark on the observed scores. Finally, we note that D 2 exhibits the same weakness as D — it varies with the number of corpus parts that enter the analysis
Econometrics · Frequency distribution · Index of dispersion · Linguistics · Natural language processing · Physics · Population · Sociology · Species evenness · Species richness · Statistics · Computational and Text Analysis Methods · Computer Science · Demography · Engineering · Mathematics · Natural Language Processing Techniques · Topic Modeling · Artificial Intelligence · Geology
Ggplot2
Word Frequency Distributions
Generalized Additive Models for Location, Scale and Shape
Ggplot2
Which Words Matter Most? Operationalizing Lexical Prevalence For Rank-Ordered Word Lists
On the (non)utility of Juilland’s D to measure lexical dispersion in large corpora
Dispersions and adjusted frequencies in corpora
Lexical dispersion and corpus design
Common Words in Spanish
Frequency Dictionary of Spanish Words
Poisson regression for linguists
The International Corpus of English (ICE) Project
| Citation velocity | historical |
|---|---|
| Highly cited | No |