A Principled Framework for Evaluating on Typologically Diverse Languages
Bibliographic Data
Beyond individual languages, multilingual natural language processing (NLP) research increasingly aims to develop models that perform well across languages generally. However, evaluating these systems on all the world’s languages is practically infeasible. To attain generalizability, representative language sampling is essential. Previous work argues that generalizable multilingual evaluation sets should contain languages with diverse typological properties. However, “typologically diverse” language samples have been found to vary considerably in this regard, and popular sampling methods are flawed and inconsistent. We present a language sampling framework for selecting highly typologically diverse languages given a sampling frame, informed by language typology. We compare sampling methods with a range of metrics and find that our systematic methods consistently retrieve more typologically diverse language selections than previous methods in NLP. Moreover, we provide evidence that this affects generalizability in multilingual model evaluation, emphasizing the importance of diverse language sampling in NLP evaluation
Computational linguistics · Generalizability theory · Language identification · Natural language · Range (aeronautics · Sample (material · Sampling (signal processing · Computational and Text Analysis Methods · Natural Language Processing Techniques · Topic Modeling
The Rise and Fall of Languages
Unsupervised Cross-lingual Representation Learning at Scale
Statistical bias control in typology
A Method of Language Sampling
Disentangling geography from genealogy
Some Principles on the Use of Macro-Areas in Typological Comparison
Languages Through the Looking Glass of BPE Compression
Large Linguistic Areas and Language Sampling
Teaching Tariana, an endangered language from Northwest Amazonia
The Greenbergian Word Order Correlations
| Citation velocity | historical |
|---|---|
| Highly cited | No |