Human tests for machine models
What lies “Beyond the Imitation Game
Bibliographic Data
| ID | 6015463 |
|---|---|
| Authors | Noya Kohavi (Copenhagen Business School Copenhagen Denmark), Anna Weichselbraun (0000-0002-3705-9968, University of Vienna Vienna Austria, corresponding author) |
| Year | 2025 |
| Volume | 36 |
| Issue | 1 |
| Publication date | 2025-11-24 |
| Peer Reviewed | Yes |
| Open Access | Yes |
| Type | ARTICLE |
| Venue | Journal of Linguistic Anthropology (JOURNAL) |
| Journal identifiers | ISSN: 1055-1360 • E-ISSN: 1548-1395 |
| Publisher | Wiley (PUBLISHER • GB) |
| DOI | 10.1111/jola.70035 |
| OpenAlex | W7106569794 |
| Language | EN |
| Citations received | 1 |
| References cited | 46 |
Benchmarking large language models (LLMs) is a key practice for evaluating their capabilities and risks. This paper considers the development of “BIG Bench,” a crowdsourced benchmark designed to test LLMs “Beyond the Imitation Game.” Drawing on linguistic anthropological and ethnographic analysis of the project's GitHub repository, we examine how contributors developed tasks based on their lay understandings of language, cognition, and intelligence. By tracing how contributors make implicit judgments about what constitutes a meaningful test of intelligence, we show how widespread language ideologies shape the evaluation of LLMs and the imaginaries that guide their development
Ethnography · Imitation · Key (lock · Nexus (standard · Test (biology · Tracing
The Cultural Logic of Computation
Voices of Modernity
The Seductions of Quantification
Standards
Hyperauthorship
Grounded Cognition: Past, Present, and Future
Dissociating language and thought in large language models
Climbing towards NLU
The debate over understanding in AI’s large language models
Minds, brains, and programs
I.—computing Machinery and Intelligence
Language Ideology
Sorting Things Out
Conversation Analysis and Online Interaction
We get the algorithms of our ground truths
Audit Cultures
Discourse and the No-thing-ness of Culture
On Semiotic Ideology
Toward cultural interpretability
Perspectives on algorithmic normativities
Agreements 'in the wild
Intelligence tests and the individual
Language is not a data set-Why overcoming ideologies of dataism is more important than ever in the age of AI
Irony in conversation
A History of the Modern Fact
Stance and Subjectivity
Commensuration as a Social Process
Accounting for Rationality
The Methodology of Racial Testing
Language at the Limits of the Human
| Unique citing works | 1 |
|---|---|
| Citations per year | 1 |
| Citation span | 2026 - 2026 (1) |
| Citation velocity | current |
| Highly cited | No |
| Citation types | Neutral: 1 |