Crowdsourced Comparative Judgement for Evaluating Learner Texts
How Reliable are Judges Recruited from an Online Crowdsourcing Platform?
Bibliographic Data
| ID | 22871285 |
|---|---|
| Authors | Peter Thwaites (0000-0003-2176-9620, Centre for English Corpus Linguistics, Institut Langage et Communication , Collège Érasme, Cardinal Mercier, 1, Louvain-la-Neuve 1348 UCLouvain), Nathan Vandeweerd (0000-0002-6498-9474, Radboud University Nijmegen), Magali Paquot (0000-0001-5687-5074, Centre for English Corpus Linguistics, Institut Langage et Communication , Collège Érasme, Cardinal Mercier, 1, Louvain-la-Neuve 1348 UCLouvain, corresponding author) |
| Year | 2025 |
| Volume | 46 |
| Issue | 4 |
| Pages | 611-628 |
| Publication date | 2025-08-28 |
| Peer Reviewed | Yes |
| Open Access | Yes |
| Type | ARTICLE |
| Venue | Applied Linguistics (JOURNAL) |
| Journal identifiers | ISSN: 0142-6001 • E-ISSN: 1477-450X |
| Publisher | Oxford University Press (OUP) (PUBLISHER) |
| DOI | 10.1093/applin/amae048 |
| OpenAlex | W4401158115 |
| Language | EN |
| Citations received | 3 |
| References cited | 46 |
Recent studies of proficiency measurement and reporting practices in applied linguists have revealed widespread use of unsatisfactory practices such as the use of proxy measures of proficiency in place of explicit tests. Learner corpus research is one specific area affected by this problem: few learner corpora contain reliable, valid evaluations of text proficiency. This has led to calls for the development of new L2 writing proficiency measures for use in research contexts. Answering this call, a recent study by Paquot et al. (2022) generated assessments of learner corpus texts using a community-driven approach in which judges, recruited from the linguistic community, conducted assessments using comparative judgement. Although the approach generated reliable assessments, its practical use is limited because linguists are not always available to contribute to data collections. This paper, therefore, explores an alternative approach, in which judges are recruited through a crowdsourcing platform. We find that assessments generated in this way can reach near identical levels of reliability and concurrent validity to those produced by members of the linguistic community.
Crowdsourcing · Inter-rater reliability · Judgement · Language proficiency · Linguistics · Mathematics education · Natural language processing · Proxy (statistics) · Rating scale · Reliability (semiconductor) · World Wide Web · Artificial Intelligence · Computer Science · Discourse Analysis in Language Studies · Interpreting and Communication in Healthcare · Natural Language Processing Techniques · Psychology
Measuring L2 Proficiency
Assessing Writing
Modern Applied Statistics with S
Emmeans
Proficiency Assessment Standards in Second Language Acquisition Research
Language Proficiency in Native and Nonnative Speakers
Data quality in online human-subjects research
Rank Analysis of Incomplete Block Designs
Proficiency Level--a Fuzzy Variable in Computer Learner Corpora
Exploring the Validity of Comparative Judgement
Validity of Comparative Judgment Scores
Research synthesis and historiography
The assessment of writing ability
Replication in Second Language Research
Coming of age
Instructional manipulation checks
(Why) Are Open Research Practices the Future for the Study of Language Learning
Crowdsourced Adaptive Comparative Judgment
Assessment of L2 Proficiency in Second Language Acquisition Research
Guidelines for Reporting Quantitative Methods and Results in Primary Research
Proficiency Reporting Practices in Research on Second Language Acquisition
Language Proficiency in Native and Non-native Speakers
| Unique citing works | 3 |
|---|---|
| Citations per year | 3 |
| Citation span | 2025 - 2025 (1) |
| Citation velocity | recent |
| Highly cited | No |
| Citation types | Neutral: 3 |