Skip to main content

ETHNOS_APP

Home • Search • Journals • List 0

Validity of Comparative Judgment Scores

How Assessors Evaluate Aspects of Text Quality When Comparing Argumentative Texts

Bibliographic Data

ID22166340
AuthorsMarije Lesterhuis (0000-0002-3808-556X, University of Antwerp, corresponding author), Renske Bouwer (0000-0003-0434-0224, Utrecht University), Tine Van Daal (0000-0001-9398-9775, University of Antwerp), Vincent Donche (0000-0002-9405-3896, University of Antwerp), Sven De Maeyer (0000-0003-2888-1631, University of Antwerp)
Year2022
Volume7
Publication date2022-05-13
Peer ReviewedYes
Open AccessYes
TypeARTICLE
VenueFrontiers in Education (JOURNAL)
Journal identifiersISSN: 2504-284X • E-ISSN: 2504-284X
PublisherFrontiers Media SA (PUBLISHER • CH)
DOI10.3389/feduc.2022.823895
OpenAlexW4280593927
LanguageEN
Citations received5
References cited27

The advantage of comparative judgment is that it is particularly suited to assess multidimensional and complex constructs as text quality. This is because assessors are asked to compare texts holistically and to make a quality judgment for each text in a pairwise comparison based upon on the most salient and critical differences. Also, the resulted rank order is based on the judgment of all assessors, representing the shared consensus. In order to be able to select the right number of assessors, the question is to what extent the conceptualization of assessors prevails in the aspects they base their judgment on, or whether comparative judgment minimizes the differences between assessors. In other words, can we detect types of assessors who tend to consider certain aspects of text quality more often than others? A total of 64 assessors compared argumentative texts, after which they provided decision statements on what aspects of text quality had informed their judgment. These decision statements were coded on six overarching themes of text quality: argumentation, organization, language use, language conventions, source use, references, and layout. Using a multilevel-latent class analysis, four different types of assessors could be distinguished: narrowly focused, broadly focused, source-focused, and language-focused. However, the analysis also showed that all assessor types mainly focused on argumentation and organization, and that assessor types only partly explained whether the aspect of text quality was mentioned in a decision statement. We conclude that comparative judgment is a strong method for comparing complex constructs like text quality. First, because the rank order combines different views on text quality, but foremost because the method of comparative judgment minimizes differences between assessors

Argumentation theory · Argumentative · Conceptualization · Epistemology · Linguistics · Natural language processing · Pairwise comparison · Computer Science · Discourse Analysis in Language Studies · Mathematics · Psychology · Software Engineering Research · Artificial Intelligence

  • Crowdsourced Comparative Judgement for Evaluating Learner Texts

    Open Access•Peter Thwaites, Nathan Vandeweerd et al.•Applied Linguistics•2025

  • Beyond reliability

    Open Access•Kjetil Egelandsdal, Jan-Ove Færstad•Frontiers in Education•2026

  • Testing crowdsourcing as a means of recruitment for the comparative judgement of L2 argumentative essays

    Open Access•Peter Thwaites, Magali Paquot•Journal of Second Language Writing•2025

  • Do cognitive processes and motives for argumentative writing converge in writer profiles

    Open Access•Fien De Smedt, Yana Landrieu et al.•The Journal of Educational Research•2022

  • Comparative Judgement for evaluating young learners’ EFL writing performances

    Open Access•Rebecca Sickinger, Tineke Brunfaut et al.•Language Testing•2025

  • Think-aloud protocols in research on essay rating

    Open Access•Khaled Barkaoui•Language Testing•2011

  • Rater types in writing performance assessments

    Open Access•Thomas Eckes•Language Testing•2008

  • Indeterminacy in the use of preset criteria for assessment and grading

    D Royce Sadler•Assessment & Evaluation in Higher…•2009

  • Meaning and Values in Test Validation

    Samuel Messick•Educational Researcher•1989

Unique citing works5
Citations per year1,25
Citation span2022 - 2026 (5)
Citation velocitycurrent
Highly citedNo
Citation typesNeutral: 5

Tools

Open DOIOpen Access
Ethnos_APP • Open Source Project • MIT License • Frontend v2.0.0 • Privacy and Cookies • API Documentation: api.ethnos.app/docs • API Source Code: GitHub • DOI: 10.5281/zenodo.17049435 • Frontend Source Code: GitHub • DOI: 10.5281/zenodo.17050053 • cruz.rio.br • Expectantes Misericordiae