Efficiency vs. understanding
A critical examination of ChatGPT’s performance in context-sensitive annotation tasks
Bibliographic Data
| ID | 19214730 |
|---|---|
| Authors | Claudia Buder (0009-0005-6265-0652, Université Paris 1 Panthéon-Sorbonne), Nina-Sophie Fritsch (0000-0001-7028-8239, Vienna University of Economics and Business), Chiara Osorio-Krauter (0009-0001-4305-9825, European University Institute), Aaron Philipp (0009-0003-0585-8794, University of Potsdam), Roland Verwiebe (0000-0002-3202-8820, University of Potsdam, corresponding author), Sarah Weissmann (0000-0001-5284-9805, University of Potsdam), Weissmann (0009-0009-7623-2549, University of Potsdam) |
| Year | 2026 |
| Volume | 56 |
| Issue | 4 |
| Pages | 1-22 |
| Publication date | 2026-05-08 |
| Peer Reviewed | Yes |
| Open Access | Yes |
| Type | ARTICLE |
| Venue | International Journal of Sociology (JOURNAL) |
| Journal identifiers | ISSN: 0020-7659 • E-ISSN: 1557-9336 |
| Publisher | Taylor & Francis (PUBLISHER • GB) |
| DOI | 10.1080/00207659.2026.2661167 |
| OpenAlex | W7160622642 |
| Language | EN |
| References cited | 76 |
This paper examines context-sensitive annotation tasks performed by OpenAI’s o3 reasoning model using content creator profiles on YouTube. It explores the inherently ambiguous and socially and culturally embedded task of annotating race. Analysing 500 annotations generated by ChatGPT, we first examine performance metrics and benchmark results against a human-annotated dataset. We further conduct a thematic analysis of the justifications provided by the model, as users increasingly rely on them making them important for understanding how the model frames social practices. Analyzing justifications also exposes inconsistencies, biases, and classification errors that remain invisible in aggregate performance metrics alone. Our findings show that despite efficiency and scalability, ChatGPT's limited cultural understanding and lack of critical reflexivity constrain its performance in complex annotation tasks. We argue that the annotation of sensitive social characteristics requires reflexive scientific practices and potentially hybrid annotation strategies to mitigate bias and preserve contextual integrity in academic research
Annotation · Baseline (sea) · Data collection · Task (project management) · Artificial Intelligence in Healthcare and Education · Explainable Artificial Intelligence (XAI · Topic Modeling
Algorithmic Bias in Education
ChatGPT
Taking race out of human genetics
Looking the Part
ChatGPT outperforms crowd workers for text-annotation tasks
Thematic analysis.
Ethics and discrimination in artificial intelligence-enabled recruitment practices
Who Counts as a “Person of Color”? The Roles of Ancestry, Phenotype, Self-Identification, and Other Factors
ChatGPT is bullshit
Reflecting on reflexive thematic analysis
Racial Non-equivalence of Socioeconomic Status and Self-rated Health among African Americans and Whites
Assessing Data Quality in the Age of Digital Social Research
ChatGPT is incredible (at being average)
Large Language Models Outperform Expert Coders and Supervised Classifiers at Annotating Political Social Media Messages
Navigating the Risks of Using Large Language Models for Text Annotation in Social Science Research
The Unceasing Significance of Colorism
Social markers of acculturation
Anchoring bias in large language models
Staying in control of technology
The Fluidity of Racial Classifications
Racial categorization of faces
Race as biology is fiction, racism as a social problem is real
What We Talk About When We Talk About Ethnicity
Integrating Generative Artificial Intelligence into Social Science Research
From Codebooks to Promptbooks
The Augmented Social Scientist
Why Sociology Matters to Race and Biosocial Science
Goodbye human annotators? Content analysis of social policy debates using ChatGPT
| Citation velocity | historical |
|---|---|
| Highly cited | No |