Skip to main content

ETHNOS_APP

Home • Search • Journals • List 0

Codebook LLMs

Evaluating LLMs as Measurement Tools for Political Science Concepts

Bibliographic Data

ID6332293
AuthorsAndrew Halterman (0000-0001-9716-9555, Michigan State University, corresponding author), Katherine A Keith (0000-0002-8101-4572, Williams College)
Year2025
Pages1-17
Publication date2025-09-19
Peer ReviewedYes
Open AccessYes
TypeARTICLE
VenuePolitical Analysis (JOURNAL)
Journal identifiersISSN: 1047-1987 • E-ISSN: 1476-4989
PublisherCambridge University Press (CUP) (PUBLISHER)
DOI10.1017/pan.2025.10017
OpenAlexW4414362097
LanguageEN
Citations received8
References cited23

Codebooks—documents that operationalize concepts and outline annotation procedures—are used almost universally by social scientists when coding political texts. To code these texts automatically, researchers are increasingly turning to generative large language models (LLMs). However, there is limited empirical evidence on whether “off-the-shelf” LLMs faithfully follow real-world codebook operationalizations and measure complex political constructs with sufficient accuracy. To address this, we gather and curate three real-world political science codebooks—covering protest events, political violence, and manifestos—along with their unstructured texts and human-coded labels. We also propose a five-stage framework for codebook-LLM measurement: Preparing a codebook for both humans and LLMs, testing LLMs’ basic capabilities on a codebook, evaluating zero-shot measurement accuracy (i.e., off-the-shelf performance), analyzing errors, and further (parameter-efficient) supervised training of LLMs. We provide an empirical demonstration of this framework using our three codebook datasets and several pre-trained 7–12 billion open-weight LLMs. We find current open-weight LLMs have limitations in following codebooks zero-shot, but that supervised instruction-tuning can substantially improve performance. Rather than suggesting the “best” LLM, our contribution lies in our codebook datasets, evaluation framework, and guidance for applied researchers who wish to implement their own codebook-LLM measurement projects

Artificial Intelligence in Law · Legal Education and Practice Innovations

  • Using LLMs for measurement in diplomatic speeches

    Open Access•Anton Peez, Johannes Scherzinger•The Review of International…•2026

  • NEClass

    Open Access•Tobias Schmidt•Communication Methods and Measures•2026

  • Piercing a Methodological Bubble

    Open Access•Dwayne Woods•Fudan Journal of the Humanities…•2026

  • The Stories Individuals “Like”

    Open Access•Tinghui Wu, Guoqiang Yan•Review of Policy Research•2026

  • Text as Data and Causal Inference in Sociology

    Open Access•N Schwitter, Robert L Bach et al.•KZfSS Kölner Zeitschrift für…•2026

  • Disaster, Distributive Politics, and the Persistence of Partisan Divides in Climate Policy

    Open Access•Adam Cayton, Brian Benjamin Crisher•Political Research Quarterly•2025

  • Cops and crypto

    Open Access•Chandler G Robinson, Matthew W Logan et al.•Journal of Criminal Justice•2026

  • Fine-tuned large language models can replicate expert coding better than trained coders

    Open Access•Dahyun Choi, Denis Peskoff et al.•Political Science Research and…•2026

  • ChatGPT outperforms crowd workers for text-annotation tasks

    Open Access•Fabrizio Gilardi, Meysam Alizadeh et al.•Proceedings of the National…•2023

  • The Measurement of Observer Agreement for Categorical Data

    J R Landis, Gary G Koch•Biometrics•1977

  • Do AIs know what the most important issue is? Using language models to code open-text social survey responses at scale

    Open Access•Jonathan Mellon, Jack Bailey et al.•Research & Politics•2024

  • Large language models as a substitute for human experts in annotating political text

    Open Access•Michael Heseltine, Bernhard Clemm Von Hohenberg•Research & Politics•2024

  • Towards Faithful Model Explanation in NLP

    Open Access•Qing Lyu, Marianna Apidianaki et al.•Computational Linguistics•2024

  • Measuring political violence in Pakistan

    Open Access•Ethan Bueno De Mesquita, Carol Christine Fair et al.•Conflict Management and Peace…•2014

  • Stance detection

    Open Access•Michael Burnham•Political Science Research and…•2025

  • Synthetically generated text for supervised text analysis

    Open Access•Andrew Halterman•Political Analysis•2025

  • Testing Causal Theories with Learned Proxies

    Open Access•David Knox, Dean Knox et al.•Annual Review of Political Science•2022

  • Can Large Language Models Transform Computational Social Science

    Open Access•Caleb Ziems, William A Held et al.•Computational Linguistics•2023

  • Text as Data

    Open Access•J Adams•Contemporary Sociology A Journal…•2024

  • Measurement Validity

    Open Access•Robert Adcock, Donald Collier et al.•American Political Science Review•2001

Unique citing works8
Citations per year8
Citation span2025 - 2026 (2)
Citation velocitycurrent
Highly citedNo
Citation typesNeutral: 8
Ethnos_APP • Open Source Project • MIT License • Frontend v2.0.0 • Privacy and Cookies • API Documentation: api.ethnos.app/docs • API Source Code: GitHub • DOI: 10.5281/zenodo.17049435 • Frontend Source Code: GitHub • DOI: 10.5281/zenodo.17050053 • cruz.rio.br • Expectantes Misericordiae