Saltar al contenido principal

ETHNOS_APP

Inicio • Búsqueda • Revistas • Lista 0

Codebook LLMs

Evaluating LLMs as Measurement Tools for Political Science Concepts

Datos Bibliográficos

ID6332293
AutoresAndrew Halterman (0000-0001-9716-9555, Michigan State University, autor de correspondencia), Katherine A Keith (0000-0002-8101-4572, Williams College)
Año2025
Páginas1-17
Fecha de publicación2025-09-19
Peer ReviewedSí
Open AccessSí
TipoARTICLE
RevistaPolitical Analysis (JOURNAL)
Identificadores de la revistaISSN: 1047-1987 • E-ISSN: 1476-4989
EditorialCambridge University Press (CUP) (PUBLISHER)
DOI10.1017/pan.2025.10017
OpenAlexW4414362097
IdiomaEN
Citas recibidas8
Referencias citadas23

Codebooks—documents that operationalize concepts and outline annotation procedures—are used almost universally by social scientists when coding political texts. To code these texts automatically, researchers are increasingly turning to generative large language models (LLMs). However, there is limited empirical evidence on whether “off-the-shelf” LLMs faithfully follow real-world codebook operationalizations and measure complex political constructs with sufficient accuracy. To address this, we gather and curate three real-world political science codebooks—covering protest events, political violence, and manifestos—along with their unstructured texts and human-coded labels. We also propose a five-stage framework for codebook-LLM measurement: Preparing a codebook for both humans and LLMs, testing LLMs’ basic capabilities on a codebook, evaluating zero-shot measurement accuracy (i.e., off-the-shelf performance), analyzing errors, and further (parameter-efficient) supervised training of LLMs. We provide an empirical demonstration of this framework using our three codebook datasets and several pre-trained 7–12 billion open-weight LLMs. We find current open-weight LLMs have limitations in following codebooks zero-shot, but that supervised instruction-tuning can substantially improve performance. Rather than suggesting the “best” LLM, our contribution lies in our codebook datasets, evaluation framework, and guidance for applied researchers who wish to implement their own codebook-LLM measurement projects

Artificial Intelligence in Law · Legal Education and Practice Innovations

  • Using LLMs for measurement in diplomatic speeches

    Open Access•Anton Peez, Johannes Scherzinger•The Review of International…•2026

  • NEClass

    Open Access•Tobias Schmidt•Communication Methods and Measures•2026

  • Piercing a Methodological Bubble

    Open Access•Dwayne Woods•Fudan Journal of the Humanities…•2026

  • The Stories Individuals “Like”

    Open Access•Tinghui Wu, Guoqiang Yan•Review of Policy Research•2026

  • Text as Data and Causal Inference in Sociology

    Open Access•N Schwitter, Robert L Bach et al.•KZfSS Kölner Zeitschrift für…•2026

  • Disaster, Distributive Politics, and the Persistence of Partisan Divides in Climate Policy

    Open Access•Adam Cayton, Brian Benjamin Crisher•Political Research Quarterly•2025

  • Cops and crypto

    Open Access•Chandler G Robinson, Matthew W Logan et al.•Journal of Criminal Justice•2026

  • Fine-tuned large language models can replicate expert coding better than trained coders

    Open Access•Dahyun Choi, Denis Peskoff et al.•Political Science Research and…•2026

  • ChatGPT outperforms crowd workers for text-annotation tasks

    Open Access•Fabrizio Gilardi, Meysam Alizadeh et al.•Proceedings of the National…•2023

  • The Measurement of Observer Agreement for Categorical Data

    J R Landis, Gary G Koch•Biometrics•1977

  • Do AIs know what the most important issue is? Using language models to code open-text social survey responses at scale

    Open Access•Jonathan Mellon, Jack Bailey et al.•Research & Politics•2024

  • Large language models as a substitute for human experts in annotating political text

    Open Access•Michael Heseltine, Bernhard Clemm Von Hohenberg•Research & Politics•2024

  • Towards Faithful Model Explanation in NLP

    Open Access•Qing Lyu, Marianna Apidianaki et al.•Computational Linguistics•2024

  • Measuring political violence in Pakistan

    Open Access•Ethan Bueno De Mesquita, Carol Christine Fair et al.•Conflict Management and Peace…•2014

  • Stance detection

    Open Access•Michael Burnham•Political Science Research and…•2025

  • Synthetically generated text for supervised text analysis

    Open Access•Andrew Halterman•Political Analysis•2025

  • Testing Causal Theories with Learned Proxies

    Open Access•David Knox, Dean Knox et al.•Annual Review of Political Science•2022

  • Can Large Language Models Transform Computational Social Science

    Open Access•Caleb Ziems, William A Held et al.•Computational Linguistics•2023

  • Text as Data

    Open Access•J Adams•Contemporary Sociology A Journal…•2024

  • Measurement Validity

    Open Access•Robert Adcock, Donald Collier et al.•American Political Science Review•2001

Obras citantes distintas8
Citas por año8
Intervalo de citas2025 - 2026 (2)
Velocidad de citacióncurrent
Altamente citadoNo
Tipos de citaNeutras: 8
Ethnos_APP • Proyecto Open Source • Licencia MIT • Frontend v2.0.0 • Privacidad y Cookies • Documentación de la API: api.ethnos.app/docs • Código de la API: GitHub • DOI: 10.5281/zenodo.17049435 • Código del Frontend: GitHub • DOI: 10.5281/zenodo.17050053 • cruz.rio.br • Expectantes Misericordiae