Codebook LLMs
Evaluating LLMs as Measurement Tools for Political Science Concepts
Bibliographic Data
| ID | 6332293 |
|---|---|
| Authors | Andrew Halterman (0000-0001-9716-9555, Michigan State University, corresponding author), Katherine A Keith (0000-0002-8101-4572, Williams College) |
| Year | 2025 |
| Pages | 1-17 |
| Publication date | 2025-09-19 |
| Peer Reviewed | Yes |
| Open Access | Yes |
| Type | ARTICLE |
| Venue | Political Analysis (JOURNAL) |
| Journal identifiers | ISSN: 1047-1987 • E-ISSN: 1476-4989 |
| Publisher | Cambridge University Press (CUP) (PUBLISHER) |
| DOI | 10.1017/pan.2025.10017 |
| OpenAlex | W4414362097 |
| Language | EN |
| Citations received | 8 |
| References cited | 23 |
Codebooks—documents that operationalize concepts and outline annotation procedures—are used almost universally by social scientists when coding political texts. To code these texts automatically, researchers are increasingly turning to generative large language models (LLMs). However, there is limited empirical evidence on whether “off-the-shelf” LLMs faithfully follow real-world codebook operationalizations and measure complex political constructs with sufficient accuracy. To address this, we gather and curate three real-world political science codebooks—covering protest events, political violence, and manifestos—along with their unstructured texts and human-coded labels. We also propose a five-stage framework for codebook-LLM measurement: Preparing a codebook for both humans and LLMs, testing LLMs’ basic capabilities on a codebook, evaluating zero-shot measurement accuracy (i.e., off-the-shelf performance), analyzing errors, and further (parameter-efficient) supervised training of LLMs. We provide an empirical demonstration of this framework using our three codebook datasets and several pre-trained 7–12 billion open-weight LLMs. We find current open-weight LLMs have limitations in following codebooks zero-shot, but that supervised instruction-tuning can substantially improve performance. Rather than suggesting the “best” LLM, our contribution lies in our codebook datasets, evaluation framework, and guidance for applied researchers who wish to implement their own codebook-LLM measurement projects
Artificial Intelligence in Law · Legal Education and Practice Innovations
Using LLMs for measurement in diplomatic speeches
NEClass
Piercing a Methodological Bubble
The Stories Individuals “Like”
Text as Data and Causal Inference in Sociology
Disaster, Distributive Politics, and the Persistence of Partisan Divides in Climate Policy
Cops and crypto
Fine-tuned large language models can replicate expert coding better than trained coders
ChatGPT outperforms crowd workers for text-annotation tasks
The Measurement of Observer Agreement for Categorical Data
Do AIs know what the most important issue is? Using language models to code open-text social survey responses at scale
Large language models as a substitute for human experts in annotating political text
Towards Faithful Model Explanation in NLP
Measuring political violence in Pakistan
Stance detection
Synthetically generated text for supervised text analysis
Testing Causal Theories with Learned Proxies
Can Large Language Models Transform Computational Social Science
Text as Data
Measurement Validity
| Unique citing works | 8 |
|---|---|
| Citations per year | 8 |
| Citation span | 2025 - 2026 (2) |
| Citation velocity | current |
| Highly cited | No |
| Citation types | Neutral: 8 |