Skip to main content

ETHNOS_APP

Home • Search • Journals • List 0

Open-source LLMs for text annotation

A Practical Guide for Model Setting and Fine-Tuning

Bibliographic Data

ID7158796
AuthorsMeysam Alizadeh (0000-0001-6696-6471, University of Zurich, corresponding author), Mael Kubli (0000-0002-5592-9648, University of Zurich), Zeynab Samei (0000-0002-6802-5090, Institute for Research in Fundamental Sciences), Shirin Dehghani (Allameh Tabataba'i University), Mohammadmasiha Zahedivafa (Iran University of Science and Technology), Juan Diego Bermeo (0009-0000-6874-4848, University of Zurich), Maria Korobeynikova (University of Zurich), Fabrizio Gilardi (0000-0002-0635-3048, University of Zurich)
Year2025
Volume8
Issue1
Pages17-17
Publication date2025-02-01
Peer ReviewedYes
Open AccessYes
TypeARTICLE
VenueJournal of Computational Social Science (JOURNAL)
Journal identifiersISSN: 2432-2725 • E-ISSN: 2432-2717
PublisherSpringer Science and Business Media LLC (PUBLISHER)
DOI10.1007/s42001-024-00345-9
PMID39712076
OpenAlexW4405552797
LanguageEN
Citations received13
References cited22

This paper studies the performance of open-source Large Language Models (LLMs) in text classification tasks typical for political science research. By examining tasks like stance, topic, and relevance classification, we aim to guide scholars in making informed decisions about their use of LLMs for text analysis and to establish a baseline performance benchmark that demonstrates the models’ effectiveness. Specifically, we conduct an assessment of both zero-shot and fine-tuned LLMs across a range of text annotation tasks using news articles and tweets datasets. Our analysis shows that fine-tuning improves the performance of open-source LLMs, allowing them to match or even surpass zero-shot GPT $$-$$ - 3.5 and GPT-4, though still lagging behind fine-tuned GPT $$-$$ - 3.5. We further establish that fine-tuning is preferable to few-shot training with a relatively modest quantity of annotated text. Our findings show that fine-tuned open-source LLMs can be effectively deployed in a broad spectrum of text annotation applications. We provide a Python notebook facilitating the application of LLMs in text annotation for other researchers

Annotation · Cartography · Geography · Open source · Computational and Text Analysis Methods · Computer Science · Natural Language Processing Techniques · Topic Modeling · Artificial Intelligence

  • How central bank independence shapes monetary policy communication

    Open Access•Lauren Leek, Simeon Bischl•European Journal of Political…•2025

  • Representation of environmental issues

    Open Access•Clara Faulí Molas•European Union Politics•2026

  • Revisiting business associations in trade politics

    Open Access•Rodrigo Fagundes Cezar•Review of International Political…•2026

  • Introducing Halc

    Open Access•Andreas Reich, Claudia Thoms et al.•Communication Methods and Measures•2026

  • Measuring partisan community dynamics

    Open Access•Giada Marino, Bruna Almeida Paroni et al.•Information Communication & Society•2026

  • Prebunking false information in the wild

    Open Access•Yeaeun Gong, Agam Goyal et al.•International Journal of…•2026

  • How backsliding government parties defend democratic backsliding in the European Parliament

    Open Access•Thomas Winzen•European Union Politics•2026

  • Navigating Ambiguities

    Open Access•Xiao Meng, Xiaohui Wang et al.•Communication Methods and Measures•2026

  • Stay Tuned

    Open Access•Max Griswold, Michael W Robbins et al.•Political Analysis•2025

  • Improving hate speech detection with large language models

    Open Access•Natalia Umansky, Mael Kubli et al.•European Journal of Political…•2026

  • Advancing AI-Assisted Policy-Text Analysis

    Open Access•Ning Wu, Binwei Yang et al.•Journal of Planning Education and…•2026

  • Public perspectives on food date labeling

    Open Access•Peng Lü, Lulu Mao et al.•Food Policy•2026

  • Updating 'The Future of Coding

    Open Access•Nga Than, Leanne Fan et al.•Sociological Methods & Research•2025

  • ChatGPT

    Open Access•Partha Pratim Ray•Internet of Things and…•2023

  • Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead

    Open Access•Cynthia Rudin•Nature Machine Intelligence•2019

  • Pre-train, Prompt, and Predict

    Open Access•Pengfei Liu, Weizhe Yuan et al.•ACM Computing Surveys•2023

  • ChatGPT

    Open Access•Eva A M van Dis, Johan Bollen et al.•Nature•2023

  • ChatGPT outperforms crowd workers for text-annotation tasks

    Open Access•Fabrizio Gilardi, Meysam Alizadeh et al.•Proceedings of the National…•2023

  • How Negative Media Coverage Impacts Platform Governance

    Open Access•Nahema Marchal, Emma Hoes et al.•Political Communication•2025

  • Automated Text Classification of News Articles

    Open Access•Pablo Barberá Aresté, Amber E Boydstun et al.•Political Analysis•2021

  • Can Large Language Models Transform Computational Social Science

    Open Access•Caleb Ziems, William A Held et al.•Computational Linguistics•2023

Unique citing works13
Citations per year13
Citation span2025 - 2026 (2)
Citation velocitycurrent
Highly citedNo
Citation typesNeutral: 11
Ethnos_APP • Open Source Project • MIT License • Frontend v2.0.0 • Privacy and Cookies • API Documentation: api.ethnos.app/docs • API Source Code: GitHub • DOI: 10.5281/zenodo.17049435 • Frontend Source Code: GitHub • DOI: 10.5281/zenodo.17050053 • cruz.rio.br • Expectantes Misericordiae