The Augmented Social Scientist
Using Sequential Transfer Learning to Annotate Millions of Texts with Human-Level Accuracy
Bibliographic Data
| ID | 2331135 |
|---|---|
| Authors | Salomé Do (0000-0002-6095-6253, ENS-Paris/PSL (LATTICE), Paris, France), E Ollion (0000-0003-3099-5240, Institut Polytechnique de Paris (CREST), Palaiseau, France, corresponding author), Rubing Shen (0000-0002-5504-6108, Sciences Po (Medialab), Paris, France) |
| Year | 2024 |
| Volume | 53 |
| Issue | 3 |
| Pages | 1167-1200 |
| Publication date | 2024-08-01 |
| Peer Reviewed | Yes |
| Open Access | Yes |
| Type | ARTICLE |
| Venue | Sociological Methods & Research (JOURNAL) |
| Journal identifiers | ISSN: 0049-1241 • E-ISSN: 1552-8294 |
| Publisher | SAGE Publications Inc (PUBLISHER) |
| DOI | 10.1177/00491241221134526 |
| OpenAlex | W4311510002 |
| Language | EN |
| Citations received | 19 |
| References cited | 33 |
The last decade witnessed a spectacular rise in the volume of available textual data. With this new abundance came the question of how to analyze it. In the social sciences, scholars mostly resorted to two well-established approaches, human annotation on sampled data on the one hand (either performed by the researcher, or outsourced to microworkers), and quantitative methods on the other. Each approach has its own merits - a potentially very fine-grained analysis for the former, a very scalable one for the latter - but the combination of these two properties has not yielded highly accurate results so far. Leveraging recent advances in sequential transfer learning, we demonstrate via an experiment that an expert can train a precise, efficient automatic classifier in a very limited amount of time. We also show that, under certain conditions, expert-trained models produce better annotations than humans themselves. We demonstrate these points using a classic research question in the sociology of journalism, the rise of a 'horse race' coverage of politics. We conclude that recent advances in transfer learning help us augment ourselves when analyzing unstructured data
Annotation · Big data · Classifier (UML) · Data mining · Data science · Machine learning · Scalability · Transfer of learning · Artificial Intelligence · Computational and Text Analysis Methods · Computer Science · Sentiment Analysis and Opinion Mining · Topic Modeling
The Insight-Inference Loop
Mining for Meaning
Theorizing and Measuring General Nonmarket Communication
Data-Campaigning on Facebook
Quantitative and Computational Approaches in the Social Studies of Economics
Navigating the Risks of Using Large Language Models for Text Annotation in Social Science Research
Aligning Emotions
Qualitative research with LLM chatbots
Au petit déjeuner, politique
Efficiency vs. understanding
Computational Text Analysis for Building and Testing Social Theory
Speculating with Machines
Time and Climate Change
Latin-American cyborg methods
Start Generating
Updating 'The Future of Coding
From Codebooks to Promptbooks
Correcting the Measurement Errors of AI-Assisted Labeling in Image Analysis Using Design-Based Supervised Learning
Large Language Models for Text Classification
The general inquirer
Exploring Textual Data
Transformers
The framing of politics as strategy and game
Deep Contextualized Word Representations
ImageNet classification with deep convolutional neural networks
Climbing towards NLU
Computational Social Science
Computational appraisal of gender representativeness in popular movies
Text as Data
Horse-Race Journalism
Automated Text Classification of News Articles
Text as Data
Big Data and historical social science
Crowd-sourced Text Analysis
How Censorship in China Allows Government Criticism but Silences Collective Expression
Ce que le big data fait à l'analyse sociologique des textes
The Future of Coding
Beyond Keywords
Networked but Commodified
| Unique citing works | 19 |
|---|---|
| Citations per year | 6,33 |
| Citation span | 2023 - 2026 (4) |
| Citation velocity | current |
| Highly cited | No |
| Citation types | Neutral: 18 |