Improving Probabilistic Models In Text Classification Via Active Learning
Bibliographic Data
| ID | 3273399 |
|---|---|
| Authors | Mitchell Bosley (0000-0002-9172-966X, University of Toronto, corresponding author), Saki Kuzushima (0000-0003-3014-5203, Harvard University, corresponding author), Ted Enamorado (0000-0002-2022-7646, Washington University in St. Louis, corresponding author), Yuki Shiraito (0000-0003-0264-1138, Michigan United, corresponding author) |
| Year | 2025 |
| Volume | 119 |
| Issue | 2 |
| Pages | 985-1002 |
| Publication date | 2025-05-01 |
| Peer Reviewed | Yes |
| Open Access | Yes |
| Type | ARTICLE |
| Venue | American Political Science Review (JOURNAL) |
| Journal identifiers | ISSN: 0003-0554 • E-ISSN: 1537-5943 |
| Publisher | Cambridge University Press (CUP) (PUBLISHER) |
| DOI | 10.1017/s0003055424000716 |
| OpenAlex | W4401340787 |
| Language | EN |
| Citations received | 3 |
| References cited | 36 |
Social scientists often classify text documents to use the resulting labels as an outcome or a predictor in empirical research. Automated text classification has become a standard tool since it requires less human coding. However, scholars still need many human-labeled documents for training. To reduce labeling costs, we propose a new algorithm for text classification that combines a probabilistic model with active learning. The probabilistic model uses both labeled and unlabeled data, and active learning concentrates labeling efforts on difficult documents to classify. Our validation study shows that with few labeled data, the classification performance of our algorithm is comparable to state-of-the-art methods at a fraction of the computational cost. We replicate the results of two published articles with only a small fraction of the original labeled data used in those studies and provide open-source software to implement our method
Active learning (machine learning · Machine learning · Probabilistic logic · Advanced Text Analysis Techniques · Artificial Intelligence · Computer Science · Machine Learning and Data Classification · Topic Modeling
Making the News
Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead
A Survey of Methods for Explaining Black Box Models
Energy and Policy Considerations for Deep Learning in NLP
Finite Mixture Models
Maximum Likelihood from Incomplete Data Via the EM Algorithm
Echo Chamber or Public Sphere? Predicting Political Orientation and Measuring Political Homophily in Twitter Using Big Data
Deadly Clerics
Electoral Reform and National Security in Japan
Computer-Assisted Topic Classification for Mixed-Methods Social Science Research
Active Learning Approaches for Labeling Text
Use of force and civil–military relations in Russia
Machine Learning Predictions as Regression Covariates
Estimating Spatial Preferences from Votes and Text
Classification Accuracy as a Substantive Quantity of Interest
Machine Learning Human Rights and Wrongs
Disaggregating Repression
Testing Causal Theories with Learned Proxies
Repression Technology
Keyword‐Assisted Topic Models
Whose Ideas? Whose Words? Authorship of Ronald Reagan's Radio Addresses
Text as Data
Electoral Accountability and Particularistic Legislation
Human Rights are (Increasingly) Plural
Who Polices the Administrative State
How the Chinese Government Fabricates Social Media Posts for Strategic Distraction, Not Engaged Argument
| Unique citing works | 3 |
|---|---|
| Citations per year | 3 |
| Citation span | 2025 - 2026 (2) |
| Citation velocity | current |
| Highly cited | No |
| Citation types | Neutral: 2 |