Investigating Opinions on Public Policies in Digital Media
Setting up a Supervised Machine Learning Tool for Stance Classification
Bibliographic Data
| ID | 12971104 |
|---|---|
| Authors | Christina Viehmann (0000-0001-6673-0987, Johannes Gutenberg University Mainz, corresponding author), Tilman Beck (0000-0002-1403-8240), Marcus Maurer (0009-0000-4582-3623, Johannes Gutenberg University Mainz), Oliver Quiring (0000-0002-6671-583X, Johannes Gutenberg University Mainz), Iryna Gurevych (0000-0003-2187-7621) |
| Year | 2022 |
| Volume | 17 |
| Issue | 2 |
| Pages | 150-184 |
| Publication date | 2022-12-12 |
| Peer Reviewed | Yes |
| Open Access | No |
| Type | ARTICLE |
| Venue | Communication Methods and Measures (JOURNAL) |
| Journal identifiers | ISSN: 1931-2458 • E-ISSN: 1931-2466 |
| Publisher | Taylor & Francis (PUBLISHER • GB) |
| DOI | 10.1080/19312458.2022.2151579 |
| OpenAlex | W4311353657 |
| Language | EN |
| Citations received | 3 |
| References cited | 58 |
Supervised machine learning (SML) provides us with tools to efficiently scrutinize large corpora of communication texts. Yet, setting up such a tool involves plenty of decisions starting with the data needed for training, the selection of an algorithm, and the details of model training. We aim at establishing a firm link between communication research tasks and the corresponding state-of-the-art in natural language processing research by systematically comparing the performance of different automatic text analysis approaches. We do this for a challenging task – stance detection of opinions on policy measures to tackle the COVID-19 pandemic in Germany voiced on Twitter. Our results add evidence that pre-trained language models such as BERT outperform feature-based and other neural network approaches. Yet, the gains one can achieve differ greatly depending on the specific merits of pre-training (i.e., use of different language models). Adding to the robustness of our conclusions, we run a generalizability check with a different use case in terms of language and topic. Additionally, we illustrate how the amount and quality of training data affect model performance pointing to potential compensation effects. Based on our results, we derive important practical recommendations for setting up such SML tools to study communication texts
Artificial neural network · Feature selection · Generalizability theory · Language model · Machine learning · Natural language processing · Robustness (evolution · Task (project management · Computational and Text Analysis Methods · Computer Science · Sentiment Analysis and Opinion Mining · Topic Modeling · Artificial Intelligence
Content analysis in communication research
Analyzing Media Messages
Enriching Word Vectors with Subword Information
Machine learning
The Value of Big Data in Digital Media Research
Deep Contextualized Word Representations
Taking Stock of the Toolkit
ImageNet classification with deep convolutional neural networks
Long Short-Term Memory
Echo Chamber or Public Sphere? Predicting Political Orientation and Measuring Political Homophily in Twitter Using Big Data
Supervised Machine Learning for Text Analysis in R
Assumptions behind Intercoder Reliability Indices
Extracting Latent Moral Information from Text Narratives
Teaching the Computer to Code Frames in News
Using Supervised Machine Learning in Automated Content Analysis
Better Crowdcoding
What’s the Tone? Easy Doesn’t Do It
More than Bags of Words
Measuring Moral Rhetoric in Text
Dictionaries, Supervised Learning, and Media Coverage of Public Policy
Are Experts (News)Worthy? Balance, Conflict, and Mass Media Coverage of Expert Consensus
What's in a Frame? A Content Analysis of Media Framing Studies in the World's Leading Communication Journals, 1990-2005
In Validations We Trust? The Impact of Imperfect Human Annotations as a Gold Standard on the Quality of Validation of Automated Content Analysis
Text as Data
| Unique citing works | 3 |
|---|---|
| Citations per year | 1 |
| Citation span | 2023 - 2026 (4) |
| Citation velocity | current |
| Highly cited | No |
| Citation types | Neutral: 3 |