Disentangling User Samples
A Supervised Machine Learning Approach to Proxy-population Mismatch in Twitter Research
Bibliographic Data
| ID | 12971121 |
|---|---|
| Authors | K Hazel Kwon (0000-0001-7414-6959, Arizona State University, corresponding author), J Hunter Priniski (0000-0001-8061-777X, Arizona State University), Mónica Chadha (0000-0003-2106-441X, Arizona State University) |
| Year | 2018 |
| Volume | 12 |
| Issue | 2-3 |
| Pages | 216-237 |
| Publication date | 2018-02-15 |
| Peer Reviewed | Yes |
| Open Access | No |
| Type | ARTICLE |
| Venue | Communication Methods and Measures (JOURNAL) |
| Journal identifiers | ISSN: 1931-2458 • E-ISSN: 1931-2466 |
| Publisher | Taylor & Francis (PUBLISHER • GB) |
| DOI | 10.1080/19312458.2018.1430755 |
| OpenAlex | W2789379401 |
| Language | EN |
| Citations received | 7 |
| References cited | 30 |
This study addresses the issue of sampling biases in social media data-driven communication research. The authors demonstrate how supervised machine learning could reduce Twitter sampling bias induced from “proxy-population mismatch”. Particularly, this study used the Random Forest (RF) classifier to disentangle tweet samples representative of general publics’ activities from non-general—or institutional—activities. By applying RF classifier models to Twitter data sets relevant to four news events and a randomly pooled dataset, the study finds systematic differences between general user samples and institutional user samples in their messaging patterns. This article calls for disentangling Twitter user samples when ordinary user behaviors are the focus of research. It also builds on the development of machine learning modeling in the context of communication research
Classifier (UML · Data science · Machine learning · Population · Proxy (statistics · Random forest · Social media · World Wide Web · Complex Network Analysis Techniques · Computer Science · Opinion Dynamics and Social Influence · Social Media and Politics · Artificial Intelligence
Understanding Twitter conversations about artificial intelligence in advertising based on natural language processing
Trans vocabularies
Crisis in Mexico
Fake thumbs in play
Automated object detection in mobile eye-tracking research
Computational Contributions
Introduction to Neural Transfer Learning With Transformers for Social Science Text Analysis
Bit by Bit
Affective Publics
Classification and regression trees
Social media for large studies of behavior
Online Human-Bot Interactions
Random Forests
Echo Chamber or Public Sphere? Predicting Political Orientation and Measuring Political Homophily in Twitter Using Big Data
Dual Screening the Political
Hijacking #myNypd
Do We Tweet Differently From Our Mobile Devices? A Study of Language Differences on Mobile and Web-Based Twitter Platforms
Partisan Selective Sharing
Are You Scared Yet? Evaluating Fear Appeal Messages in Tweets About the Tips Campaign
Assessing structural correlates to social capital in Facebook ego networks
Assessing the bias in samples of large online networks
Twittering the News
Proximity and Terrorism News in Social Media
Frequency or Skillfulness
The tweet smell of celebrity success
Birds of a Feather Tweet Together
Network Issue Agendas on Twitter During the 2012 U.S. Presidential Election
Critical Questions for Big Data
Is Bigger Always Better? Potential Biases of Big Data Derived from Social Network Sites
Big Data and the brave new world of social media research
After the crisis? Big Data and the methodological challenges of empirical sociology
Social Network Influence on Online Behavioral Choices
| Unique citing works | 7 |
|---|---|
| Citations per year | 1 |
| Citation span | 2019 - 2024 (6) |
| Citation velocity | recent |
| Highly cited | No |
| Citation types | Neutral: 7 |