Automated Hate Speech Detection and the Problem of Offensive Language
Bibliographic Data
| ID | 23328286 |
|---|---|
| Authors | Thomas R Davidson (0000-0002-5947-7490, Cornell University), Thomas Davidson (0000-0002-9393-2573), Dana Warmsley (Cornell University), Michael W Macy (0000-0003-0024-5027, Cornell University), Michael Macy, Ingmar Weber (0000-0003-4169-2579, Hamad bin Khalifa University) |
| Year | 2017 |
| Volume | 11 |
| Issue | 1 |
| Pages | 512-515 |
| Publication date | 2017-05-03 |
| Peer Reviewed | Yes |
| Open Access | Yes |
| Type | ARTICLE |
| Venue | Proceedings of the International AAAI Conference on Web and Social Media (JOURNAL) |
| Journal identifiers | ISSN: 2162-3449 • E-ISSN: 2334-0770 |
| Publisher | Association for the Advancement of Artificial Intelligence (AAAI) (PUBLISHER) |
| DOI | 10.1609/icwsm.v11i1.14955 |
| OpenAlex | W2595653137 |
| Language | EN |
| Citations received | 140 |
| References cited | 4 |
A key challenge for automatic hate-speech detection on social media is the separation of hate speech from other instances of offensive language. Lexical detection methods tend to have low precision because they classify all messages containing particular terms as hate speech and previous work using supervised learning has failed to distinguish between the two categories. We used a crowd-sourced hate speech lexicon to collect tweets containing hate speech keywords. We use crowd-sourcing to label a sample of these tweets into three categories: those containing hate speech, only offensive language, and those with neither. We train a multi-class classifier to distinguish between these different categories. Close analysis of the predictions and the errors shows when we can reliably separate hate speech from other offensive language and when this differentiation is more difficult. We find that racist and homophobic tweets are more likely to be classified as hate speech but that sexist tweets are generally classified as offensive. Tweets without explicit hate keywords are also more difficult to classify.
Classifier (UML) · Lexicon · Natural language processing · Offensive · Speech processing · Speech recognition · Voice activity detection · Artificial Intelligence · Computer Science · Hate Speech and Cyberbullying Detection · Internet Traffic Analysis and Secure E-voting · Mathematics · Spam and Phishing Detection
Language as data, hate as task
‘Toxic’ memes
Towards countering hate speech against journalists on social media
Gendered hate speech in YouTube and YouNow comments
Constructive Aggression? Multiple Roles of Aggressive Content in Political Discourse on Russian YouTube
Measuring Causal Effects of Civil Communication without Randomization
A Visual Approach to Tracking Emotional Sentiment Dynamics in Social Network Commentaries
The Dynamics of Political Incivility on Twitter
Hate and Incivilities in Hashtags against Women Candidates in Chile (2021–2022)
Exploring the evolution of posting behavior and language use in a racially and ethnically motivated extremist forum
Govor mržnje u hrvatskom medijskom prostoru
A Critical Stylistic Study of Cyber Trolls’ Comments on the Al-Jazeera Arabic TV Channel’s YouTube Videos Concerning the Israeli-Iranian 2024 Conflict
Polarización, desinformación y expresiones de odio en Twitter. Caso grupos políticos nacionalistas e independentistas en España
ArewaAgainstLGBTQ discourse
Where's the harm? Screening student evaluations of teaching for offensive, threatening or distressing comments
On the social and technical challenges of Web search autosuggestion moderation
Classifying constructive comments
Redes, equipos de monitoreo y aplicaciones móvil para combatir los discursos y delitos de odio en Europa
A critical reflection on the use of toxicity detection algorithms in proactive content moderation systems
How AI Bots Have Reinforced Gender Bias in Hate Speech
Racism in tourism reviews
No2Sectarianism
The Risk of Racial Bias in Hate Speech Detection
Social Media and Democracy
The Evolution of the Manosphere across the Web
Misinformation, Disinformation, and Online Propaganda
A Survey on Automatic Detection of Hate Speech in Text
Social Media, Echo Chambers, and Political Polarization
Whose harm gets detected? A structured review and conceptual framework for misogyny detection and Dari–Pashto marginalization in AI content moderation
A comparative analysis of machine learning algorithms for hate speech detection in social media
The Role of Victim’s Resilience and Self-Esteem in Experiencing Internet Hate
Linguistic models of abusive language
Toxic language in online incel communities
Modeling aggression propagation on social media
Understanding Large Language Model Driven Social Bots
HostileNet
Adversarial NLP for Social Network Applications
Zero-Shot Hate to Non-Hate Text Conversion Using Lexical Constraints
HateThaiSent
Backdoor Attack and Defense on Deep Learning
Sehc
BiCapsHate
Detecting Offensive Language Based on Graph Attention Networks and Fusion Features
Model-Agnostic Meta-Learning for Multilingual Hate Speech Detection
How can hate narratives be tracked in online environments? A corpus-based study
Covid-19 and Sinophobia
How Online Content Providers Moderate User‐Generated Content to Prevent Harmful Online Communication
Battle for Britain
Hebrew offensive language taxonomy and dataset
An integrated explicit and implicit offensive language taxonomy
Offensive language in media discussion forums
Inequalities and content moderation
A Feature-Based Approach to Assess Hate Speech in User Comments
From Internet Meme to the Mainstream
Negative Feedback Fuels Hate Speech
The Real Cancel Culture
Classification of discussants on cyber-violence incidents
Hate in Word and Deed
Hidden Hate
Can we predict the Billboard music chart winner? Machine learning prediction based on Twitter artist-fan interactions
Fair compensation of crowdsourcing work
Expected behavioural effects of alerts to impolite online news commenters
Study on relationship between adversarial texts and language errors
Beyond Incivility
Social media content classification and community detection using deep learning and graph analytics
Comunicación en redes y discursos de odio en el contexto español
The expression of hate speech against Afro-descendant, Roma, and LGBTQ+ communities in YouTube comments
Intersectionality and the gendered discussion around Muslim Canadian politicians on Twitter
Offensive language in reactions to public figures in polarised discourse online
From Insult to Hate Speech
Moral Foundations Twitter Corpus
Perceptions and Evaluations of Incivility in Public Online Discussions—Insights From Focus Groups With Different Online Actors
Covid-19 Induced Misinformation on YouTube
Hate Speech Directed at Spanish Female Actors
Fueling Toxicity? Studying Deceitful Opinion Leaders and Behavioral Changes of Their Followers
Developing an Incivility Dictionary for German Online Discussions – a Semi-Automated Approach Combining Human and Artificial Knowledge
Promoting Hate Speech by Dehumanizing Metaphors of Immigration
Framing Migration in Southern European Media
The Role of Minority Political Groups in the Dissemination of Disinformation. The Case of Spain
Can (and should) LLMs perform critical discourse analysis
Doxxing to destroy
Who Leaves Malicious Comments on Online News? An Empirical Study in Korea
Evolution of negative visual frames of immigrants and refugees in the main media of Southern Europe
Análisis del discurso público y la toxicidad en X
Fighting Hate Speech, Silencing Drag Queens? Artificial Intelligence in Content Moderation and Risks to LGBTQ Voices Online
Nationally Representative, Locally Misaligned
Online Polarization and Violence in the United States
An Informed Neural Network for Discovering Historical Documentation Assisting the Repatriation of Indigenous Ancestral Human Remains
Mesurer l’empreinte antisémite sur YouTube
“Moral Turn” and the Entextualization of Homosexuals as Pedophiles in Bolsonaro’s Speeches in Congress (2000 to 2018)
“Virada Moral” E Entextualização Do Homossexual Como Pedófilo Em Falas De Bolsonaro No Congresso (2000 a 2018)
Twitch aggression profile
A survey on moral foundation theory and pre-trained language models
The podcast as the centre of young Colombians’ information consumption in the digital sonosphere
Cross-cutting interaction, inter-party hostility, and partisan identity
Technology acceptance and transparency demands for toxic language classification – interviews with moderators of public online discussion fora
Opposition as “A Mould on the Fatherland
Vaniraksha
Improving Hebrew offensive language classification using LLM-assisted human-in-the-loop annotation
The Power of Narrative
| Unique citing works | 140 |
|---|---|
| Citations per year | 17,5 |
| Citation span | 2018 - 2026 (9) |
| Citation velocity | current |
| Highly cited | Yes |
| Citation types | Neutral: 118 |