Playing with machines
Using machine learning to understand automated copyright enforcement at scale
Bibliographic Data
| ID | 5260402 |
|---|---|
| Authors | Joanne E Gray, Joanne Gray (0000-0001-5425-4979, Queensland University of Technology), Nicolas Suzor (0000-0003-3029-0646, Queensland University of Technology, corresponding author) |
| Year | 2020 |
| Volume | 7 |
| Issue | 1 |
| Pages | 205395172091996 |
| Publication date | 2020-01-01 |
| Peer Reviewed | Yes |
| Open Access | Yes |
| Type | ARTICLE |
| Venue | Big Data & Society (JOURNAL) |
| Journal identifiers | ISSN: 2053-9517 • E-ISSN: 2053-9517 |
| Publisher | SAGE Publications Inc (PUBLISHER) |
| DOI | 10.1177/2053951720919963 |
| OpenAlex | W3023398882 |
| Language | EN |
| Citations received | 15 |
| References cited | 18 |
This article presents the results of methodological experimentation that utilises machine learning to investigate automated copyright enforcement on YouTube. Using a dataset of 76.7 million YouTube videos, we explore how digital and computational methods can be leveraged to better understand content moderation and copyright enforcement at a large scale.We used the BERT language model to train a machine learning classifier to identify videos in categories that reflect ongoing controversies in copyright takedowns. We use this to explore, in a granular way, how copyright is enforced on YouTube, using both statistical methods and qualitative analysis of our categorised dataset. We provide a large-scale systematic analysis of removals rates from Content ID's automated detection system and the largely automated, text search based, Digital Millennium Copyright Act notice and takedown system. These are complex systems that are often difficult to analyse, and YouTube only makes available data at high levels of abstraction. Our analysis provides a comparison of different types of automation in content moderation, and we show how these different systems play out across different categories of content. We hope that this work provides a methodological base for continued experimentation with the use of digital and computational methods to enable large-scale analysis of the operation of automated systems
Abstraction · Automation · Categorization · Computer security · Data science · Digital content · Enforcement · Law enforcement · Machine learning · World Wide Web · Computer Science · Engineering · Hate Speech and Cyberbullying Detection · Law in Society and Culture · Artificial Intelligence
Copyright Gossip
Inteligência Artificial, moderação de conteúdos no YouTube e a proteção de direitos
Digital Transformation and Cultural Policies in Europe
Uploaders' perceptions of the German implementation of the EU copyright reform and their preferences for copyright regulation
Copyright callouts and the promise of creator-driven platform governance
From content moderation to visibility moderation
Mandate to overblock? Understanding the impact of the European Union's Article 17 on copyright content moderation on YouTube
Legitimacy and space in the use of technologies for environmental and social governance
We Learn Through Mistakes”
From the virtual community to “Trust and Safety”
Steering society through algorithms
Quantitative questions on big data in translation studies
A safe, responsible, and profitable ecosystem of music”
Algorithmic regulation
Copyright Gossip
| Unique citing works | 15 |
|---|---|
| Citations per year | 2,5 |
| Citation span | 2020 - 2026 (7) |
| Citation velocity | current |
| Highly cited | No |
| Citation types | Neutral: 13 |