Skip to main content

ETHNOS_APP

Home • Search • Journals • List 0

Compounded Mediation

A Data Archaeology of the Newspaper Navigator Dataset

Bibliographic Data

ID21835751
AuthorsBenjamin Lee (0000-0002-1171-4741), Benjamin Charles Germain Lee (0000-0002-1677-6386, corresponding author)
Year2021
Volume15
Issue4
Publication date2021-12-07
Peer ReviewedYes
Open AccessYes
TypeARTICLE
VenueDigital humanities quarterly (JOURNAL)
Journal identifiersISSN: 1938-4122 • E-ISSN: 1938-4122
PublisherThe Association for Computers and the Humanities (PUBLISHER)
DOI10.63744/edw69jxpdrfy
OpenAlexW3082798741
LanguageEN
Citations received3

The increasing roles of machine learning and artificial intelligence in the construction of cultural heritage and humanities datasets necessitate critical examination of the myriad biases introduced by machines, algorithms, and the humans who build and deploy them. From image classification to optical character recognition, the effects of decisions ostensibly made by machines compound through the digitization pipeline and redouble in each step, mediating our interactions with digitally-rendered artifacts through the search and discovery process. As a result, scholars within the digital humanities community have begun advocating for the proper contextualization of cultural heritage datasets within the socio-technical systems in which they are created and utilized. One such approach to this contextualization is the data archaeology, a form of humanistic excavation of a dataset that Paul Fyfe defines as recover[ing] and reconstitut[ing] media objects within their changing ecologies . Within critical data studies, this excavation of a dataset - including its construction and mediation via machine learning - has proven to be a capacious approach. However, the data archaeology has yet to be adopted as standard practice among cultural heritage practitioners who produce such datasets with machine learning. In this article, I present a data archaeology of the Library of Congress’s Newspaper Navigator dataset, which I created as part of the Library of Congress’s Innovator in Residence program . The dataset consists of visual content extracted from 16 million historic newspaper pages in the Chronicling America database using machine learning techniques. In this case study, I examine the manifold ways in which a Chronicling America newspaper page is transmuted and decontextualized during its journey from a physical artifact to a series of probabilistic photographs, illustrations, maps, comics, cartoons, headlines, and advertisements in the Newspaper Navigator dataset . Accordingly, I draw from fields of scholarship including media archaeology, critical data studies, science and technology studies, and the autoethnography throughout. To excavate the Newspaper Navigator dataset, I consider the digitization journeys of four different pages in Black newspapers included in Chronicling America, all of which reproduce the same photograph of W.E.B. Du Bois in an article announcing the launch of The Crisis, the official magazine of the NAACP. In tracing the newspaper pages’ journeys, I unpack how each step in the Chronicling America and Newspaper Navigator pipelines, such as the imaging process and the construction of training data, not only imprints bias on the resulting Newspaper Navigator dataset but also propagates the bias through the pipeline via the machine learning algorithms employed. Along the way, I investigate the limitations of the Newspaper Navigator dataset and machine learning techniques more generally as they relate to cultural heritage, with a particular focus on marginalization and erasure via algorithmic bias, which implicitly rewrites the archive itself. In presenting this case study, I argue for the value of the data archaeology as a mechanism for contextualizing and critically examining cultural heritage datasets within the communities that create, release, and utilize them. I offer this autoethnographic investigation of the Newspaper Navigator dataset in the hope that it will be considered not only by users of this dataset in particular but also by digital humanities practitioners and end users of cultural heritage datasets writ large

Archaeology · Data science · Media studies · Mediation · Newspaper · Political science · Sociology · Computer Science · Digital Humanities and Scholarship · History · Law

  • Powell.pps

    Open Access•Trevor Owens, Benjamin Charles Germain Lee et al.•Internet Histories•2025

  • What We Didn’t Know a Recipe Could Be

    Open Access•Avery Blankenship•Journal of Cultural Analytics•2024

  • The Digital Humanities and the Ladino Press

    Open Access•Benjamin Charles Germain Lee•Jewish Studies in the Digital Age•2022

Unique citing works3
Citations per year0,75
Citation span2022 - 2025 (4)
Citation velocityrecent
Highly citedNo
Citation typesNeutral: 1
Ethnos_APP • Open Source Project • MIT License • Frontend v2.0.0 • Privacy and Cookies • API Documentation: api.ethnos.app/docs • API Source Code: GitHub • DOI: 10.5281/zenodo.17049435 • Frontend Source Code: GitHub • DOI: 10.5281/zenodo.17050053 • cruz.rio.br • Expectantes Misericordiae