Emily Maemura
Biographic Data
| ID | 6704050 |
|---|---|
| NAME | Emily Maemura |
| GIVEN NAMES | Emily |
| FAMILY NAME | Maemura |
| SIGNATURE | MAEMURA E |
| AFFILIATIONS | University of Illinois Urbana-Champaign |
| ORCID | 0000-0002-9329-7995 |
| VERIFIED | Yes |
| TOTAL WORKS | 6 |
| TOTAL CITATIONS | 2 |
| AUTHOR COUNT | 6 |
| EDITOR COUNT | 0 |
| FIRST PUBLICATION YEAR | 2023 |
| LATEST PUBLICATION YEAR | 2025 |
| H-INDEX | 1 |
From Ia_archiver to Openai: The Pasts and Futures of Automated Data Scrapers
Data scraping practices have recently come under scrutiny, as datasets scraped from the web’s social spaces are the basis of new generative AI tools like Google’s Gemini, Microsoft’s Copilot, and OpenAI’s ChatGPT. These practices of scrapers and crawlers are based on the conception of the internet as a mountain of data that’s sitting, waiting, available to be acted upon, extracted and put to use. In this paper, we examine the robots.txt exclusion…
Infrastructural consent: Robots.txt as a protocol for automated data extraction
For the past 30 years, the Robots Exclusion Protocol (REP, or robots.txt) has functioned underneath the surface of the web as the most efficient and widely used mechanism stopping web crawlers indexing or scraping data from a website. Recently, this plain text file took on a new role mediating how AI industries access massive amounts of publicly available web data. Robots.txt became touted as the best tool to prevent AI data scraping, producing w…
Conceptualizing aggregate-level description in web archives
Web archives collections are often excluded from archival science discussions, and their description instead focuses on bibliographic approaches to item-level metadata. This article argues that web archives are best understood using approaches of archival description, focusing on a case study of the Danish Netarchive, a long-running national web archive. By capturing and preserving web sites for the purposes of legal deposit, the Netarchive creat…
Web Histories in the Making: Web Archives & the Logics of Practice
Historically-situated accounts of the Web have a long history within the field of internet studies. Drawing on diverse methodologies and forms of data, web histories of platforms, cultures and communities of practice have illuminated the rich, but often transient and shifting nature of life online. Many web histories rely upon researchers capturing, collecting, and generating their own data through time, though some have also engaged with web arc…
All Warc and no playback: The materialities of data-centered web archives research
This paper examines the Web ARChive (WARC) file format, revealing how the format has come to play a central role in the development and standardization of interoperable tools and methods for the international web archiving community. In the context of emerging big data approaches, I consider the sociotechnical relationships between material construction of data and information infrastructures for collecting and research. Analysis is inspired by S…
Sorting URLs out: Seeing the web through infrastructural inversion of archival crawling
Web archives collections have become important sources for Internet scholars by documenting the past versions of web resources. Understanding how these collections are created and curated is of increasing concern and recent web archives scholarship has studied how the artefacts stored in archives represent specific curatorial choices and collecting practices. This paper takes a novel approach in studying web archiving practice, by focusing on the…
Sorting URLs out: Seeing the web through infrastructural inversion of archival crawling
Web archives collections have become important sources for Internet scholars by documenting the past versions of web resources. Understanding how these collections are created and curated is of increasing concern and recent web archives scholarship has studied how the artefacts stored in archives represent specific curatorial choices and collecting practices. This paper takes a novel approach in studying web archiving practice, by focusing on the…
Web Histories in the Making: Web Archives & the Logics of Practice
Historically-situated accounts of the Web have a long history within the field of internet studies. Drawing on diverse methodologies and forms of data, web histories of platforms, cultures and communities of practice have illuminated the rich, but often transient and shifting nature of life online. Many web histories rely upon researchers capturing, collecting, and generating their own data through time, though some have also engaged with web arc…
All Warc and no playback: The materialities of data-centered web archives research
This paper examines the Web ARChive (WARC) file format, revealing how the format has come to play a central role in the development and standardization of interoperable tools and methods for the international web archiving community. In the context of emerging big data approaches, I consider the sociotechnical relationships between material construction of data and information infrastructures for collecting and research. Analysis is inspired by S…
Sorting URLs out: Seeing the web through infrastructural inversion of archival crawling
Web archives collections have become important sources for Internet scholars by documenting the past versions of web resources. Understanding how these collections are created and curated is of increasing concern and recent web archives scholarship has studied how the artefacts stored in archives represent specific curatorial choices and collecting practices. This paper takes a novel approach in studying web archiving practice, by focusing on the…
From Ia_archiver to Openai: The Pasts and Futures of Automated Data Scrapers
Data scraping practices have recently come under scrutiny, as datasets scraped from the web’s social spaces are the basis of new generative AI tools like Google’s Gemini, Microsoft’s Copilot, and OpenAI’s ChatGPT. These practices of scrapers and crawlers are based on the conception of the internet as a mountain of data that’s sitting, waiting, available to be acted upon, extracted and put to use. In this paper, we examine the robots.txt exclusion…
Infrastructural consent: Robots.txt as a protocol for automated data extraction
For the past 30 years, the Robots Exclusion Protocol (REP, or robots.txt) has functioned underneath the surface of the web as the most efficient and widely used mechanism stopping web crawlers indexing or scraping data from a website. Recently, this plain text file took on a new role mediating how AI industries access massive amounts of publicly available web data. Robots.txt became touted as the best tool to prevent AI data scraping, producing w…
Conceptualizing aggregate-level description in web archives
Web archives collections are often excluded from archival science discussions, and their description instead focuses on bibliographic approaches to item-level metadata. This article argues that web archives are best understood using approaches of archival description, focusing on a case study of the Danish Netarchive, a long-running national web archive. By capturing and preserving web sites for the purposes of legal deposit, the Netarchive creat…
Computer Science (4 works) · World Wide Web (4 works) · Data science (3 works) · Web Data Mining and Analysis (3 works) · Digital and Traditional Archives Management (2 works) · Scientific Computing and Data Management (2 works) · Web application (2 works) · Advanced Data Storage Technologies (1 works) · Aggregate (composite (1 works) · Algorithm (1 works)