Skip to main content

ETHNOS_APP

Home • Search • Journals • List 0

Hana

A handwritten name database for offline handwritten text recognition

Bibliographic Data

ID9814747
AuthorsChristian M Dahl (0000-0001-9880-2818, University of Southern Denmark, corresponding author), Torben S D Johansen (0000-0002-5964-5052, University of Southern Denmark), Emil N Sørensen (0000-0002-2492-7850, University of Bristol), Simon Wittrock (0000-0003-2879-575X, University of Southern Denmark)
Year2023
Volume87
Pages101473
Publication date2023-01-01
Peer ReviewedYes
Open AccessYes
TypeARTICLE
VenueExplorations in Economic History (JOURNAL)
Journal identifiersISSN: 0014-4983 • E-ISSN: 1090-2457
PublisherElsevier BV (PUBLISHER)
DOI10.1016/j.eeh.2022.101473
OpenAlexW3199199880
LanguageEN
References cited8

Methods for linking individuals across historical data sets, typically in combination with AI based transcription models, are developing rapidly. Perhaps the single most important identifier for linking is personal names. However, personal names are prone to enumeration and transcription errors and although modern linking methods are designed to handle such challenges, these sources of errors are critical and should be minimized. For this purpose, improved transcription methods and large-scale databases are crucial components. This paper describes and provides documentation for HANA, a newly constructed large-scale database which consists of more than 3.3 million names. The database contains more than 105 thousand unique names with a total of more than 1.1 million images of personal names, which proves useful for transfer learning to other settings. We provide three examples hereof, obtaining significantly improved transcription accuracy on both Danish and US census data. In addition, we present benchmark results for deep learning models automatically transcribing the personal names from the scanned documents. Through making more challenging large-scale databases publicly available we hope to foster more sophisticated, accurate, and robust models for handwritten text recognition

Benchmark (surveying · Database · Documentation · Identifier · Information retrieval · Natural language processing · Transcription (linguistics · Artificial Intelligence · Computer Science · Handwritten Text Recognition Techniques · Natural Language Processing Techniques · Topic Modeling

  • How Well Do Automated Linking Methods Perform? Lessons from US Historical Data

    Martha J Bailey, Martha Bailey et al.•Journal of Economic Literature•2020

  • Automated Linking of Historical Data

    Ran Abramitzky, Leah Boustan et al.•Journal of Economic Literature•2021

  • Europe's Tired, Poor, Huddled Masses

    Ran Abramitzky, Leah Platt Boustan et al.•American Economic Review•2012

  • Have the poor always been less likely to migrate? Evidence from inheritance practices during the age of mass migration

    Ran Abramitzky, Leah Platt Boustan et al.•Journal of Development Economics•2012

  • Linking individuals across historical sources

    Ran Abramitzky, Roy Mill et al.•Historical Methods A Journal of…•2020

  • Playing with matches

    Catherine Massey, Catherine G Massey•Historical Methods A Journal of…•2017

  • A Nation of Immigrants

    Ran Abramitzky, Leah Platt Boustan et al.•Journal of Political Economy•2014

  • Multiple Measures of Historical Intergenerational Mobility

    Open Access•James J Feigenbaum•The Economic Journal•2018

Citation velocityhistorical
Highly citedNo

Tools

Open DOIOpen Access
Ethnos_APP • Open Source Project • MIT License • Frontend v2.0.0 • Privacy and Cookies • API Documentation: api.ethnos.app/docs • API Source Code: GitHub • DOI: 10.5281/zenodo.17049435 • Frontend Source Code: GitHub • DOI: 10.5281/zenodo.17050053 • cruz.rio.br • Expectantes Misericordiae