Lessons Learned Developing and Using a Machine Learning Model to Automatically Transcribe 2.3 Million Handwritten Occupation Codes
Bibliographic Data
| ID | 21244544 |
|---|---|
| Authors | Bjørn-Richard Pedersen (0009-0000-2363-0791, UiT The Arctic University of Norway, corresponding author), Einar Holsbø (0000-0002-9728-2088, UiT The Arctic University of Norway, corresponding author), Trygve Andersen (UiT The Arctic University of Norway, corresponding author), Nikita Shvetsov (UiT The Arctic University of Norway, corresponding author), Johan Ravn (corresponding author), Hilde Leikny Sommerseth (0000-0001-7070-8184, UiT The Arctic University of Norway, corresponding author), Lars Ailo Bongo (0000-0002-7544-2482, UiT The Arctic University of Norway, corresponding author) |
| Year | 2022 |
| Volume | 12 |
| Pages | 1-17 |
| Publication date | 2022-01-06 |
| Peer Reviewed | Yes |
| Open Access | Yes |
| Type | ARTICLE |
| Venue | Historical Life Course Studies (JOURNAL) |
| Journal identifiers | ISSN: 2352-6343 • E-ISSN: 2352-6343 |
| Publisher | International Institute of Social History (PUBLISHER • NL) |
| DOI | 10.51964/hlcs11331 |
| OpenAlex | W3216106108 |
| Language | EN |
| Citations received | 4 |
| References cited | 3 |
Machine learning approaches achieve high accuracy for text recognition and are therefore increasingly used for the transcription of handwritten historical sources. However, using machine learning in production requires a streamlined end-to-end pipeline that scales to the dataset size and a model that achieves high accuracy with few manual transcriptions. The correctness of the model results must also be verified. This paper describes our lessons learned developing, tuning and using the Occode end-to-end machine learning pipeline for transcribing 2.3 million handwritten occupation codes from the Norwegian 1950 population census. We achieve an accuracy of 97% for the automatically transcribed codes, and we send 3% of the codes for manual verification . We verify that the occupation code distribution found in our results matches the distribution found in our training data, which should be representative for the census as a whole. We believe our approach and lessons learned may be useful for other transcription projects that plan to use machine learning in production. The source code is available at https://github.com/uit-hdl/rhd-codes
Correctness · Machine learning · Natural language processing · Population · Programming language · Source code · Computer Science · Handwritten Text Recognition Techniques · Image Processing and 3D Reconstruction · Natural Language Processing Techniques · Artificial Intelligence
| Unique citing works | 4 |
|---|---|
| Citations per year | 1 |
| Citation span | 2022 - 2026 (5) |
| Citation velocity | current |
| Highly cited | No |
| Citation types | Neutral: 4 |