Toward a Period-specific Optimized Neural Network for OCR Error Correction of Historical Hebrew Texts-Reference-Cited by-同舟云学术

Toward a Period-specific Optimized Neural Network for OCR Error Correction of Historical Hebrew Texts

Published:2022-04-07 Issue:2 Volume:15 Page:1-20
ISSN:1556-4673
Container-title:Journal on Computing and Cultural Heritage
language:en
Short-container-title:J. Comput. Cult. Herit.

Author:

Suissa Omri¹^ORCID,Zhitomirsky-Geffet Maayan¹,Elmalech Avshalom¹

Affiliation:

1. Bar Ilan University, Ramat Gan, Israel

Abstract

Over the past few decades, large archives of paper-based historical documents, such as books and newspapers, have been digitized using the Optical Character Recognition (OCR) technology. Unfortunately, this broadly used technology is error-prone, especially when an OCRed document was written hundreds of years ago. Neural networks have shown great success in solving various text processing tasks, including OCR post-correction. The main disadvantage of using neural networks for historical corpora is the lack of sufficiently large training datasets they require to learn from, especially for morphologically rich languages like Hebrew. Moreover, it is not clear what are the optimal structure and values of hyperparameters (predefined parameters) of neural networks for OCR error correction in Hebrew due to its unique features. Furthermore, languages change across genres and periods. These changes may affect the accuracy of OCR post-correction neural network models. To overcome these challenges, we developed a new multi-phase method for generating artificial training datasets with OCR errors and hyperparameters’ optimization for building an effective neural network for OCR post-correction in Hebrew. To evaluate the proposed approach, a series of experiments using several literary Hebrew corpora from various periods and genres were conducted. The obtained results demonstrate that (1) training a network on texts from a similar period dramatically improves the network's ability to fix OCR errors, (2) using the proposed error injection algorithm, based on character-level period-specific errors, minimizes the need for manually corrected data and improves the network accuracy by 9%, (3) the optimized network design improves the accuracy by 3% compared to the state-of-the-art network, and (4) the constructed optimized network outperforms neural machine translation models and industry-leading spellcheckers. The proposed methodology may have practical implications for digital humanities projects that aim to search and analyze OCRed documents in Hebrew and potentially other morphologically rich languages.

Publisher

Association for Computing Machinery (ACM)

Subject

Computer Graphics and Computer-Aided Design,Computer Science Applications,Information Systems,Conservation

Link

https://dl.acm.org/doi/pdf/10.1145/3479159

Reference55 articles.

1. Word Error Rate Estimation for Speech Recognition: e-WER

2. Supervised OCR error detection and correction using statistical and neural machine translation methods;Amrhein C.;J. Lang. Technol. Comput. Linguist.,2018

3. Soylent

4. Two-Point Step Size Gradient Methods

5. ICDAR2017 Competition on Post-OCR Text Correction

Cited by 6 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Fire, Vulcanus , Archeus , and Alchemy: A Hybrid Close-Distant Reading of Paracelsus’s Thought on Active Agents;Ambix;2024-07-26

2. Mapping the landscape and roadmap of geospatial artificial intelligence (GeoAI) in quantitative human geography: An extensive systematic review;International Journal of Applied Earth Observation and Geoinformation;2024-04

3. From Digitization and Images to Text and Content: Transkribus as a Case Study;Manuscript Studies: A Journal of the Schoenberg Institute for Manuscript Studies;2024-03

4. Machine Translation for Historical Research: A case study of Aramaic-Ancient Hebrew Translations;Journal on Computing and Cultural Heritage;2023-10-16

5. Around the GLOBE: Numerical Aggregation Question-answering on Heterogeneous Genealogical Knowledge Graphs with Deep Neural Networks;Journal on Computing and Cultural Heritage;2023-08-09