Improving the Accuracy of Tesseract 4.0 OCR Engine Using Convolution-Based Preprocessing-Reference-Cited by-同舟云学术

Improving the Accuracy of Tesseract 4.0 OCR Engine Using Convolution-Based Preprocessing

Published:2020-05-02 Issue:5 Volume:12 Page:715
ISSN:2073-8994
Container-title:Symmetry
language:en
Short-container-title:Symmetry

Author:

Sporici Dan^ORCID,Cușnir Elena,Boiangiu Costin-Anton^ORCID

Abstract

Optical Character Recognition (OCR) is the process of identifying and converting texts rendered in images using pixels to a more computer-friendly representation. The presented work aims to prove that the accuracy of the Tesseract 4.0 OCR engine can be further enhanced by employing convolution-based preprocessing using specific kernels. As Tesseract 4.0 has proven great performance when evaluated against a favorable input, its capability of properly detecting and identifying characters in more realistic, unfriendly images is questioned. The article proposes an adaptive image preprocessing step guided by a reinforcement learning model, which attempts to minimize the edit distance between the recognized text and the ground truth. It is shown that this approach can boost the character-level accuracy of Tesseract 4.0 from 0.134 to 0.616 (+359% relative change) and the F1 score from 0.163 to 0.729 (+347% relative change) on a dataset that is considered challenging by its authors.

Funder

Unitatea Executiva pentru Finantarea Invatamantului Superior, a Cercetarii, Dezvoltarii si Inovarii

Publisher

MDPI AG

Subject

Physics and Astronomy (miscellaneous),General Mathematics,Chemistry (miscellaneous),Computer Science (miscellaneous)

Link

https://www.mdpi.com/2073-8994/12/5/715/pdf

Reference27 articles.

1. Optical Character Recognition by Open source OCR Tool Tesseract: A Case Study

2. TESSERACT(1) Manual Pagehttps://github.com/tesseract-ocr/tesseract/blob/master/doc/tesseract.1.asc

3. OCR Accuracy Improvement on Document Images through a Novel Pre-Processing Approach;Harraj;arXiv,2015

Cited by 30 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Image Recognition Method and Device for Trace Scenarios Based on LSTM Neural Network;2024 6th International Conference on Communications, Information System and Computer Engineering (CISCE);2024-05-10

2. A comprehensive dataset of environmentally contaminated sites in the state of São Paulo in Brazil;Scientific Data;2024-03-02

3. Automatic Text Recognition from Image Dataset Using Optical Character Recognition and Deep Learning Techniques;Lecture Notes in Electrical Engineering;2024

4. Reduction of Throughput Time in Digital Publishing Using AI-Based Smart Systems;Lecture Notes in Networks and Systems;2024

5. Reanalyst: Scalable Analysis of Reverse Engineering Activities;2024