Can Morphological Analyzers Improve the Quality of Optical Character Recognition?
-
Published:2015-06-17
Issue:2
Volume:
Page:45
-
ISSN:2387-3086
-
Container-title:Septentrio Conference Series
-
language:
-
Short-container-title:SCS
Author:
Silfverberg Miikka,Rueter Jack
Abstract
Optical Character Recognition (OCR) can substantially improve the usability of digitized documents. Language modeling using word lists is known to improve OCR quality for English. For morphologically rich languages, however, even large word lists do not reach high coverage on unseen text. Morphological analyzers offer a more sophisticated approach, which is useful in many language processing applications. is paper investigates language modeling in the open-source OCR engine Tesseract using morphological analyzers. We present experiments on two Uralic languages Finnish and Erzya. According to our experiments, word lists may still be superior to morphological analyzers in OCR even for languages with rich morphology. Our error analysis indicates that morphological analyzers can cause a large amount of real word OCR errors.
Publisher
UiT The Arctic University of Norway
Cited by
1 articles.
订阅此论文施引文献
订阅此论文施引文献,注册后可以免费订阅5篇论文的施引文献,订阅后可以查看论文全部施引文献