CNN-Based Page Segmentation and Object Classification for Counting Population in Ottoman Archival Documentation-Reference-Cited by-同舟云学术

CNN-Based Page Segmentation and Object Classification for Counting Population in Ottoman Archival Documentation

Published:2020-05-14 Issue:5 Volume:6 Page:32
ISSN:2313-433X
Container-title:Journal of Imaging
language:en
Short-container-title:J. Imaging

Author:

Can Yekta Said^ORCID,Kabadayı M. Erdem^ORCID

Abstract

Historical document analysis systems gain importance with the increasing efforts in the digitalization of archives. Page segmentation and layout analysis are crucial steps for such systems. Errors in these steps will affect the outcome of handwritten text recognition and Optical Character Recognition (OCR) methods, which increase the importance of the page segmentation and layout analysis. Degradation of documents, digitization errors, and varying layout styles are the issues that complicate the segmentation of historical documents. The properties of Arabic scripts such as connected letters, ligatures, diacritics, and different writing styles make it even more challenging to process Arabic script historical documents. In this study, we developed an automatic system for counting registered individuals and assigning them to populated places by using a CNN-based architecture. To evaluate the performance of our system, we created a labeled dataset of registers obtained from the first wave of population registers of the Ottoman Empire held between the 1840s and 1860s. We achieved promising results for classifying different types of objects and counting the individuals and assigning them to populated places.

Funder

H2020 European Research Council

Publisher

MDPI AG

Subject

Electrical and Electronic Engineering,Computer Graphics and Computer-Aided Design,Computer Vision and Pattern Recognition,Radiology, Nuclear Medicine and imaging

Link

https://www.mdpi.com/2313-433X/6/5/32/pdf

Reference32 articles.

1. Deep Neural Networks for Document Processing of Music Score Images

Cited by 12 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. A nineteenth-century urban Ottoman population micro dataset: Data extraction and relational database curation from the 1840s pre-census Bursa population registers;Scientific Data;2024-06-03

2. Automatic damage identification of Sanskrit palm leaf manuscripts with SegFormer;Heritage Science;2024-01-02

3. A hybrid web analytic approach through click enabled vision based page segmentation in quest software for school students;Journal of Intelligent & Fuzzy Systems;2022-09-22

4. Augmentation-based Pseudo-Ground truth Generation for Deep Learning in Historical Document Segmentation for Greater Levels of Archival Description and Access;Journal on Computing and Cultural Heritage;2022-09-16

5. Classification of handwritten annotations in mixed-media documents;2022 19th Conference on Robots and Vision (CRV);2022-05