Construction of a Character Dataset for Historical Uchen Tibetan Documents under Low-Resource Conditions-Reference-Cited by-同舟云学术

Construction of a Character Dataset for Historical Uchen Tibetan Documents under Low-Resource Conditions

Published:2022-11-27 Issue:23 Volume:11 Page:3919
ISSN:2079-9292
Container-title:Electronics
language:en
Short-container-title:Electronics

Author:

Zhang Ce^ORCID,Wang Weilan,Zhang Guowei

Abstract

The construction of a character dataset is an important part of the research on document analysis and recognition of historical Tibetan documents. The results of character segmentation research in the previous stage are presented by coloring the characters with different color values. On this basis, the characters are annotated, and the character images corresponding to the annotation are extracted to construct a character dataset. The construction of a character dataset is carried out as follows: (1) text annotation of segmented characters is performed; (2) the character image is extracted from the character block based on the real position information; (3) according to the class of annotated text, the extracted character images are classified to construct a preliminary character dataset; (4) data augmentation is used to solve the imbalance of classes and samples in the preliminary dataset; (5) research on character recognition based on the constructed dataset is performed. The experimental results show that under low-resource conditions, this paper solves the challenges in the construction of a historical Uchen Tibetan document character dataset and constructs a 610-class character dataset. This dataset lays the foundation for the character recognition of historical Tibetan documents and provides a reference for the construction of relevant document datasets.

Funder

National Natural Science Foundation of China

Science and Technology Research Program of Chongqing Education Commission

Research Program of Chongqing University of Education

Publisher

MDPI AG

Subject

Electrical and Electronic Engineering,Computer Networks and Communications,Hardware and Architecture,Signal Processing,Control and Systems Engineering

Link

https://www.mdpi.com/2079-9292/11/23/3919/pdf

Reference26 articles.

1. Automatic character recognition for Tibetan script;Kojima;J. Indian Buddh. Stud.,1991

2. Character recognition of wooden blocked Tibetan similar manuscripts by using Euclidean distance with deferential weight;Kojima;Ipsj Sig Notes,1996

3. Extraction of characteristic features in Tibetan wood-block editions;Kojima;J. Indian Buddh. Stud.,1994

4. Layout analysis for historical Tibetan documents based on convolutional denoising autoencoder;Zhang;J. Chin. Inform. Process.,2018

5. Text extraction method for historical Tibetan document images based on block projections;Duan;Optoelectron. Lett.,2017

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Survey on text analysis and recognition for multiethnic scripts;Journal of Image and Graphics;2024