Construction of a Character Dataset for Historical Uchen Tibetan Documents under Low-Resource Conditions
-
Published:2022-11-27
Issue:23
Volume:11
Page:3919
-
ISSN:2079-9292
-
Container-title:Electronics
-
language:en
-
Short-container-title:Electronics
Author:
Zhang CeORCID,
Wang Weilan,
Zhang Guowei
Abstract
The construction of a character dataset is an important part of the research on document analysis and recognition of historical Tibetan documents. The results of character segmentation research in the previous stage are presented by coloring the characters with different color values. On this basis, the characters are annotated, and the character images corresponding to the annotation are extracted to construct a character dataset. The construction of a character dataset is carried out as follows: (1) text annotation of segmented characters is performed; (2) the character image is extracted from the character block based on the real position information; (3) according to the class of annotated text, the extracted character images are classified to construct a preliminary character dataset; (4) data augmentation is used to solve the imbalance of classes and samples in the preliminary dataset; (5) research on character recognition based on the constructed dataset is performed. The experimental results show that under low-resource conditions, this paper solves the challenges in the construction of a historical Uchen Tibetan document character dataset and constructs a 610-class character dataset. This dataset lays the foundation for the character recognition of historical Tibetan documents and provides a reference for the construction of relevant document datasets.
Funder
National Natural Science Foundation of China
Science and Technology Research Program of Chongqing Education Commission
Research Program of Chongqing University of Education
Subject
Electrical and Electronic Engineering,Computer Networks and Communications,Hardware and Architecture,Signal Processing,Control and Systems Engineering
Reference26 articles.
1. Automatic character recognition for Tibetan script;Kojima;J. Indian Buddh. Stud.,1991
2. Character recognition of wooden blocked Tibetan similar manuscripts by using Euclidean distance with deferential weight;Kojima;Ipsj Sig Notes,1996
3. Extraction of characteristic features in Tibetan wood-block editions;Kojima;J. Indian Buddh. Stud.,1994
4. Layout analysis for historical Tibetan documents based on convolutional denoising autoencoder;Zhang;J. Chin. Inform. Process.,2018
5. Text extraction method for historical Tibetan document images based on block projections;Duan;Optoelectron. Lett.,2017
Cited by
1 articles.
订阅此论文施引文献
订阅此论文施引文献,注册后可以免费订阅5篇论文的施引文献,订阅后可以查看论文全部施引文献