An unsupervised automatic organization method for Professor Shirakawa’s hand-notated documents of oracle bone inscriptions
-
Published:2024-03-05
Issue:
Volume:
Page:
-
ISSN:1433-2833
-
Container-title:International Journal on Document Analysis and Recognition (IJDAR)
-
language:en
-
Short-container-title:IJDAR
Author:
Yue Xuebin,Wang Ziming,Ishibashi Ryuto,Kaneko Hayata,Meng Lin
Abstract
AbstractAs one of the most influential Chinese cultural researchers in the second half of the twentieth-century, Professor Shirakawa is active in the research field of ancient Chinese characters. He has left behind many valuable research documents, especially his hand-notated oracle bone inscriptions (OBIs) documents. OBIs are one of the world’s oldest characters and were used in the Shang Dynasty about 3600 years ago for divination and recording events. The organization of OBIs is not only helpful in better understanding Prof. Shirakawa’s research and further study of OBIs in general and their importance in ancient Chinese history. This paper proposes an unsupervised automatic organization method to organize Prof. Shirakawa’s OBIs and construct a handwritten OBIs data set for neural network learning. First, a suite of noise reduction is proposed to remove strangely shaped noise to reduce the data loss of OBIs. Secondly, a novel segmentation method based on the supervised classification of OBIs regions is proposed to reduce adverse effects between characters for more accurate OBIs segmentation. Thirdly, a unique unsupervised clustering method is proposed to classify the segmented characters. Finally, all the same characters in the hand-notated OBIs documents are organized together. The evaluation results show that noise reduction has been proposed to remove noises with an accuracy of 97.85%, which contains number information and closed-loop-like edges in the dataset. In addition, the accuracy of supervised classification of OBIs regions based on our model achieves 85.50%, which is higher than eight state-of-the-art deep learning models, and a particular preprocessing method we proposed improves the classification accuracy by nearly 11.50%. The accuracy of OBIs clustering based on supervised classification achieves 74.91%. These results demonstrate the effectiveness of our proposed unsupervised automatic organization of Prof. Shirakawa’s hand-notated OBIs documents. The code and datasets are available at http://www.ihpc.se.ritsumei.ac.jp/obidataset.html.
Funder
Ritsumeikan University
Publisher
Springer Science and Business Media LLC
Reference57 articles.
1. Guo, J., Wang, C., Roman-Rangel, E., Chao, H., Rui, Y.: Building hierarchical representations for oracle character and sketch recognition. IEEE Trans. Image Process. 25(1), 104–118 (2016). https://doi.org/10.1109/TIP.2015.2500019 2. Han, W., Ren, X., Lin, H., Fu, Y., Xue, X.: Self-supervised learning of Orc-Bert augmentator for recognizing few-shot oracle characters. In: Proceedings of the Asian Conference on Computer Vision (ACCV) (2020) 3. Fujikawa, Y., Li, H., Yue, X., Aravinda, C.V., Prabhu, G.A., Meng, L.: Recognition of oracle bone inscriptions by using two deep learning models. In: CoRR (2021). arXiv:2105.00777 4. Meng, L., Kamitoku, N., Yamazaki, K.: Recognition of oracle bone inscriptions using deep learning based on data augmentation. In: 2018 Metrology for Archaeology and Cultural Heritage (MetroArchaeo), pp. 33–38 (2018). https://doi.org/10.1109/MetroArchaeo43810.2018.9089769 5. Yue, X., Lyu, B., Li, H., Fujikawa, Y., Meng, L.: Deep learning and image processing combined organization of Shirakawa’s hand-notated documents on OBI research. In: 2021 IEEE International Conference on Networking, Sensing and Control (ICNSC), vol. 1, pp. 1–6 (2021). https://doi.org/10.1109/ICNSC52481.2021.9702164
|
|