Novel Perspectives for the Management of Multilingual and Multialphabetic Heritages through Automatic Knowledge Extraction: The DigitalMaktaba Approach-Reference-Cited by-同舟云学术

Novel Perspectives for the Management of Multilingual and Multialphabetic Heritages through Automatic Knowledge Extraction: The DigitalMaktaba Approach

Published:2022-05-25 Issue:11 Volume:22 Page:3995
ISSN:1424-8220
Container-title:Sensors
language:en
Short-container-title:Sensors

Author:

Bergamaschi Sonia^ORCID,De Nardis Stefania,Martoglia Riccardo^ORCID,Ruozzi Federico,Sala Luca,Vanzini Matteo^ORCID,Vigliermo Riccardo Amerigo

Abstract

The linguistic and social impact of multiculturalism can no longer be neglected in any sector, creating the urgent need of creating systems and procedures for managing and sharing cultural heritages in both supranational and multi-literate contexts. In order to achieve this goal, text sensing appears to be one of the most crucial research areas. The long-term objective of the DigitalMaktaba project, born from interdisciplinary collaboration between computer scientists, historians, librarians, engineers and linguists, is to establish procedures for the creation, management and cataloguing of archival heritage in non-Latin alphabets. In this paper, we discuss the currently ongoing design of an innovative workflow and tool in the area of text sensing, for the automatic extraction of knowledge and cataloguing of documents written in non-Latin languages (Arabic, Persian and Azerbaijani). The current prototype leverages different OCR, text processing and information extraction techniques in order to provide both a highly accurate extracted text and rich metadata content (including automatically identified cataloguing metadata), overcoming typical limitations of current state of the art approaches. The initial tests provide promising results. The paper includes a discussion of future steps (e.g., AI-based techniques further leveraging the extracted data/metadata and making the system learn from user feedback) and of the many foreseen advantages of this research, both from a technical and a broader cultural-preservation and sharing point of view.

Publisher

MDPI AG

Subject

Electrical and Electronic Engineering,Biochemistry,Instrumentation,Atomic and Molecular Physics, and Optics,Analytical Chemistry

Link

https://www.mdpi.com/1424-8220/22/11/3995/pdf

Reference37 articles.

1. Pearson Correlation-Based Feature Selection for Document Classification Using Balanced Training

2. Document-Image Related Visual Sensors and Machine Learning Techniques

3. Digitizing the Textual Heritage of the Premodern Islamicate World: Principles and Plans

4. Kitab Project https://kitab-project.org/about/

5. https://persdigumd.github.io/PDL/

Cited by 5 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Leveraging Automated Methods and Advanced Neural Networks for Accurate Information Extraction;2023 IEEE International Conference on Paradigm Shift in Information Technologies with Innovative Applications in Global Scenario (ICPSITIAGS);2023-12-28

2. Identifying and Resolving Conflicts Using Local Wisdom: A Qualitative Study;Journal of Intercultural Communication;2023-12-10

3. Entropy-Aware Time-Varying Graph Neural Networks with Generalized Temporal Hawkes Process: Dynamic Link Prediction in the Presence of Node Addition and Deletion;Machine Learning and Knowledge Extraction;2023-10-04

4. A Survey of OCR in Arabic Language: Applications, Techniques, and Challenges;Applied Sciences;2023-04-04

5. Sensors and Communications for the Social Good;Sensors;2023-02-22