Metadata for Efficient Management of Digital News Articles in Multilingual News Archives-Reference-Cited by-同舟云学术

Metadata for Efficient Management of Digital News Articles in Multilingual News Archives

Published:2023-10 Issue:4 Volume:13 Page:
ISSN:2158-2440
Container-title:SAGE Open
language:en
Short-container-title:SAGE Open

Author:

Khan Muzammil¹^ORCID,Alharbi Yasser²,Alferaidi Ali²,Alharbi Talal Saad²,Yadav Kusum²

Affiliation:

1. University of Swat, Pakistan

2. University of Hail, Saudi Arabia

Abstract

The digital news preservation and management of low-resource languages are challenging tasks, especially in vast collections. Unique identification of individual digital objects is possible with well-defined attributes to assure efficient management, such as access, retrieval, preservation, usability, and transformability. The metadata element set is required to maximize the available attributes related to the digital objects. To create a comprehensive metadata set that contains all the necessary attributes and data about the digital news objects. It is more challenging and complicated when the archive contains articles from low-resourced and morphologically complex languages like Urdu and Arabic, which is difficult for machines to understand. The study presents challenges in low-resource languages (LRL) and research challenges. This metadata will help to link news articles based on similarity with other news articles stored in the digital news stories archive (DNSA) and ensures accessibility. In this study, we introduced 38 metadata elements set for the digital news stories preservation (DNSP) framework, of which 16 are explicit and 12 are implicit metadata elements. The paper presents how the digital news stories archive (DNSA) is enhanced to a multilingual archive and discusses the digital news stories extractor, which addresses major issues in implementing low-resource languages and facilitates normalized format migration. The extraction results are presented in detail for high-resource languages, that is, English, and low-resource languages (HRL), that is, Urdu and Arabic. The LRL encountered a high error rate during preservation compared to HRL, 10%, and 03%, respectively. The metadata extraction results show that HRL sources support all metadata elements as compared to LRL. The LRL has good support for explicit meta elements and many implicit meta elements with low extraction percentages. The LRL needs a more detailed study for accurate news content extraction and archiving for future access.

Funder

the Scientific Research Deanship at the University of Ha’il – Saudi Arabia, through project number RG-21 090

Publisher

SAGE Publications

Subject

General Social Sciences,General Arts and Humanities

Link

http://journals.sagepub.com/doi/pdf/10.1177/21582440231201368

Reference38 articles.

1. Ancestor Hunt. (2022). Retrieved September 14, 2022, from http://www.theancestorhunt.com/blog/europe-free-online-historical-newspapers#V0exUE9SHqd (Created in 2002).

2. Machine Translation System Using Deep Learning for English to Urdu

3. The British Library. (2022). Retrieved September 14, 2022, from https://www.bl.uk/collection-guides/arabic-collections; https://archive.org/ (Created in 1973).

4. Center for Research Libraries. (2022). The international coalition on newspapers (ICON). Retrieved September 14, 2022, from, http://icon.crl.edu/digitization.php (Established in 1999).

Cited by 2 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Understanding the Research Challenges in Low-Resource Language and Linking Bilingual News Articles in Multilingual News Archive;Applied Sciences;2023-07-25

2. The Role of Transliterated Words in Linking Bilingual News Articles in an Archive;Applied Sciences;2023-03-31