Classification of Scientific Documents in the Kazakh Language Using Deep Neural Networks and a Fusion of Images and Text-Reference-Cited by-同舟云学术

Classification of Scientific Documents in the Kazakh Language Using Deep Neural Networks and a Fusion of Images and Text

Published:2022-10-24 Issue:4 Volume:6 Page:123
ISSN:2504-2289
Container-title:Big Data and Cognitive Computing
language:en
Short-container-title:BDCC

Author:

Bogdanchikov Andrey,Ayazbayev Dauren,Varlamis Iraklis^ORCID

Abstract

The rapid development of natural language processing and deep learning techniques has boosted the performance of related algorithms in several linguistic and text mining tasks. Consequently, applications such as opinion mining, fake news detection or document classification that assign documents to predefined categories have significantly benefited from pre-trained language models, word or sentence embeddings, linguistic corpora, knowledge graphs and other resources that are in abundance for the more popular languages (e.g., English, Chinese, etc.). Less represented languages, such as the Kazakh language, balkan languages, etc., still lack the necessary linguistic resources and thus the performance of the respective methods is still low. In this work, we develop a model that classifies scientific papers written in the Kazakh language using both text and image information and demonstrate that this fusion of information can be beneficial for cases of languages that have limited resources for machine learning models’ training. With this fusion, we improve the classification accuracy by 4.4499% compared to the models that use only text or only image information. The successful use of the proposed method in scientific documents’ classification paves the way for more complex classification models and more application in other domains such as news classification, sentiment analysis, etc., in the Kazakh language.

Publisher

MDPI AG

Subject

Artificial Intelligence,Computer Science Applications,Information Systems,Management Information Systems

Link

https://www.mdpi.com/2504-2289/6/4/123/pdf

Reference30 articles.

1. THESUS: Organizing Web document collections based on link semantics;Halkidi;VLDB J.,2003

2. Improving information retrieval using document clusters and semantic synonym extraction;Bharathi;J. Theor. Appl. Inf. Technol.,2012

3. Multi-label classification: An overview;Tsoumakas;Int. J. Data Warehous. Min. (IJDWM),2007

4. Kowsari, K., Jafari Meimandi, K., Heidarysafa, M., Mendu, S., Barnes, L., and Brown, D. Text classification algorithms: A survey. Information, 2019. 10.

5. The impact of deep learning on document classification using semantically rich representations;Kastrati;Inf. Process. Manag.,2019

Cited by 4 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Harnessing AI and NLP Tools for Innovating Brand Name Generation and Evaluation: A Comprehensive Review;Multimodal Technologies and Interaction;2024-07-01

2. Transforming Knowledge Management System with AI Technology for Document Archives;Day 3 Thu, May 09, 2024;2024-05-07

3. A Comparative Analysis of LSTM and BERT Models for Named Entity Recognition in Kazakh Language: A Multi-classification Approach;Communications in Computer and Information Science;2024

4. DLBCNet: A Deep Learning Network for Classifying Blood Cells;Big Data and Cognitive Computing;2023-04-14