Defining Semantically Close Words of Kazakh Language with Distributed System Apache Spark-Reference-Cited by-同舟云学术

Defining Semantically Close Words of Kazakh Language with Distributed System Apache Spark

Published:2023-09-27 Issue:4 Volume:7 Page:160
ISSN:2504-2289
Container-title:Big Data and Cognitive Computing
language:en
Short-container-title:BDCC

Author:

Ayazbayev Dauren¹,Bogdanchikov Andrey¹^ORCID,Orynbekova Kamila¹^ORCID,Varlamis Iraklis²^ORCID

Affiliation:

1. Department of Computer Science, Suleyman Demirel University, Kaskelen 040900, Kazakhstan

2. Department of Informatics and Telematics, Harokopio University of Athens, 17779 Athens, Greece

Abstract

This work focuses on determining semantically close words and using semantic similarity in general in order to improve performance in information retrieval tasks. The semantic similarity of words is an important task with many applications from information retrieval to spell checking or even document clustering and classification. Although, in languages with rich linguistic resources, the methods and tools for this task are well established, some languages do not have such tools. The first step in our experiment is to represent the words in a collection in a vector form and then define the semantic similarity of the terms using a vector similarity method. In order to tame the complexity of the task, which relies on the number of word (and, consequently, of the vector) pairs that have to be combined in order to define the semantically closest word pairs, A distributed method that runs on Apache Spark is designed to reduce the calculation time by running comparison tasks in parallel. Three alternative implementations are proposed and tested using a list of target words and seeking the most semantically similar words from a lexicon for each one of them. In a second step, we employ pre-trained multilingual sentence transformers to capture the content semantics at a sentence level and a vector-based semantic index to accelerate the searches. The code is written in MapReduce, and the experiments and results show that the proposed methods can provide an interesting solution for finding similar words or texts in the Kazakh language.

Publisher

MDPI AG

Subject

Artificial Intelligence,Computer Science Applications,Information Systems,Management Information Systems

Link

https://www.mdpi.com/2504-2289/7/4/160/pdf

Reference31 articles.

1. Means: A medical question-answering system combining NLP techniques and semantic Web technologies;Abacha;Inf. Process. Manag.,2015

2. Gong, C., He, D., Tan, X., Qin, T., Wang, L., and Liu, T.-Y. (2018). FRAGE: Frequency-Agnostic Word Representation. Adv. Neural Inf. Process. Syst., 1341–1352.

3. Chung, Y., and Glass, J. (2018). Speech2Vec: A Sequence-to-Sequence Framework for Learning Word Embeddings from Speech. arXiv.

4. Serek, A., Issabek, A., and Bogdanchikov, A. (2019, January 10–12). Distributed sentiment analysis of an agglutinative language via spark by applying machine learning methods. Proceedings of the 15th International Conference on Electronics, Computer and Computation (ICECCO), Abuja, Nigeria.

5. Bogdanchikov, A., Kariboz, D., and Meraliyev, M. (December, January 29). Face extraction and recognition from public images using hipi. Proceedings of the 14th International Conference on Electronics Computer and Computation (ICECCO), Kaskelen, Kazakhstan.

Cited by 2 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Application of Natural Language Processing and Genetic Algorithm to Fine-Tune Hyperparameters of Classifiers for Economic Activities Analysis;Big Data and Cognitive Computing;2024-06-13

2. Intent Identification by Semantically Analyzing the Search Query;Modelling;2024-02-22