Decoding Multilingual Topic Dynamics and Trend Identification through ARIMA Time Series Analysis on Social Networks: A Novel Data Translation Framework Enhanced by LDA/HDP Models-Reference-Cited by-同舟云学术

Decoding Multilingual Topic Dynamics and Trend Identification through ARIMA Time Series Analysis on Social Networks: A Novel Data Translation Framework Enhanced by LDA/HDP Models

Published:2024-01 Issue:1 Volume:2024 Page:
ISSN:2090-0147
Container-title:Journal of Electrical and Computer Engineering
language:en
Short-container-title:Journal of Electrical and Computer Engineering

Author:

Jaballi Samawel^ORCID,Hazar Manar Joundy^ORCID,Zrigui Salah^ORCID,Mahjoubi Azer^ORCID,Nicolas Henri^ORCID,Zrigui Mounir^ORCID

Abstract

In this study, the authors present a novel methodology adept at decoding multilingual topic dynamics and identifying communication trends during crises. We focus on dialogues within Tunisian social networks during the coronavirus pandemic and other notable themes like sports and politics. We start by aggregating a varied multilingual corpus of comments relevant to these subjects. This dataset undergoes rigorous refinement during data preprocessing. We then introduce our No‐English‐to‐English Machine Translation approach to handle linguistic differences. Empirical tests of this method show high accuracy and F1 scores, highlighting its suitability for linguistically coherent tasks. Delving deeper, advanced modeling techniques, specifically LDA and HDP models, are employed to extract pertinent topics from the translated content. This leads to applying ARIMA time series analysis to decode evolving topic trends. Applying our method to a multilingual Tunisian dataset, we effectively identify key topics mirroring public sentiment. Such insights prove vital for organizations and governments striving to understand public perspectives during crises. Compared to standard approaches, our model outperforms, as confirmed by metrics like coherence score, U‐mass, and topic coherence. Additionally, an in‐depth assessment of the identified topics reveals notable thematic shifts in discussions, with the proposed trends’ identification indicating impressive accuracy, backed by RMSE‐based analysis.

Publisher

Wiley

Link

https://onlinelibrary.wiley.com/doi/pdf/10.1155/2024/6669491

Reference37 articles.

1. MusaI. H. XuK. andZamitI. Multilingual document concept topic modeling Proceedings of the 2022 European Conference on Natural Language Processing and Information Retrieval (ECNLPIR) July 2022 IEEE Hangzhou China 84–91.

2. A Survey of Cross-lingual Word Embedding Models

3. YangW. Boyd-GraberJ. andResnikP. A multilingual topic model for learning weighted topic links across corpora with low comparability Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing November 2019 Hong Kong China.

4. Multilingual topic modeling for tracking COVID-19 trends based on Facebook data analysis

5. Monolingual and multilingual topic analysis using LDA and BERT embeddings