Social Media Monitoring of the COVID-19 Pandemic and Influenza Epidemic With Adaptation for Informal Language in Arabic Twitter Data: Qualitative Study-Reference-Cited by-同舟云学术

Social Media Monitoring of the COVID-19 Pandemic and Influenza Epidemic With Adaptation for Informal Language in Arabic Twitter Data: Qualitative Study

Published:2021-09-17 Issue:9 Volume:9 Page:e27670
ISSN:2291-9694
Container-title:JMIR Medical Informatics
language:en
Short-container-title:JMIR Med Inform

Author:

Alsudias Lama^ORCID,Rayson Paul^ORCID

Abstract

Background Twitter is a real-time messaging platform widely used by people and organizations to share information on many topics. Systematic monitoring of social media posts (infodemiology or infoveillance) could be useful to detect misinformation outbreaks as well as to reduce reporting lag time and to provide an independent complementary source of data compared with traditional surveillance approaches. However, such an analysis is currently not possible in the Arabic-speaking world owing to a lack of basic building blocks for research and dialectal variation. Objective We collected around 4000 Arabic tweets related to COVID-19 and influenza. We cleaned and labeled the tweets relative to the Arabic Infectious Diseases Ontology, which includes nonstandard terminology, as well as 11 core concepts and 21 relations. The aim of this study was to analyze Arabic tweets to estimate their usefulness for health surveillance, understand the impact of the informal terms in the analysis, show the effect of deep learning methods in the classification process, and identify the locations where the infection is spreading. Methods We applied the following multilabel classification techniques: binary relevance, classifier chains, label power set, adapted algorithm (multilabel adapted k-nearest neighbors [MLKNN]), support vector machine with naive Bayes features (NBSVM), bidirectional encoder representations from transformers (BERT), and AraBERT (transformer-based model for Arabic language understanding) to identify tweets appearing to be from infected individuals. We also used named entity recognition to predict the place names mentioned in the tweets. Results We achieved an F1 score of up to 88% in the influenza case study and 94% in the COVID-19 one. Adapting for nonstandard terminology and informal language helped to improve accuracy by as much as 15%, with an average improvement of 8%. Deep learning methods achieved an F1 score of up to 94% during the classifying process. Our geolocation detection algorithm had an average accuracy of 54% for predicting the location of users according to tweet content. Conclusions This study identified two Arabic social media data sets for monitoring tweets related to influenza and COVID-19. It demonstrated the importance of including informal terms, which are regularly used by social media users, in the analysis. It also proved that BERT achieves good results when used with new terms in COVID-19 tweets. Finally, the tweet content may contain useful information to determine the location of disease spread.

Publisher

JMIR Publications Inc.

Subject

Health Information Management,Health Informatics

Reference50 articles.

1. Survey of Text-based Epidemic Intelligence

2. LambAPaulMDredzeMSeparating Fact from Fear: Tracking Flu Infections on TwitterProceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies20132013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language TechnologiesJune 2013Atlanta, GA, USA789795

3. Arabic-speaking migrants’ experiences of the use of interpreters in healthcare: a qualitative explorative study

4. World Health Organization2020-03-01https://www.who.int/

Cited by 13 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Transformers and large language models in healthcare: A review;Artificial Intelligence in Medicine;2024-08

2. Social networking service (SNS) and transformer-based models for event-based surveillance for early detection of heat stroke in Aichi Prefecture, Japan;2024-07-18

3. Mapping automatic social media information disorder. The role of bots and AI in spreading misleading information in society;PLOS ONE;2024-05-31

4. An Ensemble Classification of Mental Health in Malaysia related to the Covid-19 Pandemic using Social Media Sentiment Analysis;KSII Transactions on Internet and Information Systems;2024-02-29

5. Mapping the Landscape of Misinformation Detection: A Bibliometric Approach;Information;2024-01-19