Data Exploration and Classification of News Article Reliability: Deep Learning Study-Reference-Cited by-同舟云学术

Data Exploration and Classification of News Article Reliability: Deep Learning Study

Published:2022-09-22 Issue:2 Volume:2 Page:e38839
ISSN:2564-1891
Container-title:JMIR Infodemiology
language:en
Short-container-title:JMIR Infodemiology

Author:

Zhan Kevin^ORCID,Li Yutong^ORCID,Osmani Rafay^ORCID,Wang Xiaoyu^ORCID,Cao Bo^ORCID

Abstract

BackgroundDuring the ongoing COVID-19 pandemic, we are being exposed to large amounts of information each day. This “infodemic” is defined by the World Health Organization as the mass spread of misleading or false information during a pandemic. This spread of misinformation during the infodemic ultimately leads to misunderstandings of public health orders or direct opposition against public policies. Although there have been efforts to combat misinformation spread, current manual fact-checking methods are insufficient to combat the infodemic.ObjectiveWe propose the use of natural language processing (NLP) and machine learning (ML) techniques to build a model that can be used to identify unreliable news articles online.MethodsFirst, we preprocessed the ReCOVery data set to obtain 2029 English news articles tagged with COVID-19 keywords from January to May 2020, which are labeled as reliable or unreliable. Data exploration was conducted to determine major differences between reliable and unreliable articles. We built an ensemble deep learning model using the body text, as well as features, such as sentiment, Empath-derived lexical categories, and readability, to classify the reliability.ResultsWe found that reliable news articles have a higher proportion of neutral sentiment, while unreliable articles have a higher proportion of negative sentiment. Additionally, our analysis demonstrated that reliable articles are easier to read than unreliable articles, in addition to having different lexical categories and keywords. Our new model was evaluated to achieve the following performance metrics: 0.906 area under the curve (AUC), 0.835 specificity, and 0.945 sensitivity. These values are above the baseline performance of the original ReCOVery model.ConclusionsThis paper identified novel differences between reliable and unreliable news articles; moreover, the model was trained using state-of-the-art deep learning techniques. We aim to be able to use our findings to help researchers and the public audience more easily identify false information and unreliable media in their everyday lives.

Publisher

JMIR Publications Inc.

Reference75 articles.

1. How to Fight an Infodemic: The Four Pillars of Infodemic Management

2. World Health OrganizationInfodemic20222022-06-15https://www.who.int/health-topics/infodemic

3. COVID-19 misinformation: Accuracy of articles about coronavirus prevention mostly shared on social media

4. The current state of fake news: challenges and opportunities

5. Where We Go From Here: Health Misinformation on Social Media

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Muzzling Misinformation: Drawing from Other Disciplines and Engaging Health and Science Journalists as Research Collaborators;Palgrave Handbook of Science and Health Journalism;2024