Exploiting Textual Information for Fake News Detection-Reference-Cited by-同舟云学术

Exploiting Textual Information for Fake News Detection

Published:2022-11-04 Issue:12 Volume:32 Page:
ISSN:0129-0657
Container-title:International Journal of Neural Systems
language:en
Short-container-title:Int. J. Neur. Syst.

Author:

Kasseropoulos Dimitrios Panagiotis¹,Koukaras Paraskevas¹,Tjortjis Christos¹

Affiliation:

1. The Data Mining and Analytics Research Group, School of Science and Technology International, Hellenic University, 14th km Thessaloniki — N. Moudania, 57001 Thermi, Thessaloniki, Greece

Abstract

“Fake news” refers to the deliberate dissemination of news with the purpose to deceive and mislead the public. This paper assesses the accuracy of several Machine Learning (ML) algorithms, using a style-based technique that relies on textual information extracted from news, such as part of speech counts. To expand the already proposed styled-based techniques, a new method of enhancing a linguistic feature set is proposed. It combines Named Entity Recognition (NER) with the Frequent Pattern (FP) Growth association rule mining algorithm, aiming to provide better insight into the papers’ sentence level structure. Recursive feature elimination was used to identify a subset of the highest performing linguistic characteristics, which turned out to align with the literature. Using pre-trained word embeddings, document embeddings and weighted document embeddings were constructed using each word’s TF-IDF value as the weight factor. The document embeddings were mixed with the linguistic features providing a variety of training/test feature sets. For each model, the best performing feature set was identified and fine-tuned regarding its hyper parameters to improve accuracy. ML algorithms’ results were compared with two Neural Networks: Convolutional Neural Network (CNN) and Long-Short-Term Memory (LSTM). The results indicate that CNN outperformed all other methods in terms of accuracy, when companied with pre-trained word embeddings, yet SVM performs almost the same with a wider variety of input feature sets. Although style-based technique scores lower accuracy, it provides explainable results about the author’s writing style decisions. Our work points out how new technologies and combinations of existing techniques can enhance the style-based approach capturing more information.

Publisher

World Scientific Pub Co Pte Ltd

Subject

Computer Networks and Communications,General Medicine

Link

https://www.worldscientific.com/doi/pdf/10.1142/S0129065722500587

Reference46 articles.

1. Social Media Types: introducing a data driven taxonomy

2. Social media prediction: a literature review

3. Using Twitter to Predict Chart Position for Songs

Cited by 3 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Fake News Detection via Sentiment Neutralization;2023 IEEE International Conference on Big Data (BigData);2023-12-15

2. A Modified Long Short-Term Memory Cell;International Journal of Neural Systems;2023-06-09

3. Fake News Detection Utilizing Textual Cues;IFIP Advances in Information and Communication Technology;2023