TextNetTopics: Text Classification Based Word Grouping as Topics and Topics’ Scoring-Reference-Cited by-同舟云学术

TextNetTopics: Text Classification Based Word Grouping as Topics and Topics’ Scoring

Published:2022-06-20 Issue: Volume:13 Page:
ISSN:1664-8021
Container-title:Frontiers in Genetics
language:
Short-container-title:Front. Genet.

Author:

Yousef Malik,Voskergian Daniel

Abstract

Medical document classification is one of the active research problems and the most challenging within the text classification domain. Medical datasets often contain massive feature sets where many features are considered irrelevant, redundant, and add noise, thus, reducing the classification performance. Therefore, to obtain a better accuracy of a classification model, it is crucial to choose a set of features (terms) that best discriminate between the classes of medical documents. This study proposes TextNetTopics, a novel approach that applies feature selection by considering Bag-of-topics (BOT) rather than the traditional approach, Bag-of-words (BOW). Thus our approach performs topic selections rather than words selection. TextNetTopics is based on the generic approach entitled G-S-M (Grouping, Scoring, and Modeling), developed by Yousef and his colleagues and used mainly in biological data. The proposed approach suggests scoring topics to select the top topics for training the classifier. This study applied TextNetTopics to textual data to respond to the CAMDA challenge. TextNetTopics outperforms various feature selection approaches while highly performing when applying the model to the validation data provided by the CAMDA. Additionally, we have applied our algorithm to different textual datasets.

Publisher

Frontiers Media SA

Subject

Genetics (clinical),Genetics,Molecular Medicine

Reference38 articles.

1. An Ontology-Based Two-Stage Approach to Medical Text Classification with Feature Selection by Particle Swarm Optimisation;Abdollahi,2019

2. Comparative Study of Feature Selection Methods for Medical Full Text Classification;Adriano Gonçalves,2019

3. Exploring the Impact of Short-Text Complexity and Structure on its Quality in Social Media;Al Qundus;Jeim,2020

4. A Survey of Topic Modeling in Text Mining;Alghamdi;Int. J. Adv. Comput. Sci. ApplIJACSA,2015

5. KNIME - the Konstanz Information Miner: Version 2.0 and beyond;Berthold;ACM SIGKDD Explor. Newsl.,2009

Cited by 18 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. X-News dataset for online news categorization;International Journal of Intelligent Computing and Cybernetics;2024-08-13

2. Classification of Breast Cancer Molecular Subtypes with Grouping-Scoring-Modeling Approach that Incorporates Disease-Disease Association Information;2024 32nd Signal Processing and Communications Applications Conference (SIU);2024-05-15

3. G-S-M: A Comprehensive Framework for Integrative Feature Selection in Omics Data Analysis and Beyond;2024-04-01

4. Transforming Education Policy: Evaluating UAQTE Program Implementation Through LDA, BoW and TF-IDF Techniques;2024 26th International Conference on Advanced Communications Technology (ICACT);2024-02-04

5. Rule-Based Text Classification of Dental Diagnosis;Studies in Health Technology and Informatics;2024-01-25