An ensemble scheme based on language function analysis and feature engineering for text genre classification-Reference-Cited by-同舟云学术

An ensemble scheme based on language function analysis and feature engineering for text genre classification

Published:2016-12-01 Issue:1 Volume:44 Page:28-47
ISSN:0165-5515
Container-title:Journal of Information Science
language:en
Short-container-title:Journal of Information Science

Author:

Onan Aytuğ¹

Affiliation:

1. Department of Computer Engineering, Celal Bayar University, Turkey

Abstract

Text genre classification is the process of identifying functional characteristics of text documents. The immense quantity of text documents available on the web can be properly filtered, organised and retrieved with the use of text genre classification, which may have potential use on several other tasks of natural language processing and information retrieval. Genre may refer to several aspects of text documents, such as function and purpose. The language function analysis (LFA) concentrates on single aspect of genres and it aims to classify text documents into three abstract classes, such as expressive, appellative and informative. Text genre classification is typically performed by supervised machine learning algorithms. The extraction of an efficient feature set to represent text documents is an essential task for building a robust classification scheme with high predictive performance. In addition, ensemble learning, which combines the outputs of individual classifiers to obtain a robust classification scheme, is a promising research field in machine learning research. In this regard, this article presents an extensive comparative analysis of different feature engineering schemes (such as features used in authorship attribution, linguistic features, character n-grams, part of speech n-grams and the frequency of the most discriminative words) and five different base learners (Naïve Bayes, support vector machines, logistic regression, k-nearest neighbour and Random Forest) in conjunction with ensemble learning methods (such as Boosting, Bagging and Random Subspace). Based on the empirical analysis, an ensemble classification scheme is presented, which integrates Random Subspace ensemble of Random Forest with four types of features (features used in authorship attribution, character n-grams, part of speech n-grams and the frequency of the most discriminative words). For LFA corpus, the highest average predictive performance obtained by the proposed scheme is 94.43%.

Publisher

SAGE Publications

Subject

Library and Information Sciences,Information Systems

Link

http://journals.sagepub.com/doi/pdf/10.1177/0165551516677911

Reference63 articles.

1. Han J, Kamber M. Data mining: concepts and techniques. 2nd ed.San Francisco, CA: Morgan Kaufmann Publishers, 2006, p. 800.

2. A Survey of Text Clustering Algorithms

3. Ensemble of keyword extraction methods and classifiers in text classification

4. Automatic Text Categorization in Terms of Genre and Author

Cited by 250 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Machine learning methods as auxiliary tool for effective mathematics teaching;Computer Applications in Engineering Education;2024-08-12

2. Variable-Period Estimation of Process Industry Indicators Using Working Condition Semantic Representation and Mechanism-Guided Network Groups;IEEE Transactions on Industrial Informatics;2024-08

3. Contextual classification of clinical records with bidirectional long short‐term memory (Bi‐LSTM) and bidirectional encoder representations from transformers (BERT) model;Computational Intelligence;2024-08

4. Arabic text classification based on analogical proportions;Expert Systems;2024-06-17

5. A systematic mapping to investigate the application of machine learning techniques in requirement engineering activities;CAAI Transactions on Intelligence Technology;2024-06-10