Automatic Text Categorization in Terms of Genre and Author-Reference-Cited by-同舟云学术

Automatic Text Categorization in Terms of Genre and Author

Published:2000-12 Issue:4 Volume:26 Page:471-495
ISSN:0891-2017
Container-title:Computational Linguistics
language:en
Short-container-title:Computational Linguistics

Author:

Stamatatos Efstathios¹,Fakotakis Nikos¹,Kokkinakis George¹

Affiliation:

1. University of Patras, Department of Electrical & Computer Engineering, 26500 Patras, Greece.

Abstract

The two main factors that characterize a text are its content and its style, and both can be used as a means of categorization. In this paper we present an approach to text categorization in terms of genre and author for Modern Greek. In contrast to previous stylometric approaches, we attempt to take full advantage of existing natural language processing (NLP) tools. To this end, we propose a set of style markers including analysis-level measures that represent the way in which the input text has been analyzed and capture useful stylistic information without additional cost. We present a set of small-scale but reasonable experiments in text genre detection, author identification, and author verification tasks and show that the proposed method performs better than the most popular distributional lexical measures, i.e., functions of vocabulary richness and frequencies of occurrence of the most frequent words. All the presented experiments are based on unrestricted text downloaded from the World Wide Web without any manual text preprocessing or text sampling. Various performance issues regarding the training set size and the significance of the proposed style markers are discussed. Our system can be used in any application that requires fast and easily adaptable text categorization in terms of stylistically homogeneous categories. Moreover, the procedure of defining analysis-level markers can be followed in order to extract useful stylistic information using existing text processing tools.

Publisher

MIT Press - Journals

Subject

Artificial Intelligence,Computer Science Applications,Linguistics and Language,Language and Linguistics

Link

https://www.mitpressjournals.org/doi/pdf/10.1162/089120100750105920

Reference20 articles.

1. Outside the cave of shadows: using syntactic annotation to enhance authorship attribution

2. Methodological Issues Regarding Corpus-based Analyses of Linguistic Variation

3. Representativeness in Corpus Design

Cited by 174 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Genre Classification of Books in Russian with Stylometric Features: A Case Study;Information;2024-06-07

2. Statistical analysis of the complete corpus of fiction in Russian and recognition of the author;Keldysh Institute Preprints;2024

3. Stylometry and forensic science: A literature review;Forensic Science International: Synergy;2024

4. Exploration of Text Classification Algorithms Based on Word Vector Techniques;2023 4th International Conference on Computer, Big Data and Artificial Intelligence (ICCBD+AI);2023-12-15

5. Automatic genre identification: a survey;Language Resources and Evaluation;2023-11-16