Affiliation:
1. China University of Geosciences
Abstract
The development of Internet and digital library has triggered a lot of text categorization methods. How to find desired information accurately and timely is becoming more and more important and automatic text categorization can help us achieve this goal. In general, text classifier is implemented by using some traditional classification methods such as Naive-Bayes (NB). ARC-BC (Associative Rule-based Classifier by Category) can be used for text categorization by dividing text documents into subsets in which all documents belong to the same category and generate associative classification rules for each subset. This classifier differs from previous methods in that it consists of discovered association rules between words and categories extracted from the training set. In order to train and test this classifier, we constructed training data and testing data respectively by selecting documents from Yahoo. The experimental result shows that the performance of ARC-BC based text categorization is very pretty efficient and effective and it is comparable to Naïve Bayesian algorithm based text categorization.
Publisher
Trans Tech Publications, Ltd.