A Multistage Feature Selection Model for Document Classification Using Information Gain and Rough Set

Abstract

 Huge number of documents are increasing rapidly, therefore, to organize it in digitized form text categorization becomes an challenging issue. A major issue for text categorization is its large number of features. Most of the features are noisy, irrelevant and redundant, which may mislead the classifier. Hence, it is most important to reduce dimensionality of data to get smaller subset and provide the most gain in information. Feature selection techniques reduce the dimensionality of feature space. It also improves the overall accuracy and performance. Hence, to overcome the issues of text categorization feature selection is considered as an efficient technique . Therefore, we, proposed a multistage feature selection model to improve the overall accuracy and performance of classification. In the first stage document preprocessing part is performed. Secondly, each term within the documents are ranked according to their importance for classification using the information gain. Thirdly rough set technique is applied to the terms which are ranked importantly and feature reduction is carried out. Finally a document classification is performed on the core features using Naive Bayes and KNN classifier. Experiments are carried out on three UCI datasets, Reuters 21578, Classic 04 and Newsgroup 20. Results show the better accuracy and performance of the proposed model.

Authors and Affiliations

Mrs. Leena. Patil, Dr. Mohammed Atique

Keywords

Related Articles

 Locality of Chlorophyll-A Distribution in the Intensive Study Area of the Ariake Sea, Japan in Winter Seasons based on Remote Sensing Satellite Data

 Mechanism of chlorophyll-a appearance and its locality in the intensive study area of the Ariake Sea, Japan in winter seasons is clarified by using remote sensing satellite data. Through experiments with Terra and...

A Cumulative Multi-Niching Genetic Algorithm for Multimodal Function Optimization

This paper presents a cumulative multi-niching genetic algorithm (CMN GA), designed to expedite optimization problems that have computationally-expensive multimodal objective functions. By never discarding individuals fr...

 Improved Text Reading System for Digital Open Universities

 The New Generation of Digital Open Universities (DOUNG) is a recently proposed model using m-learning and cloud computing option and based on an integrated architecture built with open networks as GSM and Internet....

 Evaluating Sentiment Analysis Methods and Identifying Scope of Negation in Newspaper Articles

 Automatic detection of linguistic negation in free text is a demanding need for many text processing applications including Sentiment Analysis. Our system uses online news archives from two different resources name...

 Method for Reducing the Number of Wild Animal Monitors by Means of Kriging

 Method for reducing the number of wild animal monitors is proposed by means of Kriging. Through wild animal route of simulations with 128 by 128 cells, the required number of wild animal monitors is clarified. Then...

Download PDF file
  • EP ID EP88970
  • DOI 10.14569/IJARAI.2014.031103
  • Views 157
  • Downloads 0

How To Cite

Mrs. Leena. Patil, Dr. Mohammed Atique (2014).  A Multistage Feature Selection Model for Document Classification Using Information Gain and Rough Set. International Journal of Advanced Research in Artificial Intelligence(IJARAI), 3(11), 14-20. https://europub.co.uk/articles/-A-88970