Correlation based feature selection with clustering for high dimensional data

2018-01-31Journal of Electrical Systems and Information Technology 2018Unverified0· sign in to hype

Smita Chormunge, Sudarson Jena

Unverified — Be the first to reproduce this paper.

Abstract

Feature selection is an essential technique to reduce the dimensionality problem in data mining task. Traditional feature selection algorithms are fail to scale on large space. This paper proposes a new method to solve dimensionality problem where clustering is integrating with correlation measure to produce good feature subset. First Irrelevant features are eliminated by using k-means clustering method and then non-redundant features are selected by correlation measure from each cluster. The proposed method is evaluate on Microarray and Text datasets and the results are compared with other renowned feature selection methods using Naïve Bayes classifier. To verify the accuracy of the proposed method with different number of relevant features, percentagewise criteria is used. The experimental results reveal the efficiency and accuracy of the proposed method.

Tasks

Clustering feature selection Vocal Bursts Intensity Prediction

Correlation based feature selection with clustering for high dimensional data

Abstract

Tasks

Reproductions