Data Distribution Aware Classification Algorithm based on K-Means

Abstract

Giving data driven decisions based on precise data analysis is widely required by different businesses. For this purpose many different data mining strategies exist. Nevertheless, existing strategies need attention by researchers so that they can be adapted to the modern data analysis needs. One of the popular algorithms is K-Means. This paper proposes a novel improvement to the classical K-Means classification algorithm. It is known that data characteristics like data distribution, high-dimensionality, the size, the sparseness of the data, etc. have a great impact on the success of the K-Means clustering, which directly affects the accuracy of classification. In this study, the K-Means algorithm was modified to remedy the algorithm’s classification accuracy degradation, which is observed when the data distribution is not suitable to be clustered by data centroids, where each centroid is represented by a single mean. Specifically, this paper proposes to intelligently include the effect of variance based on the detected data distribution nature of the data. To see the performance improvement of the proposed method, several experiments were carried out using different real datasets. The presented results, which are achieved after extensive experiments, prove that the proposed algorithm improves the classification accuracy of KMeans. The achieved performance was also compared against several recent classification studies which are based on different classification schemes.

Authors and Affiliations

Tamer Tulgar, Ali Haydar, Ibrahim Ersan

Keywords

Related Articles

A New Task Scheduling Algorithm using Firefly and Simulated Annealing Algorithms in Cloud Computing

Task scheduling is a challenging and important issue, which considering increases in data sizes and large volumes of data, has turned into an NP-hard problem. This has attracted the attention of many researchers througho...

Investigating on Mobile Ad-Hoc Network to Transfer FTP Application

Mobile Ad-hoc Network (MANET) is the collection of mobile nodes without requiring of any infrastructure. Mobile nodes in MANET are operating as a router and MANET network topology can change quickly. Due to nodes in the...

 Task Allocation Model for Rescue Disabled Persons in Disaster Area with Help of Volunteers

 In this paper, we present a task allocation model for search and rescue persons with disabilities in case of disaster. The multi agent-based simulation model is used to simulate the rescue process. Volunteers and d...

Measuring the Impact of the Blackboard System on Blended Learning Students

With the advantages of using learning management systems (LMS) such as Blackboard in the educational process, assessing the impact of such systems has become increasingly important. This study measures the impact of the...

Applying FireFly Algorithm to Solve the Problem of Balancing Curricula

The problem of assigning a balanced academic curriculum to academic periods of a curriculum, that is, the balancing curricula, represents a traditional challenge for every educational institution which look for a match a...

Download PDF file
  • EP ID EP261189
  • DOI 10.14569/IJACSA.2017.080946
  • Views 100
  • Downloads 0

How To Cite

Tamer Tulgar, Ali Haydar, Ibrahim Ersan (2017). Data Distribution Aware Classification Algorithm based on K-Means. International Journal of Advanced Computer Science & Applications, 8(9), 328-334. https://europub.co.uk/articles/-A-261189