Exploration of Data Mining Techniques in Record Deduplication

Journal Title: UNKNOWN - Year 2013, Vol 2, Issue 11

Abstract

In today?s business world, the database plays a vital role in decision making. As the organization grows, the size of the database also gets increased. This enormous growth in the database size leads to a problem of dirty data. Dirty data is the replicated data in the database which causes some issues like performance degradation, increasing operational cost and the lack of quality. This can be removed by the process of record deduplication. The record deduplication refers to identifying the same entity with different representations. Further cleaning and removing of replica in the repository become a mandatory work. Thus this paper surveys some of the record deduplication approaches. Also it compares with three approaches to record deduplication such as genetic programming, Modified BAT algorithm, and firefly algorithm approach with its limitation and advantages on all the three got discussed.

Authors and Affiliations

Keywords

Related Articles

Aspect Range Motivated Classifier Ensemble Reduction

: This paper proposed about classifier which ensemble in data mining and it constitute one of the directions in main research of machine learning. In general multiple classifiers allow better predictable performance. To...

Privacy-Preserved Search in mCL-PKE Based Secure Data Sharing over Public Clouds and Credential Trust Management through SMTP Communication

With the fast prominence of cloud computing architecture, public key cryptography exhibit certificate revocation problem and in the case of identity based encryption there exist a key escrow problem. Certificateless publ...

The Dual Activation of the Geometric Inverse Burr Distribution: Comparison Between Maximum and Minimum

In this paper, we introduced a double activation approach for the geometric inverse burr distribution. Statistical measures and their properties are derived. Explicit expression for their density, rth moment and entropy...

Prediction for Pulmonary Disease Based on Diagnostic Reciepes and Classification

In this research work we have developed a strategy in which the various parameters that influence the occurrence of pulmonary disease have been gathered from survey of doctors who specialize in diagnoses of pulmonary dis...

Design of Cream Separator Machine Using Reverse Engineering Techniques

In various industries where original design details of the product are not available, any modification or development in the product becomes a challenging task. In such cases reverse engineering can be used to develop fu...

Download PDF file
  • EP ID EP339024
  • DOI -
  • Views 88
  • Downloads 0

How To Cite

(2013). Exploration of Data Mining Techniques in Record Deduplication. UNKNOWN, 2(11), -. https://europub.co.uk/articles/-A-339024