Exploration of Data Mining Techniques in Record Deduplication

Journal Title: UNKNOWN - Year 2013, Vol 2, Issue 11

Abstract

In today?s business world, the database plays a vital role in decision making. As the organization grows, the size of the database also gets increased. This enormous growth in the database size leads to a problem of dirty data. Dirty data is the replicated data in the database which causes some issues like performance degradation, increasing operational cost and the lack of quality. This can be removed by the process of record deduplication. The record deduplication refers to identifying the same entity with different representations. Further cleaning and removing of replica in the repository become a mandatory work. Thus this paper surveys some of the record deduplication approaches. Also it compares with three approaches to record deduplication such as genetic programming, Modified BAT algorithm, and firefly algorithm approach with its limitation and advantages on all the three got discussed.

Authors and Affiliations

Keywords

Related Articles

Optimal Channel and Relay Assignment in OFDM-based System using AF and DF Relay Strategies

Optimal Channel and Relay Assignment in OFDM-based System using AF and DF Relay Strategies

Pliocene pollen and spores from Sajau Coal, Berau Basin, Northeast Kalimantan, Indonesia: Environmental and Climatic Implications

"New data on paleovegetation and paleoclimate during the Pliocene has been obtained from palynological analysis of the Pliocene age coals of Sajau Formation in the eastern part of the Berau basin, Northeast Kalimantan, I...

Management of Non Performing Assets-A Case Study

Abstract: A strong banking sector is important for an economy like our

Performance Characteristics of Rectangular Patch Antenna

Performance Characteristics of Rectangular Patch Antenna

Triple Combination Therapy – A 12 Month Case Series

The ultimate goal of periodontal therapy is complete regeneration of the periodontal attachment apparatus. The combination of various regenerative biologic agents has recently attracted the interest of researchers in the...

Download PDF file
  • EP ID EP339024
  • DOI -
  • Views 85
  • Downloads 0

How To Cite

(2013). Exploration of Data Mining Techniques in Record Deduplication. UNKNOWN, 2(11), -. https://europub.co.uk/articles/-A-339024