Exploration of Data Mining Techniques in Record Deduplication

Journal Title: International Journal of Science and Research (IJSR) - Year 2013, Vol 2, Issue 11

Abstract

In today?s business world, the database plays a vital role in decision making. As the organization grows, the size of the database also gets increased. This enormous growth in the database size leads to a problem of dirty data. Dirty data is the replicated data in the database which causes some issues like performance degradation, increasing operational cost and the lack of quality. This can be removed by the process of record deduplication. The record deduplication refers to identifying the same entity with different representations. Further cleaning and removing of replica in the repository become a mandatory work. Thus this paper surveys some of the record deduplication approaches. Also it compares with three approaches to record deduplication such as genetic programming, Modified BAT algorithm, and firefly algorithm approach with its limitation and advantages on all the three got discussed.

Authors and Affiliations

Keywords

Related Articles

Ethnomedicinal Practices of Koltribes in Shahdol Division Madhya Pradesh India

"A survey of ethnomedicinal plants of shahdol division has been carried out with co-operation of Kol tribal villagers. During study 31 ethnomedicinal plants have been identified for the treatment of various disease. Ha...

Power Quality Improvement Using Hybrid Filters for the Integration of Hybrid Distributed Generations to the Grid

Power quality improvement is an important parameter in power system. Not only to meet the demands, but also the quality power is a major goal of power system. Distributed Generations such as Solar and Fuel cell integrati...

A Comparison of Transverse Section with Arc Shaped Turbulators as an Artificial Roughness on the Absorber Plate of a Solar Air Heater

Solar energy is most important renewable energy resource due to its quantitative abundance. The simplest and most efficient way to utilize solar energy is to convert it into thermal energy for heating applications. A sol...

Modeling and Analysis of Aircraft Landing Gear: Experimental Approach

The main objective of this paper to present prototype of aircraft landing gear using higher end CAD software to study the behavior of landing gear as per actual working condition and to perform structural analysis to st...

Prevalence of Amblyopia in Children Aged from 5-15 Years in Rural Population Kurnool Dist. Andhra Pradesh, India

study of school children in rural area of Kurnool district by using E chart, streak retinoscope, and subjective correction, the children not improving refered to Govt.Regional Eye Hospital for final diagnosis. Among 1...

Download PDF file
  • EP ID EP339024
  • DOI -
  • Views 71
  • Downloads 0

How To Cite

(2013). Exploration of Data Mining Techniques in Record Deduplication. International Journal of Science and Research (IJSR), 2(11), -. https://europub.co.uk/articles/-A-339024