Logo des Repositoriums
 

Duplicate detection on GPUs

dc.contributor.authorForchhammer, Benedikt
dc.contributor.authorPapenbrock, Thorsten
dc.contributor.authorStening, Thomas
dc.contributor.authorViehmeier, Sven
dc.contributor.authorDraisbach, Uwe
dc.contributor.authorNaumann, Felix
dc.contributor.editorMarkl, Volker
dc.contributor.editorSaake, Gunter
dc.contributor.editorSattler, Kai-Uwe
dc.contributor.editorHackenbroich, Gregor
dc.contributor.editorMitschang, Bernhard
dc.contributor.editorHärder, Theo
dc.contributor.editorKöppen, Veit
dc.date.accessioned2018-10-24T09:56:17Z
dc.date.available2018-10-24T09:56:17Z
dc.date.issued2013
dc.description.abstractWith the ever increasing volume of data and the ability to integrate different data sources, data quality problems abound. Duplicate detection, as an integral part of data cleansing, is essential in modern information systems. We present a complete duplicate detection workflow that utilizes the capabilities of modern graphics processing units (GPUs) to increase the efficiency of finding duplicates in very large datasets. Our solution covers several well-known algorithms for pair selection, attribute-wise similarity comparison, record-wise similarity aggregation, and clustering. We redesigned these algorithms to run memory-efficiently and in parallel on the GPU. Our experiments demonstrate that the GPU-based workflow is able to outperform a CPU-based implementation on large, real-world datasets. For instance, the GPU-based algorithm deduplicates a dataset with 1.8m entities 10 times faster than a common CPU-based algorithm using comparably priced hardware.en
dc.identifier.isbn978-3-88579-608-4
dc.identifier.pissn1617-5468
dc.identifier.urihttps://dl.gi.de/handle/20.500.12116/17320
dc.language.isoen
dc.publisherGesellschaft für Informatik e.V.
dc.relation.ispartofDatenbanksysteme für Business, Technologie und Web (BTW) 2024
dc.relation.ispartofseriesLecture Notes in Informatics (LNI) - Proceedings, Volume P-214
dc.titleDuplicate detection on GPUsen
dc.typeText/Conference Paper
gi.citation.endPage184
gi.citation.publisherPlaceBonn
gi.citation.startPage165
gi.conference.date13.-15. März 2013
gi.conference.locationMagdeburg
gi.conference.sessiontitleRegular Research Papers

Dateien

Originalbündel
1 - 1 von 1
Lade...
Vorschaubild
Name:
165.pdf
Größe:
2.02 MB
Format:
Adobe Portable Document Format