Collecting and visualizing data lineage of Spark jobs

Schoenenwald, Alexander; Kern, Simon; Viehhauser, Josef; Schildgen, Johannes

Collecting and visualizing data lineage of Spark jobs

dc.contributor.author	Schoenenwald, Alexander
dc.contributor.author	Kern, Simon
dc.contributor.author	Viehhauser, Josef
dc.contributor.author	Schildgen, Johannes
dc.date.accessioned	2022-01-27T13:27:56Z
dc.date.available	2022-01-27T13:27:56Z
dc.date.issued	2021
dc.description.abstract	Metadata management constitutes a key prerequisite for enterprises as they engage in data analytics and governance. Today, however, the context of data is often only manually documented by subject matter experts, and lacks completeness and reliability due to the complex nature of data pipelines. Thus, collecting data lineage—describing the origin, structure, and dependencies of data—in an automated fashion increases quality of provided metadata and reduces manual effort, making it critical for the development and operation of data pipelines. In our practice report, we propose an end-to-end solution that digests lineage via (Py‑)Spark execution plans. We build upon the open-source component Spline , allowing us to reliably consume lineage metadata and identify interdependencies. We map the digested data into an expandable data model, enabling us to extract graph structures for both coarse- and fine-grained data lineage. Lastly, our solution visualizes the extracted data lineage via a modern web app, and integrates with BMW Group’s soon-to-be open-sourced Cloud Data Hub.	de
dc.identifier.doi	10.1007/s13222-021-00387-7
dc.identifier.pissn	1610-1995
dc.identifier.uri	http://dx.doi.org/10.1007/s13222-021-00387-7
dc.identifier.uri	https://dl.gi.de/handle/20.500.12116/38053
dc.publisher	Springer
dc.relation.ispartof	Datenbank-Spektrum: Vol. 21, No. 3
dc.relation.ispartofseries	Datenbank-Spektrum
dc.subject	Amazon Web Service
dc.subject	Data Engineering
dc.subject	Data Lake
dc.subject	Data Provenance
dc.subject	Metadata
dc.title	Collecting and visualizing data lineage of Spark jobs	de
dc.type	Text/Journal Article
gi.citation.endPage	189
gi.citation.startPage	179

Sammlungen

Datenbank Spektrum 21(3) - November 2021

Collecting and visualizing data lineage of Spark jobs

Dateien

Sammlungen