Dataset Discovery in Data Lakes

Alex Bogatu; Alvaro A.A. Fernandes; Norman W. Paton; Nikolaos; Konstantinou

arXiv:2011.10427·cs.DB·November 23, 2020

Dataset Discovery in Data Lakes

Alex Bogatu, Alvaro A.A. Fernandes, Norman W. Paton, Nikolaos, Konstantinou

PDF

2 Repos

TL;DR

This paper presents a novel hash-based indexing method for efficiently discovering relevant datasets in large data lakes by measuring feature similarity, significantly improving precision, recall, and discovery times.

Contribution

It introduces a new approach that uses feature-based hash indexes to identify related datasets in data lakes, enhancing discovery efficiency and accuracy.

Findings

01

Significant improvements in precision and recall.

02

Enhanced target coverage in dataset discovery.

03

Faster indexing and discovery times.

Abstract

Data analytics stands to benefit from the increasing availability of datasets that are held without their conceptual relationships being explicitly known. When collected, these datasets form a data lake from which, by processes like data wrangling, specific target datasets can be constructed that enable value-adding analytics. Given the potential vastness of such data lakes, the issue arises of how to pull out of the lake those datasets that might contribute to wrangling out a given target. We refer to this as the problem of dataset discovery in data lakes and this paper contributes an effective and efficient solution to it. Our approach uses features of the values in a dataset to construct hash-based indexes that map those features into a uniform distance space. This makes it possible to define similarity distances between features and to take those distances as measurements of…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.