Pith. sign in

REVIEW 1 cited by

Data Lakes: A Survey of Functions and Systems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.09592 v2 pith:LQ2CFEEM submitted 2021-06-17 cs.DB

classification cs.DB
keywords datalakessurveyfunctionsresearchsystemsapproacheschallenges
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Data lakes are becoming increasingly prevalent for big data management and data analytics. In contrast to traditional 'schema-on-write' approaches such as data warehouses, data lakes are repositories storing raw data in its original formats and providing a common access interface. Despite the strong interest raised from both academia and industry, there is a large body of ambiguity regarding the definition, functions and available technologies for data lakes. A complete, coherent picture of data lake challenges and solutions is still missing. This survey reviews the development, architectures, and systems of data lakes. We provide a comprehensive overview of research questions for designing and building data lakes. We classify the existing approaches and systems based on their provided functions for data lakes, which makes this survey a useful technical reference for designing, implementing and deploying data lakes. We hope that the thorough comparison of existing solutions and the discussion of open research challenges in this survey will motivate the future development of data lake research and practice.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Enhancing OLAP Resilience at LinkedIn

    cs.DB 2026-03 accept novelty 6.0 of 10

    Production resiliency framework for Apache Pinot: Query Workload Isolation, impact-free zone-aware rebalancing, and adaptive server selection that sustain 99.9% availability at 10 PB / 250k QPS.

Pith tools