REVIEW 4 cited by
Building a serverless Data Lakehouse from spare parts
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The recently proposed Data Lakehouse architecture is built on open file formats, performance, and first-class support for data transformation, BI and data science: while the vision stresses the importance of lowering the barrier for data work, existing implementations often struggle to live up to user expectations. At Bauplan, we decided to build a new serverless platform to fulfill the Lakehouse vision. Since building from scratch is a challenge unfit for a startup, we started by re-using (sometimes unconventionally) existing projects, and then investing in improving the areas that would give us the highest marginal gains for the developer experience. In this work, we review user experience, high-level architecture and tooling decisions, and conclude by sharing plans for future development.
Forward citations
Cited by 4 Pith papers
-
Not Your Usual Type(s): Data contracts as types across languages and engines
Treating data contracts as type annotations enforced at three pipeline stages lets multi-language lakehouse DAGs fail fast on schema mismatches.
-
GitLake: Git-for-data for the agentic lakehouse
GitLake lifts Iceberg snapshots into lakehouse-wide commits, branches, and merges so agents develop in isolation and multi-table pipelines publish atomically via temporary-branch merges.
-
Eudoxia: a FaaS scheduling simulator for the composable lakehouse
Eudoxia is a deterministic, open-source simulator for evaluating FaaS scheduling algorithms in composable lakehouses, with a TPC-H-based runtime validation on the Bauplan cloud platform.
-
FaaS and Furious: abstractions and differential caching for efficient data pre-processing
A columnar differential cache for lakehouse pipelines reuses overlapping scan fragments and reduces S3 bytes read by up to 30% in preliminary benchmarks.
Discussion (0). Continue with ORCID to comment.