REVIEW 7 cited by
Can Foundation Models Wrangle Your Data?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Can Foundation Models Wrangle Your Data?
read the original abstract
Foundation Models (FMs) are models trained on large corpora of data that, at very large scale, can generalize to new tasks without any task-specific finetuning. As these models continue to grow in size, innovations continue to push the boundaries of what these models can do on language and image tasks. This paper aims to understand an underexplored area of FMs: classical data tasks like cleaning and integration. As a proof-of-concept, we cast five data cleaning and integration tasks as prompting tasks and evaluate the performance of FMs on these tasks. We find that large FMs generalize and achieve SoTA performance on data cleaning and integration tasks, even though they are not trained for these data tasks. We identify specific research challenges and opportunities that these models present, including challenges with private and domain specific data, and opportunities to make data management systems more accessible to non-experts. We make our code and experiments publicly available at: https://github.com/HazyResearch/fm_data_tasks.
Forward citations
Cited by 7 Pith papers
-
LDI: Localized Data Imputation for Text-Rich Tables
LDI introduces localized LLM-based imputation for text-rich tables by selecting compact relevant subsets of attributes and tuples per missing value, reporting up to 8% accuracy gains over prior methods.
-
Managing Map Cardinality in Automatic Disease Classification Mapping: Balancing Precision, Recall and Coverage
A blocking-plus-LLM-matching method delivers higher precision and broader coverage than threshold or top-K baselines while maintaining comparable recall on ICD version mapping tasks.
-
Adaptive Graph Refinement and Label Propagation with LLMs for Cost-Effective Entity Resolution
Alper unifies entity resolution matching and clustering into an iterative graph refinement and probabilistic label propagation process that adaptively selects LLM queries via a budgeted greedy optimization to outperfo...
-
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference
DUAL-BLADE uses a dual-path KV-cache framework with NVMe-direct access to reduce prefill and decode latency by up to 33% and 42% while improving SSD utilization 2.2x under tight memory budgets.
-
TabClean: Reusable LLM-Synthesized Programs for Tabular Data Cleaning
TabClean synthesizes reusable guarded Python cleaning programs from LLM reasoning on a small development set to achieve high precision and lower recurring costs on tabular data.
-
ProfiliTable: Profiling-Driven Tabular Data Processing via Agentic Workflows
ProfiliTable is a multi-agent system with profiler, generator, and evaluator components that outperforms baselines on 18 tabular task types via dynamic profiling and closed-loop refinement.
-
ProfiliTable: Profiling-Driven Tabular Data Processing via Agentic Workflows
ProfiliTable is a profiling-driven multi-agent system that builds semantic context through exploration and closed-loop refinement to produce more reliable tabular data transformations than prior LLM approaches.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.