REVIEW 9 cited by
Data-centric Artificial Intelligence: A Survey
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Artificial Intelligence (AI) is making a profound impact in almost every domain. A vital enabler of its great success is the availability of abundant and high-quality data for building machine learning models. Recently, the role of data in AI has been significantly magnified, giving rise to the emerging concept of data-centric AI. The attention of researchers and practitioners has gradually shifted from advancing model design to enhancing the quality and quantity of the data. In this survey, we discuss the necessity of data-centric AI, followed by a holistic view of three general data-centric goals (training data development, inference data development, and data maintenance) and the representative methods. We also organize the existing literature from automation and collaboration perspectives, discuss the challenges, and tabulate the benchmarks for various tasks. We believe this is the first comprehensive survey that provides a global view of a spectrum of tasks across various stages of the data lifecycle. We hope it can help the readers efficiently grasp a broad picture of this field, and equip them with the techniques and further research ideas to systematically engineer data for building AI systems. A companion list of data-centric AI resources will be regularly updated on https://github.com/daochenzha/data-centric-AI
Forward citations
Cited by 9 Pith papers
-
Beyond Component Testing: Validating Agentic AI Systems
Agentic AI cannot be adequately validated by component tests alone; trajectory-in-context validation is required, and current practice is mature only for behavioral evaluation.
-
DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams
DataClaw0 introduces an agentic data-tailoring paradigm, a 9B model trained on a synthetically generated dataset, and a new benchmark, claiming improved downstream adaptation in video generation, VQA, and GUI navigati...
-
"Skill Issues'': Data-Centric Optimization of Lakehouse Agents
Data-centric optimization of skills for agents on a branching lakehouse improves accuracy by 31.9% on 25 tasks via state-verification evaluation.
-
Feature Shift Localization Network
FSL-Net localizes shifted features between two datasets using a network trained on 1,350 datasets, matching DataFix's F1 while being about 36x faster on average.
-
ProTAL: A Drag-and-Link Video Programming Framework for Temporal Action Localization
A drag-and-link interface lets users define rules for actions from body-part and object relations, generating frame labels to train temporal action localization models.
-
Laplace Sample Information: Data Informativeness Through a Bayesian Lens
LSI ranks training samples by informativeness using the KL divergence between Laplace-approximated posteriors with and without each sample, and the ordering transfers from a small probe to larger models.
-
DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines
DataFlow-Harness builds editable data pipelines as validated DAGs via an LLM agent, hitting 93.3% task pass rate with 72.5% lower cost than vanilla script generation.
-
MEGG: Replay via Maximally Extreme GGscore in Incremental Learning for Neural Recommendation Models
A gradient-alignment influence score (GGscore) that selects the highest- and lowest-scoring old interactions for replay improves incremental neural recommendation slightly over random replay, mainly at large replay ratios.
-
Importance of User Control in Data-Centric Steering for Healthcare Experts
Healthcare experts who manually adjusted training data improved a diabetes prediction model more than those using automated corrections, without losing trust or understanding.
Discussion (0). Sign in to comment.