Pith. sign in

REVIEW 1 cited by

What can Data-Centric AI Learn from Data and ML Engineering?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2112.06439 v1 pith:WDMIXIMZ submitted 2021-12-13 cs.LG cs.DB

classification cs.LGcs.DB
keywords dataapplicationsdata-centricengineeringinterestingmanyorganizationsrange
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Data-centric AI is a new and exciting research topic in the AI community, but many organizations already build and maintain various "data-centric" applications whose goal is to produce high quality data. These range from traditional business data processing applications (e.g., "how much should we charge each of our customers this month?") to production ML systems such as recommendation engines. The fields of data and ML engineering have arisen in recent years to manage these applications, and both include many interesting novel tools and processes. In this paper, we discuss several lessons from data and ML engineering that could be interesting to apply in data-centric AI, based on our experience building data and ML platforms that serve thousands of applications at a range of organizations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards Next Generation Data Engineering Pipelines

    cs.DB 2025-07 unverdicted novelty 5.0 of 10

    A vision paper defines three levels of next generation data engineering pipelines (optimized, self-aware, self-adapting) and proposes an architecture to realize them.

Pith tools