Pith. sign in

REVIEW 1 cited by

Informal Data Transformation Considered Harmful

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2001.00338 v1 pith:FVY4MSG4 submitted 2020-01-02 cs.DB cs.AI

classification cs.DBcs.AI
keywords dataintegrityenterprisepositiontakeachievingalgorithmsapproach
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper we take the common position that AI systems are limited more by the integrity of the data they are learning from than the sophistication of their algorithms, and we take the uncommon position that the solution to achieving better data integrity in the enterprise is not to clean and validate data ex-post-facto whenever needed (the so-called data lake approach to data management, which can lead to data scientists spending 80% of their time cleaning data), but rather to formally and automatically guarantee that data integrity is preserved as it transformed (migrated, integrated, composed, queried, viewed, etc) throughout the enterprise, so that data and programs that depend on that data need not constantly be re-validated for every particular use.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Grain Theory: Type-Level Granularity Correctness in Data Pipelines

    cs.DB 2026-01 reject novelty 3.0 of 10

    A formal notion of 'grain' (a minimal identifying type) is defined and used to infer the grain of join results from input grains, claiming compile-time verification of pipeline correctness.

Pith tools