REVIEW 1 cited by
Informal Data Transformation Considered Harmful
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this paper we take the common position that AI systems are limited more by the integrity of the data they are learning from than the sophistication of their algorithms, and we take the uncommon position that the solution to achieving better data integrity in the enterprise is not to clean and validate data ex-post-facto whenever needed (the so-called data lake approach to data management, which can lead to data scientists spending 80% of their time cleaning data), but rather to formally and automatically guarantee that data integrity is preserved as it transformed (migrated, integrated, composed, queried, viewed, etc) throughout the enterprise, so that data and programs that depend on that data need not constantly be re-validated for every particular use.
Forward citations
Cited by 1 Pith paper
-
Grain Theory: Type-Level Granularity Correctness in Data Pipelines
A formal notion of 'grain' (a minimal identifying type) is defined and used to infer the grain of join results from input grains, claiming compile-time verification of pipeline correctness.
Discussion (0). Continue with ORCID to comment.