Pith. sign in

REVIEW 6 cited by

The Effects of Data Quality on Machine Learning Performance on Tabular Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2207.14529 v6 pith:PQ2TJC3B submitted 2022-07-29 cs.DB

classification cs.DB
keywords dataqualitytrainingperformancetestapplicationsdimensionslearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Modern artificial intelligence (AI) applications require large quantities of training and test data. This need creates critical challenges not only concerning the availability of such data, but also regarding its quality. For example, incomplete, erroneous, or inappropriate training data can lead to unreliable models that produce ultimately poor decisions. Trustworthy AI applications require high-quality training and test data along many quality dimensions, such as accuracy, completeness, and consistency. We explore empirically the relationship between six data quality dimensions and the performance of 19 popular machine learning algorithms covering the tasks of classification, regression, and clustering, with the goal of explaining their performance in terms of data quality. Our experiments distinguish three scenarios based on the AI pipeline steps that were fed with polluted data: polluted training data, test data, or both. We conclude the paper with an extensive discussion of our observations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Stress-Testing ML Pipelines with Adversarial Data Corruption

    cs.LG 2025-06 conditional novelty 7.0 of 10

    SAVAGE uses dependency graphs plus beam search and Bayesian optimization to find structured data corruptions that degrade ML pipelines far more than random or manual errors.

  2. Trust and Reputation in Data Sharing: A Survey

    cs.SI 2025-08 conditional novelty 5.0 of 10

    A survey of trust and reputation management systems that introduces data-sharing-specific taxonomies and finds a consistent gap: systems evaluate entities, not data quality or consumer compliance.

  3. Geometric 2D Scene Graph Generation

    cs.CV 2026-07 reject novelty 4.0 of 10

    A three-step network predicts assembly scene graphs from geometric component images, demonstrated on a four-toy-vehicle dataset with generalization to an unseen car.

  4. Exploring Quantum Machine Learning for Weather Forecasting

    quant-ph 2025-09 conditional novelty 4.0 of 10

    A quantum neural network outperformed a classical RNN on 14-day temperature and 5-day wind forecasts for one Brazilian city, but on very small test sets.

  5. Unfolding Data Quality Dimensions in Practice: A Survey

    cs.DB 2025-07 accept novelty 4.0 of 10

    A systematic survey maps low-level data quality checks in seven open-source tools to six ISO/IEC 25012 data quality dimensions, revealing many-to-many relationships and fragmented terminology.

  6. Eco-Friendly AI: Unleashing Data Power for Green Federated Learning

    cs.LG 2025-07 conditional novelty 4.0 of 10

    Data volume reduction and node selection in simulated federated learning cut carbon emissions by an average of 56% without sacrificing accuracy.

Pith tools