REVIEW 6 cited by
The Effects of Data Quality on Machine Learning Performance on Tabular Data
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Modern artificial intelligence (AI) applications require large quantities of training and test data. This need creates critical challenges not only concerning the availability of such data, but also regarding its quality. For example, incomplete, erroneous, or inappropriate training data can lead to unreliable models that produce ultimately poor decisions. Trustworthy AI applications require high-quality training and test data along many quality dimensions, such as accuracy, completeness, and consistency. We explore empirically the relationship between six data quality dimensions and the performance of 19 popular machine learning algorithms covering the tasks of classification, regression, and clustering, with the goal of explaining their performance in terms of data quality. Our experiments distinguish three scenarios based on the AI pipeline steps that were fed with polluted data: polluted training data, test data, or both. We conclude the paper with an extensive discussion of our observations.
Forward citations
Cited by 6 Pith papers
-
Stress-Testing ML Pipelines with Adversarial Data Corruption
SAVAGE uses dependency graphs plus beam search and Bayesian optimization to find structured data corruptions that degrade ML pipelines far more than random or manual errors.
-
Trust and Reputation in Data Sharing: A Survey
A survey of trust and reputation management systems that introduces data-sharing-specific taxonomies and finds a consistent gap: systems evaluate entities, not data quality or consumer compliance.
-
Geometric 2D Scene Graph Generation
A three-step network predicts assembly scene graphs from geometric component images, demonstrated on a four-toy-vehicle dataset with generalization to an unseen car.
-
Exploring Quantum Machine Learning for Weather Forecasting
A quantum neural network outperformed a classical RNN on 14-day temperature and 5-day wind forecasts for one Brazilian city, but on very small test sets.
-
Unfolding Data Quality Dimensions in Practice: A Survey
A systematic survey maps low-level data quality checks in seven open-source tools to six ISO/IEC 25012 data quality dimensions, revealing many-to-many relationships and fragmented terminology.
-
Eco-Friendly AI: Unleashing Data Power for Green Federated Learning
Data volume reduction and node selection in simulated federated learning cut carbon emissions by an average of 56% without sacrificing accuracy.
Discussion (0). Continue with ORCID to comment.