EVian decomposes vision-language model responses into three cognitive components and audits them along consistency, coherence, and accuracy axes, showing that a small curated subset outperforms much larger training sets.
Unmasking and improving data credibility: A study with datasets for training harmless language models
4 Pith papers cite this work. Polarity classification is still indexing.
verdicts
UNVERDICTED 4representative citing papers
The ADC method automates the creation of large image classification datasets using LLMs and search engines, achieving 79% human agreement and reducing label noise on a 1 million image clothing dataset, while also releasing benchmarks for noise and bias issues.
Relabeler is an end-to-end framework that detects corrupted labels via local and global instance relationships and corrects them using feature-based estimation, reporting up to 58% better label correction precision than baselines.
CANOLA estimates label noise and performs cautious iterative soft-label refinement to correct corrupted training data, reporting 19-52% error reduction versus prior methods on six datasets.
citing papers explorer
-
Evian: Towards Explainable Visual Instruction-tuning Data Auditing
EVian decomposes vision-language model responses into three cognitive components and audits them along consistency, coherence, and accuracy axes, showing that a small curated subset outperforms much larger training sets.
-
Automatic Dataset Construction (ADC): Sample Collection, Data Curation, and Beyond
The ADC method automates the creation of large image classification datasets using LLMs and search engines, achieving 79% human agreement and reducing label noise on a 1 million image clothing dataset, while also releasing benchmarks for noise and bias issues.
-
A Data-Centric Framework for Detecting and Correcting Corrupted Labels
Relabeler is an end-to-end framework that detects corrupted labels via local and global instance relationships and corrects them using feature-based estimation, reporting up to 58% better label correction precision than baselines.
-
Noise-Aware Framework for Correcting Corrupted Labels
CANOLA estimates label noise and performs cautious iterative soft-label refinement to correct corrupted training data, reporting 19-52% error reduction versus prior methods on six datasets.