Pith. sign in

Unmasking and improving data credibility: A study with datasets for training harmless language models

4 Pith papers cite this work. Polarity classification is still indexing.

4 Pith papers citing it

years

2026 3 2024 1

verdicts

UNVERDICTED 4

representative citing papers

Evian: Towards Explainable Visual Instruction-tuning Data Auditing

cs.CV · 2026-04-22 · unverdicted · novelty 6.0

EVian decomposes vision-language model responses into three cognitive components and audits them along consistency, coherence, and accuracy axes, showing that a small curated subset outperforms much larger training sets.

A Data-Centric Framework for Detecting and Correcting Corrupted Labels

cs.LG · 2026-06-10 · unverdicted · novelty 4.0

Relabeler is an end-to-end framework that detects corrupted labels via local and global instance relationships and corrects them using feature-based estimation, reporting up to 58% better label correction precision than baselines.

Noise-Aware Framework for Correcting Corrupted Labels

cs.LG · 2026-06-10 · unverdicted · novelty 4.0

CANOLA estimates label noise and performs cautious iterative soft-label refinement to correct corrupted training data, reporting 19-52% error reduction versus prior methods on six datasets.

citing papers explorer

Showing 4 of 4 citing papers.

  • Evian: Towards Explainable Visual Instruction-tuning Data Auditing cs.CV · 2026-04-22 · unverdicted · none · ref 9

    EVian decomposes vision-language model responses into three cognitive components and audits them along consistency, coherence, and accuracy axes, showing that a small curated subset outperforms much larger training sets.

  • Automatic Dataset Construction (ADC): Sample Collection, Data Curation, and Beyond cs.AI · 2024-08-21 · unverdicted · none · ref 25

    The ADC method automates the creation of large image classification datasets using LLMs and search engines, achieving 79% human agreement and reducing label noise on a 1 million image clothing dataset, while also releasing benchmarks for noise and bias issues.

  • A Data-Centric Framework for Detecting and Correcting Corrupted Labels cs.LG · 2026-06-10 · unverdicted · none · ref 27

    Relabeler is an end-to-end framework that detects corrupted labels via local and global instance relationships and corrects them using feature-based estimation, reporting up to 58% better label correction precision than baselines.

  • Noise-Aware Framework for Correcting Corrupted Labels cs.LG · 2026-06-10 · unverdicted · none · ref 22

    CANOLA estimates label noise and performs cautious iterative soft-label refinement to correct corrupted training data, reporting 19-52% error reduction versus prior methods on six datasets.