Pith. sign in

REVIEW 1 cited by

Sample Selection with Uncertainty of Losses for Learning with Noisy Labels

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.00445 v1 pith:V63YPUET submitted 2021-06-01 cs.LG

classification cs.LG
keywords datalosseslabelslarge-lossnoisysampleselectionuncertainty
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In learning with noisy labels, the sample selection approach is very popular, which regards small-loss data as correctly labeled during training. However, losses are generated on-the-fly based on the model being trained with noisy labels, and thus large-loss data are likely but not certainly to be incorrect. There are actually two possibilities of a large-loss data point: (a) it is mislabeled, and then its loss decreases slower than other data, since deep neural networks "learn patterns first"; (b) it belongs to an underrepresented group of data and has not been selected yet. In this paper, we incorporate the uncertainty of losses by adopting interval estimation instead of point estimation of losses, where lower bounds of the confidence intervals of losses derived from distribution-free concentration inequalities, but not losses themselves, are used for sample selection. In this way, we also give large-loss but less selected data a try; then, we can better distinguish between the cases (a) and (b) by seeing if the losses effectively decrease with the uncertainty after the try. As a result, we can better explore underrepresented data that are correctly labeled but seem to be mislabeled at first glance. Experiments demonstrate that the proposed method is superior to baselines and robust to a broad range of label noise types.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Sanitizing Manufacturing Dataset Labels Using Vision-Language Models

    cs.CV 2025-06 reject novelty 4.0 of 10

    A CLIP-based pipeline for cleaning noisy multi-label manufacturing image data, tested on Factorynet, reduces the label vocabulary from 6,426 to 408 distinct labels through similarity scoring and clustering.

Pith tools