Pith. sign in

Are all training examples equally valuable?

2 Pith papers cite this work. Polarity classification is still indexing.

2 Pith papers citing it
abstract

When learning a new concept, not all training examples may prove equally useful for training: some may have higher or lower training value than others. The goal of this paper is to bring to the attention of the vision community the following considerations: (1) some examples are better than others for training detectors or classifiers, and (2) in the presence of better examples, some examples may negatively impact performance and removing them may be beneficial. In this paper, we propose an approach for measuring the training value of an example, and use it for ranking and greedily sorting examples. We test our methods on different vision tasks, models, datasets and classifiers. Our experiments show that the performance of current state-of-the-art detectors and classifiers can be improved when training on a subset, rather than the whole training set.

fields

cs.LG 2

years

2026 1 2018 1

verdicts

UNVERDICTED 2

representative citing papers

Dataset Distillation

cs.LG · 2018-11-27 · unverdicted · novelty 8.0

Dataset distillation creates a tiny synthetic training set that, when used with a fixed network initialization, produces models whose performance approximates that of models trained on the full original dataset.

citing papers explorer

Showing 2 of 2 citing papers.

  • Dataset Distillation cs.LG · 2018-11-27 · unverdicted · none · ref 100 · internal anchor

    Dataset distillation creates a tiny synthetic training set that, when used with a fixed network initialization, produces models whose performance approximates that of models trained on the full original dataset.

  • Scaling-Aware Data Selection for End-to-End Autonomous Driving Systems cs.LG · 2026-04-09 · unverdicted · none · ref 30

    MOSAIC is a scaling-aware data selection framework that outperforms baselines in training end-to-end autonomous driving planners, achieving comparable or better EPDMS scores with up to 80% less data.