Pith. sign in

REVIEW 1 cited by

Learning From Less Data: Diversified Subset Selection and Active Learning in Image Classification Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1805.11191 v1 pith:H3NTC2A3 submitted 2018-05-28 cs.CV cs.LGstat.ML

classification cs.CVcs.LGstat.ML
keywords subsetdatalearningselectiontrainingactivelabelingless
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Supervised machine learning based state-of-the-art computer vision techniques are in general data hungry and pose the challenges of not having adequate computing resources and of high costs involved in human labeling efforts. Training data subset selection and active learning techniques have been proposed as possible solutions to these challenges respectively. A special class of subset selection functions naturally model notions of diversity, coverage and representation and they can be used to eliminate redundancy and thus lend themselves well for training data subset selection. They can also help improve the efficiency of active learning in further reducing human labeling efforts by selecting a subset of the examples obtained using the conventional uncertainty sampling based techniques. In this work we empirically demonstrate the effectiveness of two diversity models, namely the Facility-Location and Disparity-Min models for training-data subset selection and reducing labeling effort. We do this for a variety of computer vision tasks including Gender Recognition, Scene Recognition and Object Recognition. Our results show that subset selection done in the right way can add 2-3% in accuracy on existing baselines, particularly in the case of less training data. This allows the training of complex machine learning models (like Convolutional Neural Networks) with much less training data while incurring minimal performance loss.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Data-Driven Novelty Score for Diverse In-Vehicle Data Recording

    cs.CV 2025-07 conditional novelty 4.0 of 10

    An online Mahalanobis-distance novelty filter, updated with streaming data, selects a smaller traffic-sign training set that can outperform the full dataset and random sampling.

Pith tools