Pith. sign in

REVIEW 1 cited by

Unsupervised Pre-Training of Image Features on Non-Curated Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1905.01278 v3 pith:HGF6D232 submitted 2019-05-03 cs.CV

classification cs.CV
keywords unsuperviseddataavailabledatasetspre-trainingapproachcuratedfeature
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Pre-training general-purpose visual features with convolutional neural networks without relying on annotations is a challenging and important task. Most recent efforts in unsupervised feature learning have focused on either small or highly curated datasets like ImageNet, whereas using uncurated raw datasets was found to decrease the feature quality when evaluated on a transfer task. Our goal is to bridge the performance gap between unsupervised methods trained on curated data, which are costly to obtain, and massive raw datasets that are easily available. To that effect, we propose a new unsupervised approach which leverages self-supervision and clustering to capture complementary statistics from large-scale data. We validate our approach on 96 million images from YFCC100M, achieving state-of-the-art results among unsupervised methods on standard benchmarks, which confirms the potential of unsupervised learning when only uncurated data are available. We also show that pre-training a supervised VGG-16 with our method achieves 74.9% top-1 classification accuracy on the validation set of ImageNet, which is an improvement of +0.8% over the same network trained from scratch. Our code is available at https://github.com/facebookresearch/DeeperCluster.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Temporal Object Captioning for Street Scene Videos from LiDAR Tracks

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A rule-based pipeline converts LiDAR tracks into template captions of traffic dynamics, and training SwinBERT on them lowers the Visual Bias Measure across three datasets.

Pith tools