Pith. sign in

REVIEW 2 cited by

Object-Aware Cropping for Self-Supervised Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2112.00319 v2 pith:5KKMUP5C submitted 2021-12-01 cs.CV cs.LG

classification cs.CVcs.LG
keywords croppingself-supervisedlearningobjectimagerandomapproachcrops
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

A core component of the recent success of self-supervised learning is cropping data augmentation, which selects sub-regions of an image to be used as positive views in the self-supervised loss. The underlying assumption is that randomly cropped and resized regions of a given image share information about the objects of interest, which the learned representation will capture. This assumption is mostly satisfied in datasets such as ImageNet where there is a large, centered object, which is highly likely to be present in random crops of the full image. However, in other datasets such as OpenImages or COCO, which are more representative of real world uncurated data, there are typically multiple small objects in an image. In this work, we show that self-supervised learning based on the usual random cropping performs poorly on such datasets. We propose replacing one or both of the random crops with crops obtained from an object proposal algorithm. This encourages the model to learn both object and scene level semantic representations. Using this approach, which we call object-aware cropping, results in significant improvements over scene cropping on classification and object detection benchmarks. For example, on OpenImages, our approach achieves an improvement of 8.8% mAP over random scene-level cropping using MoCo-v2 based pre-training. We also show significant improvements on COCO and PASCAL-VOC object detection and segmentation tasks over the state-of-the-art self-supervised learning approaches. Our approach is efficient, simple and general, and can be used in most existing contrastive and non-contrastive self-supervised learning frameworks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Self-Supervised Learning with a Multi-Task Latent Space Objective

    cs.CV 2026-02 conditional novelty 6.0 of 10

    Assigning a dedicated predictor to each view type stabilizes multi-crop Siamese SSL and, combined with asymmetric cutout views, yields consistent ImageNet gains over BYOL, SimSiam, and MoCo v3.

  2. Taming the Randomness: Towards Label-Preserving Cropping in Contrastive Learning

    cs.CV 2025-04 conditional novelty 4.0 of 10

    Gaussian-centered cropping variants are claimed to improve contrastive learning accuracy on small-image benchmarks over random cropping, with gains of 2.7 to 12.4 points on CIFAR-10.

Pith tools