Pith. sign in

REVIEW 1 cited by

Exploring Localization for Self-supervised Fine-grained Contrastive Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.15788 v4 pith:7ZESY2CO submitted 2021-06-30 cs.CV

classification cs.CV
keywords learningcontrastivefine-grainedself-supervisedforegroundalignmentclassificationcross-view
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Self-supervised contrastive learning has demonstrated great potential in learning visual representations. Despite their success in various downstream tasks such as image classification and object detection, self-supervised pre-training for fine-grained scenarios is not fully explored. We point out that current contrastive methods are prone to memorizing background/foreground texture and therefore have a limitation in localizing the foreground object. Analysis suggests that learning to extract discriminative texture information and localization are equally crucial for fine-grained self-supervised pre-training. Based on our findings, we introduce cross-view saliency alignment (CVSA), a contrastive learning framework that first crops and swaps saliency regions of images as a novel view generation and then guides the model to localize on foreground objects via a cross-view alignment loss. Extensive experiments on both small- and large-scale fine-grained classification benchmarks show that CVSA significantly improves the learned representation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PP-SSL : Priority-Perception Self-Supervised Learning for Fine-Grained Recognition

    cs.CV 2024-11 conditional novelty 6.0 of 10

    PP-SSL combines CLIP text-guided distillation with original-image GradCAM guidance to improve self-supervised fine-grained recognition, reporting state-of-the-art results on seven benchmarks.

Pith tools