Pith. sign in

REVIEW 2 cited by

Inter-Instance Similarity Modeling for Contrastive Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.12243 v3 pith:W6WPJJEB submitted 2023-06-21 cs.CV

classification cs.CV
keywords contrastiveimageslearninginter-instancemethodpatchmixaccuracyexisting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The existing contrastive learning methods widely adopt one-hot instance discrimination as pretext task for self-supervised learning, which inevitably neglects rich inter-instance similarities among natural images, then leading to potential representation degeneration. In this paper, we propose a novel image mix method, PatchMix, for contrastive learning in Vision Transformer (ViT), to model inter-instance similarities among images. Following the nature of ViT, we randomly mix multiple images from mini-batch in patch level to construct mixed image patch sequences for ViT. Compared to the existing sample mix methods, our PatchMix can flexibly and efficiently mix more than two images and simulate more complicated similarity relations among natural images. In this manner, our contrastive framework can significantly reduce the gap between contrastive objective and ground truth in reality. Experimental results demonstrate that our proposed method significantly outperforms the previous state-of-the-art on both ImageNet-1K and CIFAR datasets, e.g., 3.0% linear accuracy improvement on ImageNet-1K and 8.7% kNN accuracy improvement on CIFAR100. Moreover, our method achieves the leading transfer performance on downstream tasks, object detection and instance segmentation on COCO dataset. The code is available at https://github.com/visresearch/patchmix

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multiple Object Stitching for Unsupervised Representation Learning

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Multiple Object Stitching improves self-supervised representations by training on stitched multi-object images and reports gains on ImageNet, CIFAR, and COCO.

  2. Learning Compact Vision Tokens for Efficient Large Multimodal Models

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A learned spatial token fusion plus multi-block features lets a multimodal model use only 25% of its vision tokens while matching or exceeding baseline accuracy on eight benchmarks.

Pith tools