Pith. sign in

REVIEW 2 cited by

A Simple Recipe for Competitive Low-compute Self supervised Vision Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.09451 v1 pith:ENYPD5QE submitted 2023-01-23 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords modelsarchitecturesdistillationlargeself-supervisedsupervisedbranchclevr
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Self-supervised methods in vision have been mostly focused on large architectures as they seem to suffer from a significant performance drop for smaller architectures. In this paper, we propose a simple self-supervised distillation technique that can train high performance low-compute neural networks. Our main insight is that existing joint-embedding based SSL methods can be repurposed for knowledge distillation from a large self-supervised teacher to a small student model. Thus, we call our method Replace one Branch (RoB) as it simply replaces one branch of the joint-embedding training with a large teacher model. RoB is widely applicable to a number of architectures such as small ResNets, MobileNets and ViT, and pretrained models such as DINO, SwAV or iBOT. When pretraining on the ImageNet dataset, RoB yields models that compete with supervised knowledge distillation. When applied to MSN, RoB produces students with strong semi-supervised capabilities. Finally, our best ViT-Tiny models improve over prior SSL state-of-the-art on ImageNet by $2.3\%$ and are on par or better than a supervised distilled DeiT on five downstream transfer tasks (iNaturalist, CIFAR, Clevr/Count, Clevr/Dist and Places). We hope RoB enables practical self-supervision at smaller scale.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. On the Surprising Effectiveness of Attention Transfer for Vision Transformers

    cs.LG 2024-11 conditional novelty 7.0 of 10

    Transferring only the attention maps of a pre-trained ViT recovers the accuracy gain of full fine-tuning on ImageNet-1K classification.

  2. Distilling foundation models for robust and efficient models in digital pathology

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A distilled 86M-parameter pathology model reaches near state-of-the-art performance on EVA and HEST benchmarks and shows strong robustness to scanner and staining variation.

Pith tools