Pith. sign in

REVIEW 3 cited by

Contrastive Learning Rivals Masked Image Modeling in Fine-tuning via Feature Distillation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.14141 v3 pith:II6BHMF2 submitted 2022-05-27 cs.CV cs.LG

classification cs.CVcs.LG
keywords fine-tuningrepresentationsimagecontrastivedistillationfeatureimprovedlearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Masked image modeling (MIM) learns representations with remarkably good fine-tuning performances, overshadowing previous prevalent pre-training approaches such as image classification, instance contrastive learning, and image-text alignment. In this paper, we show that the inferior fine-tuning performance of these pre-training approaches can be significantly improved by a simple post-processing in the form of feature distillation (FD). The feature distillation converts the old representations to new representations that have a few desirable properties just like those representations produced by MIM. These properties, which we aggregately refer to as optimization friendliness, are identified and analyzed by a set of attention- and optimization-related diagnosis tools. With these properties, the new representations show strong fine-tuning performance. Specifically, the contrastive self-supervised learning methods are made as competitive in fine-tuning as the state-of-the-art masked image modeling (MIM) algorithms. The CLIP models' fine-tuning performance is also significantly improved, with a CLIP ViT-L model reaching 89.0% top-1 accuracy on ImageNet-1K classification. On the 3-billion-parameter SwinV2-G model, the fine-tuning accuracy is improved by +1.5 mIoU / +1.1 mAP to 61.4 mIoU / 64.2 mAP on ADE20K semantic segmentation and COCO object detection, respectively, creating new records on both benchmarks. More importantly, our work provides a way for the future research to focus more effort on the generality and scalability of the learnt representations without being pre-occupied with optimization friendliness since it can be enhanced rather easily. The code will be available at https://github.com/SwinTransformer/Feature-Distillation.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Pretraining Multiple Instance Learning Networks with Multi-Teacher Distillation from Pathology Slide Foundation Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Distilling TITAN and CARE slide embeddings into MIL aggregators gives reusable pretrained weights that beat from-scratch training on most of 15 pathology tasks, with the largest gains in few-shot and linear-probing settings.

  2. Make me an Expert: Distilling from Generalist Black-Box Models into Specialized Models for Semantic Segmentation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    ATGC selects the best input scale for a black-box open-vocabulary segmentation API, using DINOv2 attention entropy, improving one-hot-label distillation on Cityscapes and ACDC.

  3. PiPViT: Patch-based Visual Interpretable Prototypes for Retinal Image Analysis

    cs.CV 2025-06 conditional novelty 4.0 of 10

    PiPViT combines vision transformers and prototype learning to classify retinal OCT scans while showing the spatial extent of the biomarker that drove the decision.

Pith tools