Pith. sign in

REVIEW 3 cited by

Convolutional Bypasses Are Better Vision Transformer Adapters

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2207.07039 v3 pith:W3PL5WW4 submitted 2022-07-14 cs.CV

classification cs.CV
keywords modulesadaptationvisionconvolutionalconvpassbypassesfinetunelanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The pretrain-then-finetune paradigm has been widely adopted in computer vision. But as the size of Vision Transformer (ViT) grows exponentially, the full finetuning becomes prohibitive in view of the heavier storage overhead. Motivated by parameter-efficient transfer learning (PETL) on language transformers, recent studies attempt to insert lightweight adaptation modules (e.g., adapter layers or prompt tokens) to pretrained ViT and only finetune these modules while the pretrained weights are frozen. However, these modules were originally proposed to finetune language models and did not take into account the prior knowledge specifically for visual tasks. In this paper, we propose to construct Convolutional Bypasses (Convpass) in ViT as adaptation modules, introducing only a small amount (less than 0.5% of model parameters) of trainable parameters to adapt the large ViT. Different from other PETL methods, Convpass benefits from the hard-coded inductive bias of convolutional layers and thus is more suitable for visual tasks, especially in the low-data regime. Experimental results on VTAB-1K benchmark and few-shot learning datasets show that Convpass outperforms current language-oriented adaptation modules, demonstrating the necessity to tailor vision-oriented adaptation modules for adapting vision models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads

    cs.CV 2025-06 conditional novelty 6.0 of 10

    ViT-Split freezes a vision foundation model and adds a copied task head plus a multi-scale prior head, matching or beating prior adapters with fewer parameters and up to 4x faster training.

  2. AI-driven Remote Facial Skin Hydration and TEWL Assessment from Selfie Images: A Systematic Solution

    cs.CV 2025-09 conditional novelty 5.0 of 10

    Selfie images can be used to predict skin hydration and water loss with R2 up to about 0.35, using a new dataset of 336 panelists and an adapter-based vision transformer.

  3. Weight Spectra Induced Efficient Model Adaptation

    cs.LG 2025-05 reject novelty 4.0 of 10

    Fine-tuning mostly amplifies and reorients the top singular directions of weight matrices, and SpecLoRA learns to rescale a top-left block plus LoRA to improve PEFT performance.

Pith tools