Pith. sign in

REVIEW 38 cited by

Vision Transformer Adapter for Dense Predictions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.08534 v4 pith:Z2D2QS46 submitted 2022-05-17 cs.CV

classification cs.CV
keywords densevit-adapteradapterplaintasksvision-specificbiasesdata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This work investigates a simple yet powerful dense prediction task adapter for Vision Transformer (ViT). Unlike recently advanced variants that incorporate vision-specific inductive biases into their architectures, the plain ViT suffers inferior performance on dense predictions due to weak prior assumptions. To address this issue, we propose the ViT-Adapter, which allows plain ViT to achieve comparable performance to vision-specific transformers. Specifically, the backbone in our framework is a plain ViT that can learn powerful representations from large-scale multi-modal data. When transferring to downstream tasks, a pre-training-free adapter is used to introduce the image-related inductive biases into the model, making it suitable for these tasks. We verify ViT-Adapter on multiple dense prediction tasks, including object detection, instance segmentation, and semantic segmentation. Notably, without using extra detection data, our ViT-Adapter-L yields state-of-the-art 60.9 box AP and 53.0 mask AP on COCO test-dev. We hope that the ViT-Adapter could serve as an alternative for vision-specific transformers and facilitate future research. The code and models will be released at https://github.com/czczup/ViT-Adapter.

Discussion (0). Sign in to comment.

Forward citations

Cited by 38 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Timage: A Generative Text-in-Image Paradigm for Fine-Tuning Vision-Language Models

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    Timage generates text query overlays on images via Constrained Schrödinger Bridge to boost fine-grained spatial reasoning in vision-language models, outperforming larger systems on VMCBench with a 7B backbone.

  2. VFM$^{4}$SDG: Unveiling the Power of VFMs for Single-Domain Generalized Object Detection

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    VFM4SDG is a dual-prior framework that distills cross-domain stable relations from VFMs into DETR encoders and injects semantic-contextual priors into decoder queries to reduce missed detections in single-domain gener...

  3. DinoRADE: Full Spectral Radar-Camera Fusion with Vision Foundation Model Features for Multi-class Object Detection in Adverse Weather

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    DinoRADE reports a radar-centered multi-class detection pipeline that fuses dense radar tensors with DINOv3 features via deformable attention and outperforms prior radar-camera methods by 12.1% on the K-Radar dataset ...

  4. Delineate Anything Flow: Fast, Country-Level Field Boundary Detection from Any Source

    cs.CV 2025-11 unverdicted novelty 7.0 of 10

    DelAnyFlow combines a YOLOv11 model trained on the FBIS 22M dataset with post-processing to generate accurate vector field boundaries from multi-resolution satellite imagery, enabling country-scale mapping in hours.

  5. LiLa-WAM: Lightweight Latent Reasoning World-Action Model for Robotic Manipulation

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A 0.5B parameter latent world-action model trains end-to-end on a single GPU and reaches 90.48% average success on 50 RoboTwin 2.0 tasks with a language-free Visual Transition Token for task specification.

  6. From Multi-Resolution Cells to Gigapixel Whole Slide Images Foundation Model for Computational Pathology

    cs.CV 2026-08 conditional novelty 6.0 of 10

    MRPT, a multi-resolution hierarchical transformer pre-trained on 36K whole-slide images, is reported to outperform prior pathology foundation models on 34 classification, captioning, and VQA datasets.

  7. iFAN: Inference-Aware Learning for Plain Mask Transformers

    cs.CV 2026-08 conditional novelty 6.0 of 10

    iFAN improves query-based mask transformers by training a mask-quality head for score ranking and distilling better intermediate-layer predictions into the final layer, yielding consistent gains at no inference-time cost.

  8. FBLayout: Optimizing Memory Layout for Efficient LLM Finetuning on Mobile GPUs

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A tile-based memory layout for mobile GPUs that unifies forward and backward data access, eliminating most transpose/reshape overhead and speeding LLM fine-tuning 2.2–5.7× in the paper's measurements.

  9. REAL-OW: Rehearsal-free Open World Object Detection with Low-Rank Adaptation and Dual-Stage Objectness Modeling

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A rehearsal-free open-world detector using collaborative LoRA adapters and dual-stage objectness modeling outperforms exemplar-replay OWOD methods on standard benchmarks.

  10. UNITY: Attention Flow Networks for Adaptive Conditioning in Diffusion

    cs.CV 2026-06 conditional novelty 6.0 of 10

    A two-stage universal-then-specialize adapter with Morphable Attention Flow networks delivers SOTA FID/CLIP under single or composite conditioning at constant parameter and memory cost.

  11. Selective, Regularized, and Calibrated: Harnessing Vision Foundation Models for Cross-Domain Few-Shot Semantic Segmentation

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    HERA is a select-regularize-calibrate framework adapting frozen vision foundation models for cross-domain few-shot semantic segmentation via hierarchical layer selection with ETR, prior-guided regularization, and pixe...

  12. VFM$^{4}$SDG: Unveiling the Power of VFMs for Single-Domain Generalized Object Detection

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    VFM⁴SDG uses a frozen vision foundation model to inject cross-domain stability priors into both the encoding and decoding stages of object detectors, reducing missed detections in unseen environments.

  13. HAMSA: Scanning-Free Vision State Space Models via SpectralPulseNet

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    HAMSA achieves 85.7% ImageNet-1K top-1 accuracy as a spectral-domain SSM with 2.2x faster inference and lower memory than transformers or scanning-based SSMs.

  14. Memory-Efficient Transfer Learning with Fading Side Networks via Masked Dual Path Distillation

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    MDPD mutually distills knowledge between a frozen backbone and a learnable side network during fine-tuning, then discards the side network at inference to accelerate speed by at least 25% while preserving accuracy.

  15. MP-ISMoE: Mixed-Precision Interactive Side Mixture-of-Experts for Efficient Transfer Learning

    cs.LG 2026-04 unverdicted novelty 6.0 of 10

    MP-ISMoE uses Gaussian noise perturbed iterative quantization and interactive side mixture-of-experts to deliver higher accuracy than prior memory-efficient transfer learning methods while keeping similar parameter an...

  16. Frequency-Enhanced Diffusion Models: Curriculum-Guided Semantic Alignment for Zero-Shot Skeleton Action Recognition

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    FDSM recovers fine-grained motion details in zero-shot skeleton action recognition by integrating semantic-guided spectral residual, timestep-adaptive spectral loss, and curriculum-based semantic abstraction, reaching...

  17. Foundation Model-Driven Semantic Change Detection in Remote Sensing Imagery

    cs.CV 2026-02 unverdicted novelty 6.0 of 10

    PerASCD sets new state-of-the-art Sek scores on SECOND and LandsatSCD datasets by using a modular cascaded gated decoder on PerA foundation model features plus a new consistency loss.

  18. Exploring the Rashomon Set for Concept-Based Models

    cs.LG 2025-11 conditional novelty 6.0 of 10

    A shared frozen backbone plus per-model LoRA adapters and a concept-diversity loss trains a set of accurate CBMs that reason through different concepts.

  19. Live(r) Die: Predicting Survival in Colorectal Liver Metastasis

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A fully automated pre/post-contrast MRI framework, combining prompt-based segmentation with autoencoder multiple-instance survival analysis, improves CRLM post-surgery survival prediction over clinical and genomic bio...

  20. FoMo4Wheat: Toward reliable crop vision foundation models with globally curated data

    cs.CV 2025-09 conditional novelty 6.0 of 10

    Training a vision transformer on 2.5 million wheat images outperforms general-domain backbones across ten crop vision tasks.

  21. MPT: Motion Prompt Tuning for Micro-Expression Recognition

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    Motion Prompt Tuning with motion magnification and Gaussian tokenization claims state-of-the-art micro-expression recognition on three benchmarks.

  22. Latest Object Memory Management for Temporally Consistent Video Instance Segmentation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    LOMM achieves 54.0 AP on YouTube-VIS 2022 (offline) and 48.2 AP online, via foreground-probability-weighted memory and occupancy-guided decoupled association.

  23. Mamba Guided Boundary Prior Matters: A New Perspective for Generalized Polyp Segmentation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    SAM-MaGuP, a SAM-based polyp segmentation model with a 1D-2D Mamba adapter and boundary distillation, reports state-of-the-art mDice/mIoU on five public colonoscopy datasets.

  24. Radar-Guided Polynomial Fitting for Metric Depth Estimation

    cs.CV 2025-03 unverdicted novelty 6.0 of 10

    POLAR converts scaleless monocular depth maps to metric scale via radar-guided polynomial fitting and first-derivative regularization, claiming 24.9% MAE and 33.2% RMSE gains over prior methods on three datasets.

  25. T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models

    cs.CV 2023-02 unverdicted novelty 6.0 of 10

    T2I-Adapters are lightweight modules that enable fine-grained control over color and structure in text-to-image diffusion models by aligning external conditions with the frozen model's internal knowledge.

  26. Adapting Prithvi-EO for Fallow Detection for Food-Water Nexus: ViT-Adapter Necks and Parameter-Efficient Backbone tuning of Geospatial Foundation Model

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    Lite ViT-Adapter with LoRA on Prithvi-EO reaches mAP@50 of 0.9479 for fallow detection, improving the baseline adapter-free approach by 25.70%.

  27. Unleashing Vision Transformer Potential In Image Quality Assessment via Global-Local Adaptive Interaction

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    Proposes GLIA framework to adapt Vision Transformers for blind image quality assessment via dual-stream global-local interaction, claiming higher accuracy and robustness with reduced parameters.

  28. Beyond ViT Tokens: Masked-Diffusion Pretrained Convolutional Pathology Foundation Model for Cell-Level Dense Prediction

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    A masked-diffusion pretrained convolutional model outperforms ViT pathology foundation models on cell-level dense prediction tasks in histology.

  29. Frequency-Enhanced Diffusion Models: Curriculum-Guided Semantic Alignment for Zero-Shot Skeleton Action Recognition

    cs.CV 2026-04 unverdicted novelty 5.0 of 10

    FDSM adds spectral residual, timestep-adaptive spectral loss, and curriculum semantic abstraction to diffusion models for zero-shot skeleton-text action recognition and claims SOTA on NTU, PKU-MMD, and Kinetics-skeleton.

  30. Foundation Model-Driven Semantic Change Detection in Remote Sensing Imagery

    cs.CV 2026-02 conditional novelty 5.0 of 10

    PerASCD, a foundation-model-driven cascaded gated decoder with a soft semantic consistency loss, reports state-of-the-art Sek scores of 26.11% on SECOND and 65.21% on LandsatSCD.

  31. AI-driven Remote Facial Skin Hydration and TEWL Assessment from Selfie Images: A Systematic Solution

    cs.CV 2025-09 conditional novelty 5.0 of 10

    Selfie images can be used to predict skin hydration and water loss with R2 up to about 0.35, using a new dataset of 336 panelists and an adapter-based vision transformer.

  32. Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Concatenating monocular depth maps as an extra input channel improves video instance segmentation and reaches 56.2 AP, a new state of the art on OVIS.

  33. Uncertainty in Real-Time Semantic Segmentation on Embedded Systems

    cs.CV 2022-12 unverdicted novelty 5.0 of 10

    Combines pre-trained features, Bayesian regression, and moment propagation to enable real-time epistemic uncertainty for semantic segmentation on embedded systems while preserving accuracy.

  34. UNITY: Attention Flow Networks for Adaptive Conditioning in Diffusion

    cs.CV 2026-06 unverdicted novelty 4.0 of 10

    UNITY is a two-stage adapter with Morphable Attention Flow networks for efficient single and composite conditioning in diffusion-based image generation.

  35. Colorectal Cancer Tumor Grade Segmentation in Digital Histopathology Images: From Giga to Mini Challenge

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A challenge summary showing top-performing deep learning methods improve colorectal cancer tumor grade segmentation on the METU CCTGS dataset, with the winner reaching 70.2 macro F-score.

  36. SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A three-stage coarse-to-fine training recipe for vision backbones produces consistent benchmark gains for lightweight multimodal LLMs.

  37. Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey

    cs.LG 2024-03 accept novelty 4.0 of 10

    A comprehensive survey of PEFT algorithms for large models, covering their performance, overhead, applications, and real-world system implementations.

  38. State Space Models Meet Remote Sensing: A Survey

    cs.CV 2026-06 unverdicted novelty 2.0 of 10

    A literature survey of State Space Model methods applied to remote sensing tasks, architectures, and challenges since their introduction to the field.

Pith tools