REVIEW 20 cited by
Unified Vision and Language Prompt Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Prompt tuning, a parameter- and data-efficient transfer learning paradigm that tunes only a small number of parameters in a model's input space, has become a trend in the vision community since the emergence of large vision-language models like CLIP. We present a systematic study on two representative prompt tuning methods, namely text prompt tuning and visual prompt tuning. A major finding is that none of the unimodal prompt tuning methods performs consistently well: text prompt tuning fails on data with high intra-class visual variances while visual prompt tuning cannot handle low inter-class variances. To combine the best from both worlds, we propose a simple approach called Unified Prompt Tuning (UPT), which essentially learns a tiny neural network to jointly optimize prompts across different modalities. Extensive experiments on over 11 vision datasets show that UPT achieves a better trade-off than the unimodal counterparts on few-shot learning benchmarks, as well as on domain generalization benchmarks. Code and models will be released to facilitate future research.
Forward citations
Cited by 20 Pith papers
-
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval
CIR-LVLM fine-tunes Qwen-VL-Chat with LoRA and hybrid task and instance-specific prompts to produce query and target embeddings, achieving new state-of-the-art recall on Fashion-IQ, Shoes, and CIRR.
-
MuRA: Multi-Rank Adaptation for Efficient and Effective Test-Time Vision-Language Generalization
MuRA improves label-free test-time adaptation of CLIP by routing each image's tokens to a weighted mix of rank-2 through rank-32 LoRA experts at the deepest visual layer.
-
ProPy: Building Interactive Prompt Pyramids upon CLIP for Partially Relevant Video Retrieval
A hierarchical prompt pyramid over CLIP with ancestor-descendant attention improves partially relevant video retrieval.
-
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding
DynImg represents a video snippet as a keyframe plus four resized neighboring frames as temporal prompts, with a 4D rotary position embedding, and reports improved video QA accuracy.
-
HOLa: Zero-Shot HOI Detection with Low-Rank Decomposed VLM Feature Adaptation
HOLa achieves state-of-the-art zero-shot human-object interaction detection on HICO-DET by low-rank decomposing VLM text features and using LLM-generated action descriptions to regularize weight adaptation.
-
One Last Attention for Your Vision-Language Model
RAda attaches one lightweight attention layer to the end of a VLM to learn a mask that reweights the final fused image-text representation, improving fine-tuning across FFT, EFT, and TTT settings.
-
Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models
MuGCP adapts CLIP by decoding instance-specific prompts from a frozen MLLM's KV cache and fusing them with visual prompts, achieving state-of-the-art few-shot classification on 14 datasets.
-
From Local Details to Global Context: Advancing Vision-Language Models with Attention-Based Selection
Attention-guided cropping in raw and feature space plus soft matching improves CLIP zero-shot classification and out-of-distribution generalization by 0.3 to 6.2 points, with no training.
-
Diff-Prompt: Diffusion-Driven Prompt Generator with Mask Supervision
Diff-Prompt uses a mask-supervised diffusion model to generate input-specific visual and textual prompts, improving GLIP on referring expression comprehension beyond existing prompt tuning methods.
-
ProKeR: A Kernel Perspective on Few-Shot Adaptation of Large Vision-Language Models
ProKeR treats CLIP cache models as Nadaraya-Watson estimators and fits a proximally regularized kernel ridge regression in an RKHS, reporting state-of-the-art training-free few-shot accuracy on 11 datasets.
-
Differentiable Prompt Learning for Vision Language Models
DPL searches the per-layer prompt length for CLIP with differentiable architecture search, and reports higher few-shot accuracy than fixed-length prompt baselines.
-
Efficient Policy Adaptation with Contrastive Prompt Ensemble for Embodied Agents
Contrastively learned visual prompts, combined through guided attention, improve zero-shot visual domain adaptation of embodied RL policies.
-
C-GAP: Class-Aware and Online Prompting Improves Vision-Language Models on Imbalanced Classes
Iterative LLM caption refinement guided by minority-class AP@0.5 lifts rare-object detection on frozen open-vocabulary detectors without labels or weight updates.
-
MMLoP: Multi-Modal Low-Rank Prompting for Efficient Vision-Language Adaptation
MMLoP compresses deep multi-modal prompts into a rank-1 shared subspace, reaching a 79.70% base-to-novel harmonic mean with 11.5K trainable parameters.
-
Handling Imbalanced Pseudolabels for Vision-Language Models with Concept Alignment and Confusion-Aware Calibrated Margin
A concept-alignment and confusion-aware margin framework improves CLIP fine-tuning with imbalanced pseudolabels, but the headline 6.29% gain is limited to the unsupervised setting.
-
INT: Instance-Specific Negative Mining for Task-Generic Promptable Segmentation
INT improves task-generic promptable segmentation by progressively mining negative candidates, using VLM output differences after masking to select and refine instance-specific prompts.
-
A Wander Through the Multimodal Landscape: Efficient Transfer Learning via Low-rank Sequence Multimodal Adapter
Wander is a low-rank sequence adapter that fuses token-level features across an arbitrary number of modalities with far fewer parameters than full fine-tuning.
-
Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning
SPIDER updates only parameters whose fine-tuning gradient importance exceeds their pre-trained weight importance, reducing catastrophic forgetting and improving downstream performance in multimodal LLM fine-tuning.
-
CLIP-Powered Domain Generalization and Domain Adaptation: A Comprehensive Survey
CLIP-powered domain generalization and domain adaptation methods are surveyed and categorized into prompt-learning versus backbone use, and source-available versus source-free settings.
-
Generalizing vision-language models to novel domains: A comprehensive survey
A survey of VLM generalization literature organized by transferred module, with benchmark tables and a review of multimodal LLMs.
Discussion (0). Continue with ORCID to comment.