Pith. sign in

REVIEW 12 cited by

Unified Vision and Language Prompt Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.07225 v1 pith:ZA27ZAYE submitted 2022-10-13 cs.CV cs.AI

classification cs.CVcs.AI
keywords prompttuninglearningvisionvisualbenchmarksmethodsmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Prompt tuning, a parameter- and data-efficient transfer learning paradigm that tunes only a small number of parameters in a model's input space, has become a trend in the vision community since the emergence of large vision-language models like CLIP. We present a systematic study on two representative prompt tuning methods, namely text prompt tuning and visual prompt tuning. A major finding is that none of the unimodal prompt tuning methods performs consistently well: text prompt tuning fails on data with high intra-class visual variances while visual prompt tuning cannot handle low inter-class variances. To combine the best from both worlds, we propose a simple approach called Unified Prompt Tuning (UPT), which essentially learns a tiny neural network to jointly optimize prompts across different modalities. Extensive experiments on over 11 vision datasets show that UPT achieves a better trade-off than the unimodal counterparts on few-shot learning benchmarks, as well as on domain generalization benchmarks. Code and models will be released to facilitate future research.

Discussion (0). Sign in to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. BadBone: Backdoor Attacks Against Backbone Models in Visual Prompt Learning

    cs.CR 2026-05 unverdicted novelty 7.0 of 10

    BadBone backdoors backbone models with bi-level optimization to make prompt learning on downstream tasks vulnerable while preserving model utility.

  2. LAGO: Language-Guided Adaptive Object-Region Focus for Zero-Shot Visual-Text Alignment

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    LAGO achieves state-of-the-art zero-shot performance with fewer image regions by using class-agnostic object discovery followed by confidence-controlled language-guided refinement and dual-channel aggregation.

  3. MuRA: Multi-Rank Adaptation for Efficient and Effective Test-Time Vision-Language Generalization

    cs.CV 2026-08 conditional novelty 6.0 of 10

    MuRA improves label-free test-time adaptation of CLIP by routing each image's tokens to a weighted mix of rank-2 through rank-32 LoRA experts at the deepest visual layer.

  4. Plug-and-play Class-aware Knowledge Injection for Prompt Learning with Visual-Language Model

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    CAKI generates class-specific prompts from few-shot samples of the same class, stores them in a knowledge bank, and uses query-key matching to inject relevant class knowledge into test instance predictions for improve...

  5. ProPy: Building Interactive Prompt Pyramids upon CLIP for Partially Relevant Video Retrieval

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A hierarchical prompt pyramid over CLIP with ancestor-descendant attention improves partially relevant video retrieval.

  6. DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding

    cs.CV 2025-07 conditional novelty 6.0 of 10

    DynImg represents a video snippet as a keyframe plus four resized neighboring frames as temporal prompts, with a 4D rotary position embedding, and reports improved video QA accuracy.

  7. HOLa: Zero-Shot HOI Detection with Low-Rank Decomposed VLM Feature Adaptation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    HOLa achieves state-of-the-art zero-shot human-object interaction detection on HICO-DET by low-rank decomposing VLM text features and using LLM-generated action descriptions to regularize weight adaptation.

  8. One Last Attention for Your Vision-Language Model

    cs.CV 2025-07 conditional novelty 6.0 of 10

    RAda attaches one lightweight attention layer to the end of a VLM to learn a mask that reweights the final fused image-text representation, improving fine-tuning across FFT, EFT, and TTT settings.

  9. Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    MuGCP adapts CLIP by decoding instance-specific prompts from a frozen MLLM's KV cache and fusing them with visual prompts, achieving state-of-the-art few-shot classification on 14 datasets.

  10. Robust Adaptation of Foundation Models with Black-Box Visual Prompting

    cs.CV 2024-07 unverdicted novelty 6.0 of 10

    BlackVIP adapts foundation models via a Coordinator for input-dependent visual prompts and SPSA-GC for gradient estimation, enabling robust transfer on 19 datasets with low memory use and a link to randomized smoothin...

  11. C-GAP: Class-Aware and Online Prompting Improves Vision-Language Models on Imbalanced Classes

    cs.CV 2026-07 conditional novelty 5.5 of 10

    Iterative LLM caption refinement guided by minority-class AP@0.5 lifts rare-object detection on frozen open-vocabulary detectors without labels or weight updates.

  12. MMLoP: Multi-Modal Low-Rank Prompting for Efficient Vision-Language Adaptation

    cs.CV 2026-02 conditional novelty 5.0 of 10

    MMLoP compresses deep multi-modal prompts into a rank-1 shared subspace, reaching a 79.70% base-to-novel harmonic mean with 11.5K trainable parameters.

Pith tools