Pith. sign in

REVIEW 3 cited by

Multitask Vision-Language Prompt Tuning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.11720 v3 pith:NQ7C5VQS submitted 2022-11-21 cs.CV cs.CL

classification cs.CVcs.CL
keywords prompttuningtasksvision-languagemvlpttaskcross-taskknowledge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Prompt Tuning, conditioning on task-specific learned prompt vectors, has emerged as a data-efficient and parameter-efficient method for adapting large pretrained vision-language models to multiple downstream tasks. However, existing approaches usually consider learning prompt vectors for each task independently from scratch, thereby failing to exploit the rich shareable knowledge across different vision-language tasks. In this paper, we propose multitask vision-language prompt tuning (MVLPT), which incorporates cross-task knowledge into prompt tuning for vision-language models. Specifically, (i) we demonstrate the effectiveness of learning a single transferable prompt from multiple source tasks to initialize the prompt for each target task; (ii) we show many target tasks can benefit each other from sharing prompt vectors and thus can be jointly learned via multitask prompt tuning. We benchmark the proposed MVLPT using three representative prompt tuning methods, namely text prompt tuning, visual prompt tuning, and the unified vision-language prompt tuning. Results in 20 vision tasks demonstrate that the proposed approach outperforms all single-task baseline prompt tuning methods, setting the new state-of-the-art on the few-shot ELEVATER benchmarks and cross-task generalization benchmarks. To understand where the cross-task knowledge is most effective, we also conduct a large-scale study on task transferability with 20 vision tasks in 400 combinations for each prompt tuning method. It shows that the most performant MVLPT for each prompt tuning method prefers different task combinations and many tasks can benefit each other, depending on their visual similarity and label similarity. Code is available at https://github.com/sIncerass/MVLPT.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SynBridge: Bridging Reaction States via Discrete Flow for Bidirectional Reaction Prediction

    cs.LG 2025-07 conditional novelty 5.0 of 10

    A bidirectional discrete flow matching model, SynBridge, predicts reaction products and reactants on graph representations and reports state-of-the-art Top-k accuracy on USPTO-50K, USPTO-MIT, and Pistachio.

  2. Context-Based Semantic-Aware Alignment for Semi-Supervised Multi-Label Learning

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A CLIP-based semi-supervised multi-label method that aligns class prompts with label-specific image features and adds a context-identification auxiliary task, reporting SOTA mAP on COCO, VOC, and NUS-WIDE.

  3. Enhancing Parameter-Efficient Fine-Tuning of Vision Transformers through Frequency-Based Adaptation

    cs.CV 2024-11 reject novelty 5.0 of 10

    FreqFit is a frequency-domain filter module that, when inserted between ViT blocks, improves the accuracy of existing PEFT methods on most but not all evaluated benchmarks.

Pith tools