Pith. sign in

REVIEW 2 cited by

Prompting Large Pre-trained Vision-Language Models For Compositional Concept Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.05077 v1 pith:DDMSHCGN submitted 2022-11-09 cs.CV

classification cs.CV
keywords learningcompositionalpromptcompvltextitachievesczsllargemodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This work explores the zero-shot compositional learning ability of large pre-trained vision-language models(VLMs) within the prompt-based learning framework and propose a model (\textit{PromptCompVL}) to solve the compositonal zero-shot learning (CZSL) problem. \textit{PromptCompVL} makes two design choices: first, it uses a soft-prompting instead of hard-prompting to inject learnable parameters to reprogram VLMs for compositional learning. Second, to address the compositional challenge, it uses the soft-embedding layer to learn primitive concepts in different combinations. By combining both soft-embedding and soft-prompting, \textit{PromptCompVL} achieves state-of-the-art performance on the MIT-States dataset. Furthermore, our proposed model achieves consistent improvement compared to other CLIP-based methods which shows the effectiveness of the proposed prompting strategies for CZSL.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EVA: Mixture-of-Experts Semantic Variant Alignment for Compositional Zero-Shot Learning

    cs.CV 2025-06 reject novelty 5.0 of 10

    EVA reports SOTA on MIT-States, UT-Zappos, and C-GQA in closed- and open-world CZSL by combining MoE adapters with semantic variant alignment, but missing backbone-matched baselines weaken the claim.

  2. Learning Clustering-based Prototypes for Compositional Zero-shot Learning

    cs.CV 2025-02 conditional novelty 5.0 of 10

    ClusPro improves compositional zero-shot learning by representing each primitive with multiple online-clustered prototypes and adding prototype-anchored contrastive and decorrelation losses.

Pith tools