Pith. sign in

REVIEW 2 cited by

What does a platypus look like? Generating customized prompts for zero-shot image classification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.03320 v3 pith:GBKCXDJB submitted 2022-09-07 cs.CV cs.LG

classification cs.CVcs.LG
keywords modelsimageclassificationlanguagepromptsopen-vocabularyzero-shotaccuracy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Open-vocabulary models are a promising new paradigm for image classification. Unlike traditional classification models, open-vocabulary models classify among any arbitrary set of categories specified with natural language during inference. This natural language, called "prompts", typically consists of a set of hand-written templates (e.g., "a photo of a {}") which are completed with each of the category names. This work introduces a simple method to generate higher accuracy prompts, without relying on any explicit knowledge of the task domain and with far fewer hand-constructed sentences. To achieve this, we combine open-vocabulary models with large language models (LLMs) to create Customized Prompts via Language models (CuPL, pronounced "couple"). In particular, we leverage the knowledge contained in LLMs in order to generate many descriptive sentences that contain important discriminating characteristics of the image categories. This allows the model to place a greater importance on these regions in the image when making predictions. We find that this straightforward and general approach improves accuracy on a range of zero-shot image classification benchmarks, including over one percentage point gain on ImageNet. Finally, this simple baseline requires no additional training and remains completely zero-shot. Code available at https://github.com/sarahpratt/CuPL.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. COBRA: COmBinatorial Retrieval Augmentation for Few-Shot Adaptation

    cs.LG 2024-12 conditional novelty 6.0 of 10

    A diversity-aware combinatorial mutual information retrieval objective (COBRA) outperforms nearest-neighbor retrieval for few-shot CLIP adaptation.

  2. Does VLM Classification Benefit from LLM Description Semantics?

    cs.CV 2024-12 conditional novelty 6.0 of 10

    LLM-generated descriptions improve VLM classification only when selected to discriminate among ambiguous classes, not when simply ensembled.

Pith tools