Pith. sign in

REVIEW 10 cited by

A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.12980 v1 pith:WKYYBME3 submitted 2023-07-24 cs.CV

A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

classification cs.CV
keywords modelsengineeringpromptmodelvision-languagelanguagenaturalpre-trained
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Prompt engineering is a technique that involves augmenting a large pre-trained model with task-specific hints, known as prompts, to adapt the model to new tasks. Prompts can be created manually as natural language instructions or generated automatically as either natural language instructions or vector representations. Prompt engineering enables the ability to perform predictions based solely on prompts without updating model parameters, and the easier application of large pre-trained models in real-world tasks. In past years, Prompt engineering has been well-studied in natural language processing. Recently, it has also been intensively studied in vision-language modeling. However, there is currently a lack of a systematic overview of prompt engineering on pre-trained vision-language models. This paper aims to provide a comprehensive survey of cutting-edge research in prompt engineering on three types of vision-language models: multimodal-to-text generation models (e.g. Flamingo), image-text matching models (e.g. CLIP), and text-to-image generation models (e.g. Stable Diffusion). For each type of model, a brief model summary, prompting methods, prompting-based applications, and the corresponding responsibility and integrity issues are summarized and discussed. Furthermore, the commonalities and differences between prompting on vision-language models, language models, and vision models are also discussed. The challenges, future directions, and research opportunities are summarized to foster future research on this topic.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media

    cs.CL 2026-05 unverdicted novelty 7.0

    PluRule is a new multimodal multilingual benchmark showing that state-of-the-art vision-language models perform only marginally better than a trivial baseline at detecting specific rule violations in pluralistic onlin...

  2. Visual prompt engineering for video models

    cs.CV 2026-07 conditional novelty 6.0

    Automatically converting task images to photorealistic variants (visual prompt engineering) improves video-model reasoning performance, often beating text prompt engineering and test-time scaling.

  3. ZEBRA: Zero-Shot Entropy-Regularized Prompt Learning for Base-to-Novel Generalization in Audio-Language Models

    cs.SD 2026-06 unverdicted novelty 6.0

    ZEBRA reduces the base-to-novel generalization gap in audio-language models by fusing zero-shot and prompt-learning logits with entropy regularization.

  4. AERMANI-VLM: Structured Prompting and Reasoning for Aerial Manipulation with Vision Language Models

    cs.RO 2025-11 conditional novelty 6.0

    Structured prompting plus a discrete skill library lets a frozen VLM direct aerial manipulation, reaching 87.5% simulated and 80% hardware success in pick-and-place tasks.

  5. Modality Alignment across Trees on Heterogeneous Hyperbolic Manifolds

    cs.CV 2025-10 reject novelty 6.0

    A VLM method aligns hierarchical image and text feature trees across hyperbolic manifolds of different curvatures via an intermediate manifold, but the theoretical justification is flawed.

  6. EPIG: Emotion-Based Prompting for Personalised Image Generation

    cs.AI 2026-06 unverdicted novelty 5.0

    EPIG is a training-free prompt enrichment technique using valence-arousal representations that reduces mean arousal error by 14% versus naive insertion and 12% versus LLM expansion on a 10-prompt benchmark while prese...

  7. Does Your VFM Speak Plant? The Botanical Grammar of Vision Foundation Models for Object Detection

    cs.CV 2026-04 unverdicted novelty 5.0

    Optimized prompts for vision foundation models improve cowpea detection accuracy by over 0.35 mAP on synthetic data and transfer effectively to real fields without manual annotations.

  8. Are vision-language models ready to zero-shot replace supervised classification models in agriculture?

    cs.CV 2025-12 unverdicted novelty 4.0

    Zero-shot VLMs reach at most 62% accuracy on agricultural classification tasks while supervised models like YOLO11 perform markedly higher, indicating they are not ready to replace task-specific systems.

  9. A Survey of Personalized Federated Foundation Models for Privacy-Preserving Recommendation

    cs.LG 2025-06 unverdicted novelty 3.0

    A survey of personalization techniques and foundation model adaptations in federated settings for privacy-preserving recommendations, emphasizing their architectural intersection.

  10. Multilingual and Multimodal LLMs in the Wild: Building for Low-Resource Languages

    cs.CL 2026-05 unverdicted novelty 2.0

    A tutorial synthesizing foundations, recent models such as PALO and Maya, and low-cost methods for tri-modal multilingual AI in resource-constrained settings.