Pith. sign in

REVIEW 2 cited by

Review of Large Vision Models and Visual Prompt Engineering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.00855 v1 pith:XO55IQUZ submitted 2023-07-03 cs.CV cs.AI

classification cs.CVcs.AI
keywords visualengineeringpromptmodelslargevisionmethodsreview
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Visual prompt engineering is a fundamental technology in the field of visual and image Artificial General Intelligence, serving as a key component for achieving zero-shot capabilities. As the development of large vision models progresses, the importance of prompt engineering becomes increasingly evident. Designing suitable prompts for specific visual tasks has emerged as a meaningful research direction. This review aims to summarize the methods employed in the computer vision domain for large vision models and visual prompt engineering, exploring the latest advancements in visual prompt engineering. We present influential large models in the visual domain and a range of prompt engineering methods employed on these models. It is our hope that this review provides a comprehensive and systematic description of prompt engineering methods based on large visual models, offering valuable insights for future researchers in their exploration of this field.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Generating floorplans for various building functionalities via latent diffusion model

    cs.LG 2024-12 conditional novelty 5.0 of 10

    Fine-tuning a latent diffusion model on 500 footprint-text-floorplan triples yields plausible floorplans for several building types, though the quantitative evaluation is weak and irreproducible.

  2. Dspy-based Neural-Symbolic Pipeline to Enhance Spatial Reasoning in LLMs

    cs.AI 2024-11 conditional novelty 4.0 of 10

    A DSPy-orchestrated LLM plus Answer Set Programming pipeline reports 82% average accuracy on StepGame and 69% on SparQA, well above direct prompting baselines.

Pith tools