Pith. sign in

REVIEW 4 major objections 5 minor 61 references

This paper claims that visually editing a video model's input image, converting abstract sketches into photorealistic scenes, systematically improves its reasoning performance, often more than text prompt engineering or test-time scaling.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 02:07 UTC pith:IBFVXUTG

load-bearing objection The VIPE concept is worth taking seriously, but the headline VPCT comparison is confounded and needs a clean replication before the central claim is trustworthy. the 4 major comments →

arxiv 2607.25537 v1 pith:IBFVXUTG submitted 2026-07-28 cs.CV cs.AI

Visual prompt engineering for video models

classification cs.CV cs.AI
keywords visual prompt engineeringvideo modelsvisual reasoningrealism biasprompt engineeringtest-time scalingimage editingphysics reasoning
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper asks whether video models benefit from engineering their input prompt the way language models do, except the prompt is an image. It finds that transforming an abstract task image into a realistic-looking version, a process it calls visual prompt engineering (VIPE), improves reasoning accuracy across physics, maze-solving, puzzle, and sorting tasks. The improvement can exceed what classic text-prompt engineering or self-consistency sampling delivers, and the two strategies can be combined. The authors explain this through a realism bias: video models reason more reliably in photorealistic scenes, where objects stay consistent, so abstract benchmarks systematically understate the models' true ability. A sympathetic reader would care because VIPE is a cheap, automatic, model-agnostic lever for eliciting better visual reasoning from video generators, and it reframes benchmark scores as lower bounds on competence.

Core claim

The central discovery is that video models have a systematic preference for photorealistic visual input: given the same underlying task, replacing an abstract sketch with a realistic rendering of the same scene improves the model's ability to solve it. In the paper's flagship physics experiment, transforming a ball-and-ramp sketch into a photograph-like scene moves accuracy from near chance to well above chance across several video models, and a step-by-step realism ladder shows that each step toward realism increases generation consistency. The authors argue that the root cause is a representation gap: video models are trained predominantly on realistic footage, so abstract inputs force rea

What carries the argument

The central mechanism is visual prompt engineering (VIPE), a pipeline that starts with an ideator proposing a visual edit in natural language, an image editing model that applies the edit to the task image, and an optional filter that selects the most faithful variant. The paper attributes VIPE's effectiveness to the realism bias, and the key isolating evidence is a graded realism ladder: as a synthetic scene is progressively upgraded with railway tracks, then rail wagons, then a realistic background, human-rated scene consistency in the model's generated video rises from 0% to 59%. This establishes photorealism, rather than mere 3D structure, as the causal driver of the reasoning improvemen

Load-bearing premise

The load-bearing premise is that the edited visual prompt preserves the task's logic and difficulty, so that accuracy gains reflect better reasoning rather than an easier or altered task; the paper's own verification of this is qualitative, and in the headline comparison the VIPE condition also changed the text prompt, removed the buckets, and used a different evaluator.

What would settle it

Run the physics task with the text prompt, evaluator, and task setup held identical while changing only the image style; if the accuracy gain vanishes or reverses, the improvement is not attributable to the visual prompt. A complementary check is to have human raters judge whether the photorealistic version is objectively easier to solve than the sketch; if it is, the model's competence gain is partly an artifact of easier inputs.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Abstract or synthetic benchmarks understate video models' reasoning abilities; re-rendering the same task photorealistically gives a fairer measure of competence.
  • Visual prompt engineering is a cost-effective test-time scaling strategy: spending a fixed budget on more visual variants often beats generating more videos on a single prompt, and the two can be combined for compounded gains.
  • Automated ideation, whether freeform with a vision-language model or structured step-by-step concept edits, can find effective visual prompts without relying on human intuition.
  • The realism-bias finding generalizes across tasks and video models, suggesting a practical recommendation: when evaluating video reasoning, prefer realistically rendered task versions over sketch-like ones.
  • Native image generation models benefit much less from VIPE on the same task, indicating the effect is tied to a representational gap specific to video models rather than a universal property of all visual models.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the representation-gap explanation is correct, VIPE's gains should shrink as video models are trained on more stylistically diverse data; the paper does not test this, but it is a directly testable trajectory.
  • The same framing suggests an inverse effect for models trained on abstract or synthetic domains, such as simulation-trained agents: those models might reason better when realistic scenes are simplified into sketches, and VIPE could be applied in the reverse direction.
  • The headline physics comparison confounds the image edit with changes to the text prompt, removal of the buckets from view, and a different evaluator; a clean replication holding everything else fixed while varying only the image would strengthen the causal claim.
  • A small internal result, where overlaying a static grid on maze videos improved a vision-language autorater's agreement with humans, hints that VIPE-style edits could also improve video understanding rather than just generation.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper introduces Visual Prompt Engineering (VIPE): before querying a video generation model on a visual reasoning task, an image-editing model transforms the task image (e.g., an abstract sketch into a photorealistic scene), and the edited image is used as the first frame. The authors claim that VIPE consistently improves video reasoning across tasks; that it can be automated by freeform VLM ideation or by step-by-step Atomic Concept Editing (ACE); that it is more cost-effective than test-time scaling via self-consistency; that it can outperform text-based prompt engineering; and that video models exhibit a 'realism bias' whereby abstract tasks underestimate true competence. Experiments cover VPCT, mazes, RushHour, Conjunctive Search, Sort 3 Numbers, and Connect the Dots, using Veo 3.1, Wan2.2, Omni Flash, and two Nano Banana image models.

Significance. If the causal claim is correct, VIPE is an important, low-cost technique and has immediate implications for video-reasoning benchmark design. The paper's strengths include a large body of experiments (18,160 videos), human-validated autoraters, an unnatural-texture ablation, and a cost model. However, the headline VPCT result and the automated-VIPE results currently do not establish the claim because of confounding changes and post-hoc selection. The paper is worth publishing after a revision that isolates the effect of visual editing from evaluator/text-prompt/task changes and reports selection-corrected estimates.

major comments (4)
  1. [§3, Fig. 2, App. C] The headline VPCT comparison changes three variables at once. App. C specifies that baseline samples are scored with an MSE-based container-entry detector on the original sketches, while VIPE videos are scored with a color-based tracker that assigns the nearest container if the ball is on the final frame, and the VIPE images have the buckets removed. The text prompt also changes from '...ends up in one of the three containers at the bottom' to '...it finally drops to the ground ... NO CAMERA MOVEMENT AT ALL'. The only equivalence check is the sentence in §3 that the authors verified the edits did not make the problem easier; no quantitative difficulty measure, geometric-fidelity check, or inter-rater protocol is provided. Because the VIPE prompt no longer requires the ball to enter a container and the final-frame evaluator is more permissive, the observed improvement (e.g., Veo 3.1 from
  2. [§4, Fig. 3, App. E] The automated-VIPE results report the best variant per task/split among n=20 proposals, selected after the variants' downstream accuracies were observed. Solid bars are thus maxima over 20 correlated draws, which overstates expected improvement from a randomly chosen or independently generated variant. The hatched 'oracle per-sample' bars are explicitly unachievable. With only 5–10 samples per split, the variance of the max is substantial. Please report the distribution over variants (mean, median, worst), apply a selection correction or holdout procedure, and state how many of the 20 variants beat baseline. This is load-bearing for the claim that VIPE can be automated.
  3. [§5, Fig. 4 and Fig. 5] The comparison with test-time scaling uses a single, already-selected best visual prompt from §3 against self-consistency on the unmodified sketch. This is not an unbiased comparison of methods: the VIPE curve is the outcome of prior selection, whereas the baseline self-consistency curve is not. The budget analysis in Fig. 5 relies on the assumed cost ratio ($0.40 per VIPE vs $3.20 per video); a sensitivity analysis over this ratio would be needed to support the 'more cost-effective' takeaway. Please compare average or freshly ideated VIPE variants to self-consistency, and report the cost model's sensitivity.
  4. [§6, Fig. 6] The VPCT baseline in Fig. 6 (57%) is inconsistent with the 41.3% reported for Veo 3.1 on the full VPCT in Fig. 2. The subset definition, text prompt, and evaluator used in §6 are not specified, making the text-vs-image comparison and the 'notable exception' (57%→73%) difficult to interpret. Please specify the exact subset/evaluation protocol, or use the full dataset for a consistent comparison.
minor comments (5)
  1. [App. H.1] The numbers in the text disagree with Table 9: the text reports 85.6% agreement, Cohen's κ=0.711, and 93.8% / κ=0.874 for clean-cut cases, while Table 9 lists 87.0% / 0.732 and 94.9% / 0.896. Please reconcile.
  2. [§3] Typo: 'performes' should be 'performs' in the sentence about Wan.
  3. [App. E, Fig. 9] The appendix shows freeform VIPE results with Wan2.2 TI2V, but §4 states 'we consistently perform all of the following experiments with Veo 3.1'. Please clarify whether the main-text statement applies only to Sec. 4 or to the whole paper.
  4. [App. F] For Maze and RushHour, pass rate is used instead of majority-vote accuracy. Please state this explicitly in the main text when comparing to Fig. 5, since the two metrics are not directly comparable.
  5. [Fig. 3 caption] The phrase 'Best variant per task and split' could be misinterpreted as a per-split selection procedure; clarify that it is the overall maximum over n=20 variants after evaluation.

Circularity Check

0 steps flagged

No circularity: the VIPE claim rests on external benchmarks and human-validated autoraters; the Sec. 3 comparison is confounded but not self-referential.

full rationale

The paper's central claim—that VIPE improves video reasoning—is an empirical result, not a derivation: accuracy is measured on VPCT, RushHour, and other tasks with autoraters validated against human ratings (App. H, Cohen's kappa roughly 0.71–0.90), and the freeform ideator is explicitly open-loop. The only selection steps, Eq. (3) (argmax over m proposals) and Fig. 3's 'best variant' bars, are transparently labeled as filtered or oracle choices, not as fitted parameters later renamed as predictions. Self-citations [9], [13], and [37] supply tasks and the ACE procedure, but are not invoked as uniqueness theorems or as definitions of the target result. The genuine weakness is external validity, not circularity: Sec. 3 changes the image, the text prompt, and the evaluator simultaneously, and App. C relaxes the scoring (assigning the nearest container if the ball is on the final frame), with only the unquantified sentence 'verified by the authors' as an equivalence check. That is a confound that may explain part of the gain, but it does not make the reported improvement true by construction. Hence no circular step is present.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 0 invented entities

The paper introduces no new physical entities. Its free parameters are experimental design choices: search budgets, sampling counts, and cost assumptions. The main unstated load is the task-equivalence assumption and the reliance on same-stack models for editing, filtering, and rating.

free parameters (5)
  • number of ideation proposals n = 20
    The automated-VIPE headline is the best of 20 proposals per task; the reported gains depend on this search budget, and a fixed single-proposal pipeline would likely show much smaller gains.
  • editor proposals m = 5
    The filter selects the best of 5 image-editing attempts; the cost model (USD 0.40 per VIPE) and the claimed 1/8-video cost ratio assume this value.
  • videos per sample = 10 (freeform), 3 (ACE)
    Performance per variant is averaged over few videos; small-sample variance is present even with reported confidence intervals.
  • ACE branching factors = (6, 5, 3)
    Tree-search depth and branching are hand-chosen; ACE results depend on this exploration budget.
  • cost ratio assumptions = USD 3.20/video, USD 0.40/VIPE
    The cost-effectiveness conclusion in Sec. 5 rests on API-pricing assumptions and the m=5 filtering budget.
axioms (5)
  • domain assumption VIPE edits preserve the underlying task logic and difficulty
    Sec. 3 states the authors manually verified that VIPE did not make the problem easier; no quantitative equivalence check is provided.
  • domain assumption VLM autoraters produce valid correctness labels
    Autoraters are validated against humans on subsets (85-94% agreement), but not on all tasks and not for the realism dimension specifically.
  • domain assumption Video-model outputs are interpretable as task solution attempts
    The analysis relies on the video-reasoning framing of Wiedemer et al. [9]; generated videos are parsed as solutions by a rater or parser.
  • domain assumption Performance differences reflect reasoning ability rather than rater preference for photorealistic content
    No control eliminates the possibility that VLM autoraters score realistic videos more favorably independent of task correctness.
  • domain assumption The image editor faithfully implements the prescribed edits without changing geometry or spatial layout
    The filter prompt enforces essential structure, but faithfulness is evaluated by the same model family that generates the edits.

pith-pipeline@v1.3.0-alltime-deepseek · 30720 in / 11337 out tokens · 121071 ms · 2026-08-01T02:07:47.292475+00:00 · methodology

0 comments
read the original abstract

In the age of foundation models, a model is only as good as its prompt. For this reason, prompt engineering has become an essential technique for improving language model performance. Since video models are currently becoming foundation models for visual tasks (e.g., visual reasoning), we here ask whether they similarly benefit from visual prompt engineering: automatically modifying the task image to improve model performance. For example, for a visual physics reasoning task ("Where does the ball land, after passing a set of obstacles?"), an abstract sketch-like scene can be turned into a photorealistic version with a simple call to an image editing model. We find that visual prompt engineering, or VIPE for short, improves video reasoning performance across tasks. In fact, for video models, visual prompt engineering can be even more effective than classic text-based prompt engineering or test-time scaling. Ultimately, just as text-based prompt engineering systematically improves language model performance, visual prompt engineering can serve as a simple, compute-efficient approach to elicit better visual reasoning performance from video models. Example videos on our project page at https://visual-prompt-engineering.github.io/.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

61 extracted references · 13 linked inside Pith

  1. [1]

    Demonstrate-search-predict: Composing retrieval and language models for knowledge-intensive NLP.arXiv preprint arXiv:2212.14024, 2022

    Omar Khattab, Keshav Santhanam, Xiang Lisa Li, David Hall, Percy Liang, Christopher Potts, and Matei Zaharia. Demonstrate-search-predict: Composing retrieval and language models for knowledge-intensive NLP.arXiv preprint arXiv:2212.14024, 2022

  2. [2]

    A prompt pattern catalog to enhance prompt engineering with ChatGPT.arXiv preprint arXiv:2302.11382, 2023

    Jules White, Quchen Fu, Sam Hays, Michael Sandborn, Carlos Olea, Henry Gilbert, Ashraf Elnashar, Jesse Spencer-Smith, and Douglas C Schmidt. A prompt pattern catalog to enhance prompt engineering with ChatGPT.arXiv preprint arXiv:2302.11382, 2023

  3. [3]

    Prompt engineering with ChatGPT: a guide for academic writers.Annals of biomedical engineering, 51(12):2629–2633, 2023

    Louie Giray. Prompt engineering with ChatGPT: a guide for academic writers.Annals of biomedical engineering, 51(12):2629–2633, 2023

  4. [4]

    Prompt programming for large language models: Beyond the few-shot paradigm

    Laria Reynolds and Kyle McDonell. Prompt programming for large language models: Beyond the few-shot paradigm. InExtended abstracts of the 2021 CHI conference on human factors in computing systems, pages 1–7, 2021

  5. [5]

    Dspy: compiling declarative language model calls into state-of-the-art pipelines

    Omar Khattab, Arnav Singhvi, Paridhi Maheshwari, Zhiyuan Zhang, Keshav Santhanam, Saiful Haq, Ashutosh Sharma, Thomas Joshi, Hanna Moazam, Heather Miller, et al. Dspy: compiling declarative language model calls into state-of-the-art pipelines. InInternational Conference on Learning Representations, volume 2024, pages 54928–54958, 2024

  6. [6]

    Textgrad: Automatic differentiation via text.arXiv preprint arXiv:2406.07496, 2024

    Mert Yuksekgonul, Federico Bianchi, Joseph Boen, Sheng Liu, Zhi Huang, Carlos Guestrin, and James Zou. Textgrad: Automatic differentiation via text.arXiv preprint arXiv:2406.07496, 2024

  7. [7]

    Prompt engineering in large language models

    Ggaliwango Marvin, Nakayiza Hellen, Daudi Jjingo, and Joyce Nakatumba-Nabende. Prompt engineering in large language models. InInternational conference on data intelligence and cognitive informatics, pages 387–402. Springer, 2023

  8. [8]

    Prompt engineering as an important emerging skill for medical professionals: tutorial.Journal of medical Internet research, 25:e50638, 2023

    Bertalan Meskó. Prompt engineering as an important emerging skill for medical professionals: tutorial.Journal of medical Internet research, 25:e50638, 2023. 12 Visual prompt engineering for video models

  9. [9]

    Video models are zero-shot learners and reasoners

    Thaddäus Wiedemer, Yuxuan Li, Paul Vicol, Shixiang Shane Gu, Nick Matarese, Kevin Swersky, Been Kim, Priyank Jaini, and Robert Geirhos. Video models are zero-shot learners and reasoners. arXiv preprint arXiv:2509.20328, 2025

  10. [10]

    Video as the new language for real-world decision making.arXiv preprint arXiv:2402.17139, 2024

    Sherry Yang, Jacob Walker, Jack Parker-Holder, Yilun Du, Jake Bruce, Andre Barreto, Pieter Abbeel, and Dale Schuurmans. Video as the new language for real-world decision making.arXiv preprint arXiv:2402.17139, 2024

  11. [11]

    Rethinking visual intelligence: Insights from video pretraining.arXiv preprint arXiv:2510.24448, 2025

    Pablo Acuaviva, Aram Davtyan, Mariam Hassan, Sebastian Stapf, Ahmad Rahimi, Alexandre Alahi, and Paolo Favaro. Rethinking visual intelligence: Insights from video pretraining.arXiv preprint arXiv:2510.24448, 2025

  12. [12]

    A very big video reasoning suite.arXiv preprint arXiv:2602.20159, 2026

    Maijunxian Wang, Ruisi Wang, Juyi Lin, Ran Ji, Thaddäus Wiedemer, Qingying Gao, Dezhi Luo, Yaoyao Qian, Lianyu Huang, Zelong Hong, et al. A very big video reasoning suite.arXiv preprint arXiv:2602.20159, 2026

  13. [13]

    MENTISOCULI: Revealing the limits of reasoning with mental imagery.arXiv preprint arXiv:2602.02465, 2026

    Jana Zeller, Thaddäus Wiedemer, Fanfei Li, Thomas Klein, Prasanna Mayilvahanan, Matthias Bethge, Felix Wichmann, Ryan Cotterell, and Wieland Brendel. MENTISOCULI: Revealing the limits of reasoning with mental imagery.arXiv preprint arXiv:2602.02465, 2026

  14. [14]

    Video models reason early: Exploiting plan commitment for maze solving.arXiv preprint arXiv:2603.30043, 2026

    Kaleb Newman, Tyler Zhu, and Olga Russakovsky. Video models reason early: Exploiting plan commitment for maze solving.arXiv preprint arXiv:2603.30043, 2026

  15. [15]

    Are video models ready as zero-shot reasoners? an empirical study with the mme-cof benchmark

    ZiyuGuo, XinyanChen, RenruiZhang, RuichuanAn, YuQi, DongzhiJiang, XiangtaiLi, Manyuan Zhang, Hongsheng Li, and Pheng-Ann Heng. Are video models ready as zero-shot reasoners? an empirical study with the mme-cof benchmark. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9175–9184, 2026

  16. [16]

    Demystifying video reasoning.arXiv preprint arXiv:2603.16870, 2026

    Ruisi Wang, Zhongang Cai, Fanyi Pu, Junxiang Xu, Wanqi Yin, Maijunxian Wang, Ran Ji, Chenyang Gu, Bo Li, Ziqi Huang, et al. Demystifying video reasoning.arXiv preprint arXiv:2603.16870, 2026

  17. [17]

    Thinking in frames: How visual context and test-time scaling empower video reasoning.arXiv preprint arXiv:2601.21037, 2026

    Chengzu Li, Zanyi Wang, Jiaang Li, Yi Xu, Han Zhou, Huanyu Zhang, Ruichuan An, Dengyang Jiang, Zhaochong An, Ivan Vulić, et al. Thinking in frames: How visual context and test-time scaling empower video reasoning.arXiv preprint arXiv:2601.21037, 2026

  18. [18]

    VLMs are good teachers for video reasoning via adaptive test-time optimization.arXiv preprint arXiv:2606.02564, 2026

    Junhao Cheng, Liang Hou, Tianxiong Zhong, Xin Tao, Pengfei Wan, Kun Gai, and Jing Liao. VLMs are good teachers for video reasoning via adaptive test-time optimization.arXiv preprint arXiv:2606.02564, 2026

  19. [19]

    Thinking with video: Video generation as a promising multimodal reasoning paradigm

    Jingqi Tong, Yurong Mou, Hangcheng Li, Mingzhe Li, Yongzhuo Yang, Ming Zhang, Qiguang Chen, Tianyi Liang, Xiaomeng Hu, Yining Zheng, et al. Thinking with video: Video generation as a promising multimodal reasoning paradigm. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 41121–41129, 2026

  20. [20]

    Review of large vision models and visual prompt engineering.Meta-Radiology, 1(3):100047, 2023

    Jiaqi Wang, Zhengliang Liu, Lin Zhao, Zihao Wu, Chong Ma, Sigang Yu, Haixing Dai, Qiushi Yang, Yiheng Liu, Songyao Zhang, et al. Review of large vision models and visual prompt engineering.Meta-Radiology, 1(3):100047, 2023

  21. [21]

    A systematic survey of prompt engineering on vision-language foundation models.arXiv preprint arXiv:2307.12980, 2023

    Jindong Gu, Zhen Han, Shuo Chen, Ahmad Beirami, Bailan He, Gengyuan Zhang, Ruotong Liao, Yao Qin, Volker Tresp, and Philip Torr. A systematic survey of prompt engineering on vision-language foundation models.arXiv preprint arXiv:2307.12980, 2023

  22. [22]

    What does clip know about a red circle? visual prompt engineering for vlms

    Aleksandar Shtedritski, Christian Rupprecht, and Andrea Vedaldi. What does clip know about a red circle? visual prompt engineering for vlms. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 11987–11997, 2023

  23. [23]

    Cpt: Colorful prompt tuning for pre-trained vision-language models.AI Open, 5:30–38, 2024

    Yuan Yao, Ao Zhang, Zhengyan Zhang, Zhiyuan Liu, Tat-Seng Chua, and Maosong Sun. Cpt: Colorful prompt tuning for pre-trained vision-language models.AI Open, 5:30–38, 2024. 13 Visual prompt engineering for video models

  24. [24]

    Exploring visual prompts for adapting large-scale models.arXiv preprint arXiv:2203.17274, 2022

    Hyojin Bahng, Ali Jahber, Prithvijit Chakrabarty, and Phillip Isola. Exploring visual prompts for adapting large-scale models.arXiv preprint arXiv:2203.17274, 2022

  25. [25]

    Highlight: Learning visual prompts for vision-language models, 2024

    Jana Ricarda Zeller, Aleksandar Shtedritski, and Christian Rupprecht. Highlight: Learning visual prompts for vision-language models, 2024

  26. [26]

    Visual prompt tuning

    Menglin Jia, Luming Tang, Bor-Chun Chen, Claire Cardie, Serge Belongie, Bharath Hariharan, and Ser-Nam Lim. Visual prompt tuning. InEuropean Conference on Computer Vision (ECCV), pages 709–727. Springer, 2022

  27. [27]

    Visual promptingviaimageinpainting

    Amir Bar, Yossi Gandelsman, Trevor Darrell, Amir Globerson, and Alexei A Efros. Visual promptingviaimageinpainting. InAdvancesinNeuralInformationProcessingSystems, volume35, pages 25005–25017, 2022

  28. [28]

    Images speak in images: A generalist painter for in-context visual learning

    Xinlong Wang, Wen Wang, Yue Cao, Chunhua Shen, and Tiejun Huang. Images speak in images: A generalist painter for in-context visual learning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6830–6839, 2023

  29. [29]

    Visualcloze: A universal image generation framework via visual in-context learning

    Zhong-Yu Li, Ruoyi Du, Juncheng Yan, Le Zhuo, Zhen Li, Peng Gao, Zhanyu Ma, and Ming-Ming Cheng. Visualcloze: A universal image generation framework via visual in-context learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 18969–18979, 2025

  30. [30]

    Draw-and-understand: Leveraging visual prompts to enable mllms to comprehend what you want

    Weifeng Lin, Xinyu Wei, Ruichuan An, Gao Peng, Bocheng Zou, Yulin Luo, Siyuan Huang, Shanghang Zhang, and Hongsheng Li. Draw-and-understand: Leveraging visual prompts to enable mllms to comprehend what you want. InInternational Conference on Learning Representations, volume 2025, pages 46374–46403, 2025

  31. [31]

    Visual Physics Comprehension Test (VPCT) Dataset.https://huggingface

    camelCase12. Visual Physics Comprehension Test (VPCT) Dataset.https://huggingface. co/datasets/camelCase12/vpct-1, 2025

  32. [32]

    Nano Banana 2: Gemini Image Generation Overview.https://gemini.google/ov erview/image-generation/, 2026

    Google. Nano Banana 2: Gemini Image Generation Overview.https://gemini.google/ov erview/image-generation/, 2026. Accessed: June 17, 2026

  33. [33]

    Gemini 3.1 Pro

    Google. Gemini 3.1 Pro. https://deepmind.google/models/gemini/pro/ , 2026. Accessed: June 17, 2026

  34. [34]

    Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314, 2025

    Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, Jianyuan Zeng, Jiayu Wang, Jingfeng Zhang, Jingren Zhou, Jinkai Wang, Jixuan Chen, Kai Zhu, Kang Zhao, Keyu Yan, Lianghua Huang, Mengyang Feng, Ningyi Zhang, Pandeng Li, Pingyu Wu, Ruihang Chu, Ruili Feng, Shiwei Zhang, Siyang Sun, Tao Fang, T...

  35. [35]

    Google. Veo 3.1. https://deepmind.google/models/veo/ , 2026. Accessed: June 17, 2026

  36. [36]

    Omni Flash Model Card.https://deepmind.google/models/model-cards/g emini-omni-flash/, 2026

    Google. Omni Flash Model Card.https://deepmind.google/models/model-cards/g emini-omni-flash/, 2026. Accessed: July 1, 2026

  37. [37]

    Interpreting and controlling model behavior via constitutions for atomic concept edits

    Neha Kalibhat, Zi Wang, Prasoon Bajpai, Drew Proud, Wenjun Zeng, Been Kim, and Mani Malek. Interpreting and controlling model behavior via constitutions for atomic concept edits. InAnnual Conference on Artificial Intelligence and Statistics (AISTATS), 2026. 14 Visual prompt engineering for video models

  38. [38]

    Self-consistency improves chain of thought reasoning in language models.arXiv preprint arXiv:2203.11171, 2022

    Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models.arXiv preprint arXiv:2203.11171, 2022

  39. [39]

    Gemini developer API pricing.https://ai.google.dev/gemini-api/docs/pr icing#veo-3.1, 2026

    Google. Gemini developer API pricing.https://ai.google.dev/gemini-api/docs/pr icing#veo-3.1, 2026. URL https://ai.google.dev/gemini-api/docs/pricing# veo-3.1. Accessed: 2026-07-01

  40. [40]

    Performance vs

    Chaz Firestone. Performance vs. competence in human–machine comparisons.Proceedings of the National Academy of Sciences, 117(43):26562–26571, 2020

  41. [41]

    How can we know what language models know?Transactions of the Association for Computational Linguistics, 8:423–438, 2020

    Zhengbao Jiang, Frank F Xu, Jun Araki, and Graham Neubig. How can we know what language models know?Transactions of the Association for Computational Linguistics, 8:423–438, 2020

  42. [42]

    Inducing relational knowledge from BERT

    Zied Bouraoui, Jose Camacho-Collados, and Steven Schockaert. Inducing relational knowledge from BERT. InProceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 7456–7463, 2020

  43. [43]

    Shortcut learning in deep neural networks.Nature Machine Intelligence, 2(11):665–673, 2020

    Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann. Shortcut learning in deep neural networks.Nature Machine Intelligence, 2(11):665–673, 2020

  44. [44]

    Unmasking Clever Hans predictors and assessing what machines really learn.Nature communications, 10(1):1096, 2019

    Sebastian Lapuschkin, Stephan Wäldchen, Alexander Binder, Grégoire Montavon, Wojciech Samek, and Klaus-Robert Müller. Unmasking Clever Hans predictors and assessing what machines really learn.Nature communications, 10(1):1096, 2019

  45. [45]

    Unbiased look at dataset bias

    Antonio Torralba and Alexei A Efros. Unbiased look at dataset bias. InCVPR 2011, pages 1521–1528. IEEE, 2011

  46. [46]

    Cosmos 3: Omnimodal world models for physical AI.arXiv preprint arXiv:2606.02800, 2026

    Niket Agarwal, Arslan Ali, Jon Allen, Martin Antolini, Adeline Aubame, Alisson Azzolini, Junjie Bai, Maciej Bala, Yogesh Balaji, Josh Bapst, et al. Cosmos 3: Omnimodal world models for physical AI.arXiv preprint arXiv:2606.02800, 2026

  47. [47]

    this vertical drop leads directly into the first bucket on the left

    Google. Gemini Omni. https://gemini.google/overview/video-generation/, 2026. Accessed: June 17, 2026. 15 Visual prompt engineering for video models Appendix A. Datasets VPCTMIT license, dataset on HuggingFace by camelCase12 [31], 100 samples. For the results in Sec. 4, we use the first ten samples. For Sec. 3, the entire dataset is used. Conjunctive Searc...

  48. [48]

    Each vehicle can only move forward or backward with straight sliding motion along its own axis

  49. [49]

    No rotation is allowed at any time

  50. [50]

    A vehicle continues to move in the chosen direction until it touches another vehicle or a boundary

  51. [51]

    Only one vehicle moves per action

  52. [52]

    The goal is for the red car to reach the exit located on the edge of the grid

  53. [53]

    Vehicle shapes, colors, exit, and outlines must not change throughout the solution

  54. [54]

    No camera motion: no zoom, no pan, no rotate, no tilt, no dolly

  55. [55]

    Do not add or remove anything: no new objects, labels, lights, shadows, reflections, textures, markings, or UI elements

  56. [56]

    Static shot, no zoom or pan

    The background, grid, exit, and all pieces remain perfectly static, except for the piece currently sliding. Task: Plan the minimal sequence of moves needed to free the red car and allow it to exit the parking lot. Output: A video demonstrating the full solution to the puzzle, one move at a time. Example proposals from freeform prompt engineering InVPCT,Co...

  57. [57]

    “success”: A boolean indicating if the runner successfully reached the goal without any rule violations

  58. [59]

    justification

    “justification”: A text explanation of your analysis. Explain why the video was marked valid/invalid or success/failure. Only output the JSON object, nothing else. Do not wrap it in markdown block. VLM-based autorater with grid overlayWe experiment with using overlaying the entire video with a static16× 9grid with black-and-white2px grid lines, see Fig. 2...

  59. [60]

    moves”: A list of strings representing the extracted moves in order, e.g., [“A N

    “moves”: A list of strings representing the extracted moves in order, e.g., [“A N”, “B NW”]. If no moves occur before the scene becomes invalid, this should be []

  60. [61]

    invalid_after_seconds

    “invalid_after_seconds”: A float representing the timestamp (in seconds) when the scene first became invalid. If the video remains fully valid and does not violate any rules until the end, set this to null

  61. [62]

    justification

    “justification”: A text explanation of your analysis. Explain why the video was marked valid/invalid (e.g., if an object morphed, specify which object and at what time), and describe the moves you observed. Only output the JSON object, nothing else. Do not wrap it in markdown block. 41 Visual prompt engineering for video models Table 12|Autorater-human ag...