REVIEW 4 major objections 5 minor 61 references
This paper claims that visually editing a video model's input image, converting abstract sketches into photorealistic scenes, systematically improves its reasoning performance, often more than text prompt engineering or test-time scaling.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 02:07 UTC pith:IBFVXUTG
load-bearing objection The VIPE concept is worth taking seriously, but the headline VPCT comparison is confounded and needs a clean replication before the central claim is trustworthy. the 4 major comments →
Visual prompt engineering for video models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central discovery is that video models have a systematic preference for photorealistic visual input: given the same underlying task, replacing an abstract sketch with a realistic rendering of the same scene improves the model's ability to solve it. In the paper's flagship physics experiment, transforming a ball-and-ramp sketch into a photograph-like scene moves accuracy from near chance to well above chance across several video models, and a step-by-step realism ladder shows that each step toward realism increases generation consistency. The authors argue that the root cause is a representation gap: video models are trained predominantly on realistic footage, so abstract inputs force rea
What carries the argument
The central mechanism is visual prompt engineering (VIPE), a pipeline that starts with an ideator proposing a visual edit in natural language, an image editing model that applies the edit to the task image, and an optional filter that selects the most faithful variant. The paper attributes VIPE's effectiveness to the realism bias, and the key isolating evidence is a graded realism ladder: as a synthetic scene is progressively upgraded with railway tracks, then rail wagons, then a realistic background, human-rated scene consistency in the model's generated video rises from 0% to 59%. This establishes photorealism, rather than mere 3D structure, as the causal driver of the reasoning improvemen
Load-bearing premise
The load-bearing premise is that the edited visual prompt preserves the task's logic and difficulty, so that accuracy gains reflect better reasoning rather than an easier or altered task; the paper's own verification of this is qualitative, and in the headline comparison the VIPE condition also changed the text prompt, removed the buckets, and used a different evaluator.
What would settle it
Run the physics task with the text prompt, evaluator, and task setup held identical while changing only the image style; if the accuracy gain vanishes or reverses, the improvement is not attributable to the visual prompt. A complementary check is to have human raters judge whether the photorealistic version is objectively easier to solve than the sketch; if it is, the model's competence gain is partly an artifact of easier inputs.
If this is right
- Abstract or synthetic benchmarks understate video models' reasoning abilities; re-rendering the same task photorealistically gives a fairer measure of competence.
- Visual prompt engineering is a cost-effective test-time scaling strategy: spending a fixed budget on more visual variants often beats generating more videos on a single prompt, and the two can be combined for compounded gains.
- Automated ideation, whether freeform with a vision-language model or structured step-by-step concept edits, can find effective visual prompts without relying on human intuition.
- The realism-bias finding generalizes across tasks and video models, suggesting a practical recommendation: when evaluating video reasoning, prefer realistically rendered task versions over sketch-like ones.
- Native image generation models benefit much less from VIPE on the same task, indicating the effect is tied to a representational gap specific to video models rather than a universal property of all visual models.
Where Pith is reading between the lines
- If the representation-gap explanation is correct, VIPE's gains should shrink as video models are trained on more stylistically diverse data; the paper does not test this, but it is a directly testable trajectory.
- The same framing suggests an inverse effect for models trained on abstract or synthetic domains, such as simulation-trained agents: those models might reason better when realistic scenes are simplified into sketches, and VIPE could be applied in the reverse direction.
- The headline physics comparison confounds the image edit with changes to the text prompt, removal of the buckets from view, and a different evaluator; a clean replication holding everything else fixed while varying only the image would strengthen the causal claim.
- A small internal result, where overlaying a static grid on maze videos improved a vision-language autorater's agreement with humans, hints that VIPE-style edits could also improve video understanding rather than just generation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Visual Prompt Engineering (VIPE): before querying a video generation model on a visual reasoning task, an image-editing model transforms the task image (e.g., an abstract sketch into a photorealistic scene), and the edited image is used as the first frame. The authors claim that VIPE consistently improves video reasoning across tasks; that it can be automated by freeform VLM ideation or by step-by-step Atomic Concept Editing (ACE); that it is more cost-effective than test-time scaling via self-consistency; that it can outperform text-based prompt engineering; and that video models exhibit a 'realism bias' whereby abstract tasks underestimate true competence. Experiments cover VPCT, mazes, RushHour, Conjunctive Search, Sort 3 Numbers, and Connect the Dots, using Veo 3.1, Wan2.2, Omni Flash, and two Nano Banana image models.
Significance. If the causal claim is correct, VIPE is an important, low-cost technique and has immediate implications for video-reasoning benchmark design. The paper's strengths include a large body of experiments (18,160 videos), human-validated autoraters, an unnatural-texture ablation, and a cost model. However, the headline VPCT result and the automated-VIPE results currently do not establish the claim because of confounding changes and post-hoc selection. The paper is worth publishing after a revision that isolates the effect of visual editing from evaluator/text-prompt/task changes and reports selection-corrected estimates.
major comments (4)
- [§3, Fig. 2, App. C] The headline VPCT comparison changes three variables at once. App. C specifies that baseline samples are scored with an MSE-based container-entry detector on the original sketches, while VIPE videos are scored with a color-based tracker that assigns the nearest container if the ball is on the final frame, and the VIPE images have the buckets removed. The text prompt also changes from '...ends up in one of the three containers at the bottom' to '...it finally drops to the ground ... NO CAMERA MOVEMENT AT ALL'. The only equivalence check is the sentence in §3 that the authors verified the edits did not make the problem easier; no quantitative difficulty measure, geometric-fidelity check, or inter-rater protocol is provided. Because the VIPE prompt no longer requires the ball to enter a container and the final-frame evaluator is more permissive, the observed improvement (e.g., Veo 3.1 from
- [§4, Fig. 3, App. E] The automated-VIPE results report the best variant per task/split among n=20 proposals, selected after the variants' downstream accuracies were observed. Solid bars are thus maxima over 20 correlated draws, which overstates expected improvement from a randomly chosen or independently generated variant. The hatched 'oracle per-sample' bars are explicitly unachievable. With only 5–10 samples per split, the variance of the max is substantial. Please report the distribution over variants (mean, median, worst), apply a selection correction or holdout procedure, and state how many of the 20 variants beat baseline. This is load-bearing for the claim that VIPE can be automated.
- [§5, Fig. 4 and Fig. 5] The comparison with test-time scaling uses a single, already-selected best visual prompt from §3 against self-consistency on the unmodified sketch. This is not an unbiased comparison of methods: the VIPE curve is the outcome of prior selection, whereas the baseline self-consistency curve is not. The budget analysis in Fig. 5 relies on the assumed cost ratio ($0.40 per VIPE vs $3.20 per video); a sensitivity analysis over this ratio would be needed to support the 'more cost-effective' takeaway. Please compare average or freshly ideated VIPE variants to self-consistency, and report the cost model's sensitivity.
- [§6, Fig. 6] The VPCT baseline in Fig. 6 (57%) is inconsistent with the 41.3% reported for Veo 3.1 on the full VPCT in Fig. 2. The subset definition, text prompt, and evaluator used in §6 are not specified, making the text-vs-image comparison and the 'notable exception' (57%→73%) difficult to interpret. Please specify the exact subset/evaluation protocol, or use the full dataset for a consistent comparison.
minor comments (5)
- [App. H.1] The numbers in the text disagree with Table 9: the text reports 85.6% agreement, Cohen's κ=0.711, and 93.8% / κ=0.874 for clean-cut cases, while Table 9 lists 87.0% / 0.732 and 94.9% / 0.896. Please reconcile.
- [§3] Typo: 'performes' should be 'performs' in the sentence about Wan.
- [App. E, Fig. 9] The appendix shows freeform VIPE results with Wan2.2 TI2V, but §4 states 'we consistently perform all of the following experiments with Veo 3.1'. Please clarify whether the main-text statement applies only to Sec. 4 or to the whole paper.
- [App. F] For Maze and RushHour, pass rate is used instead of majority-vote accuracy. Please state this explicitly in the main text when comparing to Fig. 5, since the two metrics are not directly comparable.
- [Fig. 3 caption] The phrase 'Best variant per task and split' could be misinterpreted as a per-split selection procedure; clarify that it is the overall maximum over n=20 variants after evaluation.
Circularity Check
No circularity: the VIPE claim rests on external benchmarks and human-validated autoraters; the Sec. 3 comparison is confounded but not self-referential.
full rationale
The paper's central claim—that VIPE improves video reasoning—is an empirical result, not a derivation: accuracy is measured on VPCT, RushHour, and other tasks with autoraters validated against human ratings (App. H, Cohen's kappa roughly 0.71–0.90), and the freeform ideator is explicitly open-loop. The only selection steps, Eq. (3) (argmax over m proposals) and Fig. 3's 'best variant' bars, are transparently labeled as filtered or oracle choices, not as fitted parameters later renamed as predictions. Self-citations [9], [13], and [37] supply tasks and the ACE procedure, but are not invoked as uniqueness theorems or as definitions of the target result. The genuine weakness is external validity, not circularity: Sec. 3 changes the image, the text prompt, and the evaluator simultaneously, and App. C relaxes the scoring (assigning the nearest container if the ball is on the final frame), with only the unquantified sentence 'verified by the authors' as an equivalence check. That is a confound that may explain part of the gain, but it does not make the reported improvement true by construction. Hence no circular step is present.
Axiom & Free-Parameter Ledger
free parameters (5)
- number of ideation proposals n =
20
- editor proposals m =
5
- videos per sample =
10 (freeform), 3 (ACE)
- ACE branching factors =
(6, 5, 3)
- cost ratio assumptions =
USD 3.20/video, USD 0.40/VIPE
axioms (5)
- domain assumption VIPE edits preserve the underlying task logic and difficulty
- domain assumption VLM autoraters produce valid correctness labels
- domain assumption Video-model outputs are interpretable as task solution attempts
- domain assumption Performance differences reflect reasoning ability rather than rater preference for photorealistic content
- domain assumption The image editor faithfully implements the prescribed edits without changing geometry or spatial layout
read the original abstract
In the age of foundation models, a model is only as good as its prompt. For this reason, prompt engineering has become an essential technique for improving language model performance. Since video models are currently becoming foundation models for visual tasks (e.g., visual reasoning), we here ask whether they similarly benefit from visual prompt engineering: automatically modifying the task image to improve model performance. For example, for a visual physics reasoning task ("Where does the ball land, after passing a set of obstacles?"), an abstract sketch-like scene can be turned into a photorealistic version with a simple call to an image editing model. We find that visual prompt engineering, or VIPE for short, improves video reasoning performance across tasks. In fact, for video models, visual prompt engineering can be even more effective than classic text-based prompt engineering or test-time scaling. Ultimately, just as text-based prompt engineering systematically improves language model performance, visual prompt engineering can serve as a simple, compute-efficient approach to elicit better visual reasoning performance from video models. Example videos on our project page at https://visual-prompt-engineering.github.io/.
Reference graph
Works this paper leans on
-
[1]
Omar Khattab, Keshav Santhanam, Xiang Lisa Li, David Hall, Percy Liang, Christopher Potts, and Matei Zaharia. Demonstrate-search-predict: Composing retrieval and language models for knowledge-intensive NLP.arXiv preprint arXiv:2212.14024, 2022
Pith/arXiv arXiv 2022
-
[2]
Jules White, Quchen Fu, Sam Hays, Michael Sandborn, Carlos Olea, Henry Gilbert, Ashraf Elnashar, Jesse Spencer-Smith, and Douglas C Schmidt. A prompt pattern catalog to enhance prompt engineering with ChatGPT.arXiv preprint arXiv:2302.11382, 2023
Pith/arXiv arXiv 2023
-
[3]
Prompt engineering with ChatGPT: a guide for academic writers.Annals of biomedical engineering, 51(12):2629–2633, 2023
Louie Giray. Prompt engineering with ChatGPT: a guide for academic writers.Annals of biomedical engineering, 51(12):2629–2633, 2023
2023
-
[4]
Prompt programming for large language models: Beyond the few-shot paradigm
Laria Reynolds and Kyle McDonell. Prompt programming for large language models: Beyond the few-shot paradigm. InExtended abstracts of the 2021 CHI conference on human factors in computing systems, pages 1–7, 2021
2021
-
[5]
Dspy: compiling declarative language model calls into state-of-the-art pipelines
Omar Khattab, Arnav Singhvi, Paridhi Maheshwari, Zhiyuan Zhang, Keshav Santhanam, Saiful Haq, Ashutosh Sharma, Thomas Joshi, Hanna Moazam, Heather Miller, et al. Dspy: compiling declarative language model calls into state-of-the-art pipelines. InInternational Conference on Learning Representations, volume 2024, pages 54928–54958, 2024
2024
-
[6]
Textgrad: Automatic differentiation via text.arXiv preprint arXiv:2406.07496, 2024
Mert Yuksekgonul, Federico Bianchi, Joseph Boen, Sheng Liu, Zhi Huang, Carlos Guestrin, and James Zou. Textgrad: Automatic differentiation via text.arXiv preprint arXiv:2406.07496, 2024
Pith/arXiv arXiv 2024
-
[7]
Prompt engineering in large language models
Ggaliwango Marvin, Nakayiza Hellen, Daudi Jjingo, and Joyce Nakatumba-Nabende. Prompt engineering in large language models. InInternational conference on data intelligence and cognitive informatics, pages 387–402. Springer, 2023
2023
-
[8]
Prompt engineering as an important emerging skill for medical professionals: tutorial.Journal of medical Internet research, 25:e50638, 2023
Bertalan Meskó. Prompt engineering as an important emerging skill for medical professionals: tutorial.Journal of medical Internet research, 25:e50638, 2023. 12 Visual prompt engineering for video models
2023
-
[9]
Video models are zero-shot learners and reasoners
Thaddäus Wiedemer, Yuxuan Li, Paul Vicol, Shixiang Shane Gu, Nick Matarese, Kevin Swersky, Been Kim, Priyank Jaini, and Robert Geirhos. Video models are zero-shot learners and reasoners. arXiv preprint arXiv:2509.20328, 2025
Pith/arXiv arXiv 2025
-
[10]
Video as the new language for real-world decision making.arXiv preprint arXiv:2402.17139, 2024
Sherry Yang, Jacob Walker, Jack Parker-Holder, Yilun Du, Jake Bruce, Andre Barreto, Pieter Abbeel, and Dale Schuurmans. Video as the new language for real-world decision making.arXiv preprint arXiv:2402.17139, 2024
Pith/arXiv arXiv 2024
-
[11]
Pablo Acuaviva, Aram Davtyan, Mariam Hassan, Sebastian Stapf, Ahmad Rahimi, Alexandre Alahi, and Paolo Favaro. Rethinking visual intelligence: Insights from video pretraining.arXiv preprint arXiv:2510.24448, 2025
arXiv 2025
-
[12]
A very big video reasoning suite.arXiv preprint arXiv:2602.20159, 2026
Maijunxian Wang, Ruisi Wang, Juyi Lin, Ran Ji, Thaddäus Wiedemer, Qingying Gao, Dezhi Luo, Yaoyao Qian, Lianyu Huang, Zelong Hong, et al. A very big video reasoning suite.arXiv preprint arXiv:2602.20159, 2026
arXiv 2026
-
[13]
Jana Zeller, Thaddäus Wiedemer, Fanfei Li, Thomas Klein, Prasanna Mayilvahanan, Matthias Bethge, Felix Wichmann, Ryan Cotterell, and Wieland Brendel. MENTISOCULI: Revealing the limits of reasoning with mental imagery.arXiv preprint arXiv:2602.02465, 2026
Pith/arXiv arXiv 2026
-
[14]
Kaleb Newman, Tyler Zhu, and Olga Russakovsky. Video models reason early: Exploiting plan commitment for maze solving.arXiv preprint arXiv:2603.30043, 2026
arXiv 2026
-
[15]
Are video models ready as zero-shot reasoners? an empirical study with the mme-cof benchmark
ZiyuGuo, XinyanChen, RenruiZhang, RuichuanAn, YuQi, DongzhiJiang, XiangtaiLi, Manyuan Zhang, Hongsheng Li, and Pheng-Ann Heng. Are video models ready as zero-shot reasoners? an empirical study with the mme-cof benchmark. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9175–9184, 2026
2026
-
[16]
Demystifying video reasoning.arXiv preprint arXiv:2603.16870, 2026
Ruisi Wang, Zhongang Cai, Fanyi Pu, Junxiang Xu, Wanqi Yin, Maijunxian Wang, Ran Ji, Chenyang Gu, Bo Li, Ziqi Huang, et al. Demystifying video reasoning.arXiv preprint arXiv:2603.16870, 2026
Pith/arXiv arXiv 2026
-
[17]
Chengzu Li, Zanyi Wang, Jiaang Li, Yi Xu, Han Zhou, Huanyu Zhang, Ruichuan An, Dengyang Jiang, Zhaochong An, Ivan Vulić, et al. Thinking in frames: How visual context and test-time scaling empower video reasoning.arXiv preprint arXiv:2601.21037, 2026
arXiv 2026
-
[18]
Junhao Cheng, Liang Hou, Tianxiong Zhong, Xin Tao, Pengfei Wan, Kun Gai, and Jing Liao. VLMs are good teachers for video reasoning via adaptive test-time optimization.arXiv preprint arXiv:2606.02564, 2026
Pith/arXiv arXiv 2026
-
[19]
Thinking with video: Video generation as a promising multimodal reasoning paradigm
Jingqi Tong, Yurong Mou, Hangcheng Li, Mingzhe Li, Yongzhuo Yang, Ming Zhang, Qiguang Chen, Tianyi Liang, Xiaomeng Hu, Yining Zheng, et al. Thinking with video: Video generation as a promising multimodal reasoning paradigm. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 41121–41129, 2026
2026
-
[20]
Review of large vision models and visual prompt engineering.Meta-Radiology, 1(3):100047, 2023
Jiaqi Wang, Zhengliang Liu, Lin Zhao, Zihao Wu, Chong Ma, Sigang Yu, Haixing Dai, Qiushi Yang, Yiheng Liu, Songyao Zhang, et al. Review of large vision models and visual prompt engineering.Meta-Radiology, 1(3):100047, 2023
2023
-
[21]
Jindong Gu, Zhen Han, Shuo Chen, Ahmad Beirami, Bailan He, Gengyuan Zhang, Ruotong Liao, Yao Qin, Volker Tresp, and Philip Torr. A systematic survey of prompt engineering on vision-language foundation models.arXiv preprint arXiv:2307.12980, 2023
Pith/arXiv arXiv 2023
-
[22]
What does clip know about a red circle? visual prompt engineering for vlms
Aleksandar Shtedritski, Christian Rupprecht, and Andrea Vedaldi. What does clip know about a red circle? visual prompt engineering for vlms. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 11987–11997, 2023
2023
-
[23]
Cpt: Colorful prompt tuning for pre-trained vision-language models.AI Open, 5:30–38, 2024
Yuan Yao, Ao Zhang, Zhengyan Zhang, Zhiyuan Liu, Tat-Seng Chua, and Maosong Sun. Cpt: Colorful prompt tuning for pre-trained vision-language models.AI Open, 5:30–38, 2024. 13 Visual prompt engineering for video models
2024
-
[24]
Exploring visual prompts for adapting large-scale models.arXiv preprint arXiv:2203.17274, 2022
Hyojin Bahng, Ali Jahber, Prithvijit Chakrabarty, and Phillip Isola. Exploring visual prompts for adapting large-scale models.arXiv preprint arXiv:2203.17274, 2022
Pith/arXiv arXiv 2022
-
[25]
Highlight: Learning visual prompts for vision-language models, 2024
Jana Ricarda Zeller, Aleksandar Shtedritski, and Christian Rupprecht. Highlight: Learning visual prompts for vision-language models, 2024
2024
-
[26]
Visual prompt tuning
Menglin Jia, Luming Tang, Bor-Chun Chen, Claire Cardie, Serge Belongie, Bharath Hariharan, and Ser-Nam Lim. Visual prompt tuning. InEuropean Conference on Computer Vision (ECCV), pages 709–727. Springer, 2022
2022
-
[27]
Visual promptingviaimageinpainting
Amir Bar, Yossi Gandelsman, Trevor Darrell, Amir Globerson, and Alexei A Efros. Visual promptingviaimageinpainting. InAdvancesinNeuralInformationProcessingSystems, volume35, pages 25005–25017, 2022
2022
-
[28]
Images speak in images: A generalist painter for in-context visual learning
Xinlong Wang, Wen Wang, Yue Cao, Chunhua Shen, and Tiejun Huang. Images speak in images: A generalist painter for in-context visual learning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6830–6839, 2023
2023
-
[29]
Visualcloze: A universal image generation framework via visual in-context learning
Zhong-Yu Li, Ruoyi Du, Juncheng Yan, Le Zhuo, Zhen Li, Peng Gao, Zhanyu Ma, and Ming-Ming Cheng. Visualcloze: A universal image generation framework via visual in-context learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 18969–18979, 2025
2025
-
[30]
Draw-and-understand: Leveraging visual prompts to enable mllms to comprehend what you want
Weifeng Lin, Xinyu Wei, Ruichuan An, Gao Peng, Bocheng Zou, Yulin Luo, Siyuan Huang, Shanghang Zhang, and Hongsheng Li. Draw-and-understand: Leveraging visual prompts to enable mllms to comprehend what you want. InInternational Conference on Learning Representations, volume 2025, pages 46374–46403, 2025
2025
-
[31]
Visual Physics Comprehension Test (VPCT) Dataset.https://huggingface
camelCase12. Visual Physics Comprehension Test (VPCT) Dataset.https://huggingface. co/datasets/camelCase12/vpct-1, 2025
2025
-
[32]
Nano Banana 2: Gemini Image Generation Overview.https://gemini.google/ov erview/image-generation/, 2026
Google. Nano Banana 2: Gemini Image Generation Overview.https://gemini.google/ov erview/image-generation/, 2026. Accessed: June 17, 2026
2026
-
[33]
Gemini 3.1 Pro
Google. Gemini 3.1 Pro. https://deepmind.google/models/gemini/pro/ , 2026. Accessed: June 17, 2026
2026
-
[34]
Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314, 2025
Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, Jianyuan Zeng, Jiayu Wang, Jingfeng Zhang, Jingren Zhou, Jinkai Wang, Jixuan Chen, Kai Zhu, Kang Zhao, Keyu Yan, Lianghua Huang, Mengyang Feng, Ningyi Zhang, Pandeng Li, Pingyu Wu, Ruihang Chu, Ruili Feng, Shiwei Zhang, Siyang Sun, Tao Fang, T...
Pith/arXiv arXiv 2025
-
[35]
Google. Veo 3.1. https://deepmind.google/models/veo/ , 2026. Accessed: June 17, 2026
2026
-
[36]
Omni Flash Model Card.https://deepmind.google/models/model-cards/g emini-omni-flash/, 2026
Google. Omni Flash Model Card.https://deepmind.google/models/model-cards/g emini-omni-flash/, 2026. Accessed: July 1, 2026
2026
-
[37]
Interpreting and controlling model behavior via constitutions for atomic concept edits
Neha Kalibhat, Zi Wang, Prasoon Bajpai, Drew Proud, Wenjun Zeng, Been Kim, and Mani Malek. Interpreting and controlling model behavior via constitutions for atomic concept edits. InAnnual Conference on Artificial Intelligence and Statistics (AISTATS), 2026. 14 Visual prompt engineering for video models
2026
-
[38]
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models.arXiv preprint arXiv:2203.11171, 2022
Pith/arXiv arXiv 2022
-
[39]
Gemini developer API pricing.https://ai.google.dev/gemini-api/docs/pr icing#veo-3.1, 2026
Google. Gemini developer API pricing.https://ai.google.dev/gemini-api/docs/pr icing#veo-3.1, 2026. URL https://ai.google.dev/gemini-api/docs/pricing# veo-3.1. Accessed: 2026-07-01
2026
-
[40]
Performance vs
Chaz Firestone. Performance vs. competence in human–machine comparisons.Proceedings of the National Academy of Sciences, 117(43):26562–26571, 2020
2020
-
[41]
How can we know what language models know?Transactions of the Association for Computational Linguistics, 8:423–438, 2020
Zhengbao Jiang, Frank F Xu, Jun Araki, and Graham Neubig. How can we know what language models know?Transactions of the Association for Computational Linguistics, 8:423–438, 2020
2020
-
[42]
Inducing relational knowledge from BERT
Zied Bouraoui, Jose Camacho-Collados, and Steven Schockaert. Inducing relational knowledge from BERT. InProceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 7456–7463, 2020
2020
-
[43]
Shortcut learning in deep neural networks.Nature Machine Intelligence, 2(11):665–673, 2020
Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann. Shortcut learning in deep neural networks.Nature Machine Intelligence, 2(11):665–673, 2020
2020
-
[44]
Unmasking Clever Hans predictors and assessing what machines really learn.Nature communications, 10(1):1096, 2019
Sebastian Lapuschkin, Stephan Wäldchen, Alexander Binder, Grégoire Montavon, Wojciech Samek, and Klaus-Robert Müller. Unmasking Clever Hans predictors and assessing what machines really learn.Nature communications, 10(1):1096, 2019
2019
-
[45]
Unbiased look at dataset bias
Antonio Torralba and Alexei A Efros. Unbiased look at dataset bias. InCVPR 2011, pages 1521–1528. IEEE, 2011
2011
-
[46]
Cosmos 3: Omnimodal world models for physical AI.arXiv preprint arXiv:2606.02800, 2026
Niket Agarwal, Arslan Ali, Jon Allen, Martin Antolini, Adeline Aubame, Alisson Azzolini, Junjie Bai, Maciej Bala, Yogesh Balaji, Josh Bapst, et al. Cosmos 3: Omnimodal world models for physical AI.arXiv preprint arXiv:2606.02800, 2026
Pith/arXiv arXiv 2026
-
[47]
this vertical drop leads directly into the first bucket on the left
Google. Gemini Omni. https://gemini.google/overview/video-generation/, 2026. Accessed: June 17, 2026. 15 Visual prompt engineering for video models Appendix A. Datasets VPCTMIT license, dataset on HuggingFace by camelCase12 [31], 100 samples. For the results in Sec. 4, we use the first ten samples. For Sec. 3, the entire dataset is used. Conjunctive Searc...
2026
-
[48]
Each vehicle can only move forward or backward with straight sliding motion along its own axis
-
[49]
No rotation is allowed at any time
-
[50]
A vehicle continues to move in the chosen direction until it touches another vehicle or a boundary
-
[51]
Only one vehicle moves per action
-
[52]
The goal is for the red car to reach the exit located on the edge of the grid
-
[53]
Vehicle shapes, colors, exit, and outlines must not change throughout the solution
-
[54]
No camera motion: no zoom, no pan, no rotate, no tilt, no dolly
-
[55]
Do not add or remove anything: no new objects, labels, lights, shadows, reflections, textures, markings, or UI elements
-
[56]
Static shot, no zoom or pan
The background, grid, exit, and all pieces remain perfectly static, except for the piece currently sliding. Task: Plan the minimal sequence of moves needed to free the red car and allow it to exit the parking lot. Output: A video demonstrating the full solution to the puzzle, one move at a time. Example proposals from freeform prompt engineering InVPCT,Co...
-
[57]
“success”: A boolean indicating if the runner successfully reached the goal without any rule violations
-
[59]
justification
“justification”: A text explanation of your analysis. Explain why the video was marked valid/invalid or success/failure. Only output the JSON object, nothing else. Do not wrap it in markdown block. VLM-based autorater with grid overlayWe experiment with using overlaying the entire video with a static16× 9grid with black-and-white2px grid lines, see Fig. 2...
-
[60]
moves”: A list of strings representing the extracted moves in order, e.g., [“A N
“moves”: A list of strings representing the extracted moves in order, e.g., [“A N”, “B NW”]. If no moves occur before the scene becomes invalid, this should be []
-
[61]
invalid_after_seconds
“invalid_after_seconds”: A float representing the timestamp (in seconds) when the scene first became invalid. If the video remains fully valid and does not violate any rules until the end, set this to null
-
[62]
justification
“justification”: A text explanation of your analysis. Explain why the video was marked valid/invalid (e.g., if an object morphed, specify which object and at what time), and describe the moves you observed. Only output the JSON object, nothing else. Do not wrap it in markdown block. 41 Visual prompt engineering for video models Table 12|Autorater-human ag...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.