REVIEW 5 major objections 7 minor 41 references
A reference video can specify physical behavior that text prompts cannot, and VIPER transfers that behavior to new scenes.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-30 21:24 UTC pith:OUD45OYC
load-bearing objection Solid systems paper on reference-as-physics for I2V; real packaging contribution, but the headline VLM-as-Judge gap is partly compromised by Qwen reuse across data, conditioning, and scoring. the 5 major comments →
VIPER: Visual In-Context Physics Reasoning for Physically Plausible Video Generation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper claims that physically plausible image-to-video generation can be cast as visual in-context physics reasoning: given a target image, a brief prompt, and a reference video, an MLLM can extract transferable physical cues from the reference and, via hierarchical training into a DiT image-to-video backbone, produce a target video that inherits the reference’s material response, trajectory, and impact behavior while preserving the base generator’s visual prior—without exhaustive physics prompts.
What carries the argument
VIPER’s reference-video physical conditioning path: learnable physics query tokens read the reference through an MLLM, a connector projects those hidden states into DiT condition tokens, and a three-stage hierarchical training schedule first aligns self-reference conditioning, then trains cross-video physical transfer on VIPER-19K pairs, then lightly adapts the generator with LoRA so physics tokens matter without overwriting the pretrained visual prior.
Load-bearing premise
The method assumes that compact MLLM query tokens plus pair training isolate transferable physical dynamics rather than appearance shortcuts, semantic overlap, or artifacts of how the same model family helps build pairs, write baseline prompts, and judge physical similarity.
What would settle it
On truly held-out reference–target pairs with matched high-level labels but deliberately inverted dynamics (for example opposite spin direction, shatter versus bounce, or mismatched impact scale), measure whether VIPER’s physical-similarity and human preference advantages disappear while baselines catch up when given equally detailed text physics descriptions.
If this is right
- Reference video becomes a practical control interface for material response, contact, and trajectory without hand-written physics prompts.
- Pretrained image-to-video models can be steered for physical behavior transfer while keeping their existing visual quality prior via staged alignment and light LoRA adaptation.
- Datasets organized by material, trajectory, and physical-impact labels with filtered cross-appearance pairs become the natural supervision unit for this task.
- Qualitative transfer of melting, deformation, and related processes to new target scenes is achievable from brief target prompts plus a demonstration clip.
Where Pith is reading between the lines
- If reference-as-physics-spec works at scale, interactive tools could let users pick a short demo clip instead of engineering force or material language for each shot.
- Coarse taxonomy gaps called out in the limitations (for example spin without direction) suggest the next bottleneck is finer causal labels and compatibility checks, not only bigger generators.
- Heavy reuse of one multimodal model family across pair filtering, baseline textification, and judging implies future benchmarks may need an independent physics scorer to separate method gains from judge alignment.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes VIPER, a reference-guided image-to-video framework in which a reference video serves as a dense visual demonstration of desired physical behavior (material response, trajectory, contact, impact). An MLLM (Qwen3-VL-4B) with learnable physics query tokens encodes the reference into compact condition tokens injected into a Wan2.2-I2V-14B DiT generator; a three-stage hierarchical training scheme (self-alignment → cross-video transfer → LoRA adaptation) is used to avoid pixel-copying and preserve the base prior. The authors also construct VIPER-19K, a dataset with material/trajectory/impact annotations and reference–target pairs filtered by Qwen3-VL-32B. On 75 sampled validation pairs, VIPER reports VLM-as-Judge physical similarity 3.42 vs ≤2.97 for baselines (Wan variants, VAP, VACE), competitive VBench scores, and 58.8% rank-1 human preference, plus ablations on queries, training stages, and query count.
Significance. If the results hold, the work makes two useful contributions: (1) a clearly formulated task — visual in-context physics transfer for I2V — that is distinct from prior video-as-prompt semantic control, and (2) VIPER-19K, a paired physical-transfer dataset whose annotation and pair-filtering prompts are fully disclosed in the appendix (Figs. 10–11), which is genuinely valuable for reuse and scrutiny. The paper also ships complete training/inference hyperparameters (Table 3), ablations over training stages and query-token count (Tables 2, 4), explicit failure-mode analysis with a known label-granularity limitation (App. F), and a qualitative demonstration that synthetic references can transfer (App. E). These are real strengths. However, the headline quantitative claim rests on a single lenient VLM judge over 75 pairs, and the judge–data relationship needs disentangling before the improvement over video-conditioned baselines (VAP-Wan: 2.87) can be considered established.
major comments (5)
- [§5.2, App. D, Fig. 12] The judge model is never identified, and this matters. §3.2 uses Qwen3-VL-32B to decide which pairs count as transferable physical behavior; Qwen3-VL-4B is both VIPER's physics encoder and the text extractor for baselines (§5.1). If the Fig. 12 judge is also Qwen-family, the evaluation risks measuring co-adaptation between the pair filter's notion of transferability and the judge's rubric — a bias that would credit VIPER specifically, since only VIPER trains on the retained pairs. Please (a) state the judge model explicitly, and (b) re-score the same 75 pairs with at least one judge from a different model family and report agreement. This is the load-bearing metric of the paper.
- [§5.1–5.3, Table 1, Fig. 12] The headline gap (3.42 vs 2.97/2.87) is reported on 75 pairs randomly sampled from a 500-pair validation set, with no confidence intervals or significance testing. The Fig. 12 prompt explicitly instructs leniency ('default expectation should be a positive score (3, 4, or 5)'), compressing scores upward so that a 0.45 gap on a 1–5 scale may not be robust. Please report paired statistics (bootstrap CIs or Wilcoxon signed-rank over per-pair score differences), justify the 75/500 subsample, and ideally evaluate on the full 500 pairs or show the result is stable across resamples.
- [§5.5, Table 2] The claim that 'both the learnable physics queries and the hierarchical training strategy are essential' is not supported at the reported precision: w/o Stage 1 scores 3.41 vs Full 3.42 on physical similarity — indistinguishable without error bars. Either soften the claim to what the data support, or add significance testing. Relatedly, the w/o-queries variant's aesthetic quality collapses to 34.17 (vs ~50 for all other rows, Table 2); this large, unexplained drop suggests instability in the no-query condition and deserves analysis rather than silence.
- [§5.1 (Baselines), Table 1] Non-video baselines receive the reference only as text extracted by Qwen3-VL-4B. Since the paper's motivating premise is that text is a lossy channel for physics, this comparison partly measures the premise rather than VIPER's transfer quality, and it leaves open whether carefully engineered prompts (Fig. 1b) close the gap. A control with best-effort human-written physical prompts on a subset would directly test the 'without carefully engineered prompts' claim. Note VAP and VACE natively accept video conditioning, so the text-conversion caveat applies mainly to the Wan rows — this should be stated precisely.
- [§5.2–5.3, Table 1, Fig. 6, App. D] The human study is the only evaluation leg independent of the Qwen pipeline, so its reporting needs strengthening. (i) Table 1 reports 58.8% rank-1 for VIPER while Fig. 6 reports per-criterion rates of 56.0/68.5/48.0% — the aggregation relating these numbers is never explained, and the 'Preference Rank↓' column mixes a mean rank with a parenthesized percentage without definition. (ii) §5.2 says 10–20 comparison groups per participant; App. D says 15–20 questions. (iii) No inter-rater agreement is reported. Please reconcile the numbers, define the column, and report agreement (e.g., Krippendorff's alpha or rank-1 vote concentration).
minor comments (7)
- [Fig. 12] Typo: 'You are a and pragmatic Physics Engine Evaluator' in the judge prompt.
- [§2.1, Fig. 3, App. B] Typos: 'Other methodss' (§2.1); 'materia' in the Fig. 3 caption; unresolved cross-reference 'Figure?? shows the collection and annotation prompt'.
- [§4.2, Eq. (1)] The system prompt p_sys in Eq. (1) that biases the MLLM toward physical behavior is never shown (Figs. 10–11 cover annotation and pair filtering, not extraction). Please include it in the appendix for reproducibility.
- [§4.2, App. C, Table 4] The number of learnable query tokens used in the main model is only discoverable via the App. C ablation; state it in §4.2 or §5.1. Also, Table 4 omits physical-similarity scores, so the insensitivity claim covers VBench only — say so explicitly in the main text.
- [§5.2, Table 1] VBench motion smoothness and temporal flickering are near-saturated across all methods (95.9–98.7) and do not discriminate physical plausibility. Consider adding a physics-oriented external benchmark or at least discussing the limited discriminative range of these metrics for this paper's claims.
- [§3, §6] Release plans for VIPER-19K, the trained weights, and code are not stated. Given that the dataset and prompts are a major part of the contribution, an explicit release statement would strengthen the paper.
- [§5.1 (Implementation details)] Stage-step budgets (3K/6K/6K) and per-stage data mixes are stated but never justified or ablated; a brief remark on sensitivity would help practitioners.
Circularity Check
Empirical systems paper with no by-construction derivation circularity; MLLM reuse is a validity concern, not a forced prediction.
specific steps
-
other
[Sec. 3.2 pair filtering; Sec. 5.1 baseline protocol; Sec. 5.2 VLM-as-Judge]
"Second, we use Qwen3-VL-32B to assess whether the physical behavior in the reference video is transferable to the target video. ... for methods that do not natively support in-context video conditioning, we use Qwen3-VL-4B-Instruct to extract textual physical descriptions from the reference video and append these descriptions to the generation text prompt. ... we use a VLM-as-Judge protocol to assess physical similarity between the reference video and the generated video on a 1–5 scale."
Not definitional or fit-as-prediction circularity: retained pairs and the physical-similarity rubric share an MLLM-family notion of transferable physics, and baselines are given only lossy text from the same family—so the VLM-as-Judge gap can be inflated by pipeline co-adaptation. The result is still not forced by construction (target-video supervision and human preference remain independent), so this is mild evaluation self-reinforcement rather than a circular derivation step.
full rationale
VIPER’s central claim is empirical (higher reference-video physical similarity and human preference than baselines on held-out pairs), not a first-principles derivation. The generator is supervised by target-video reconstruction under hierarchical training (Stages 1–3), not by the evaluation metric itself, so success is not definitionally guaranteed. Pair filtering with Qwen3-VL-32B (Sec. 3.2), textification of references for non-video baselines (Sec. 5.1), and VIPER’s own Qwen3-VL-4B physics encoder share a model family and create correlated conditioning/evaluation risk, but that is methodological co-adaptation rather than X-defined-as-Y, a fitted parameter renamed as prediction, or a self-cited uniqueness theorem forcing the result. Independent legs remain: VBench quality metrics, a human preference study (25 volunteers), and qualitative transfer cases. No equations equate a reported score to a fitted input; no author-overlapping uniqueness/ansatz citation load-bears the claim. Score 1 reflects only mild pipeline self-reinforcement, not circular derivation.
Axiom & Free-Parameter Ledger
free parameters (5)
- Number of learnable physics query tokens =
Not uniquely fixed in main results; appendix explores 64–512
- Hierarchical stage step budgets and data mixes =
3K + 6K + 6K steps; lr 1e-5
- LoRA rank/alpha and injected DiT layers =
rank 64, alpha 32
- Inference CFG scale and denoising steps =
CFG 6.0, 50 steps
- MLLM pair-filter keep threshold / few-shot decision boundary =
Prompted keep true/false via Qwen3-VL-32B few-shot rubric
axioms (5)
- domain assumption Large pretrained I2V generators already encode useful physical priors that can be elicited by richer conditioning rather than only by more text.
- domain assumption An MLLM with system prompts and learnable queries can extract generation-relevant physical dynamics while suppressing object identity and scene appearance.
- ad hoc to paper Material, trajectory, and physical-impact label compatibility plus MLLM filtering yields pairs whose shared structure is physical transfer rather than semantic/appearance shortcut.
- ad hoc to paper Three-stage training (self-alignment → cross-video transfer → LoRA) is necessary/sufficient to avoid target-image overfitting and preserve base visual priors.
- domain assumption VBench + VLM physical-similarity + small human preference study are adequate proxies for “physically plausible” reference-guided generation.
invented entities (3)
-
Learnable physics query tokens (q_learnable) as compact visual-physics interface
no independent evidence
-
VIPER-19K reference–target pair dataset with material/trajectory/impact taxonomy
no independent evidence
-
Visual in-context physics reasoning task formulation
no independent evidence
read the original abstract
Modern video generation models can synthesize visually compelling and temporally coherent clips, yet controlling their physical behavior remains difficult with standard text and image conditions. The core challenge is a conditioning bottleneck: material response, contact interaction, deformation, and motion trajectory are continuous and relational physical cues that are hard to specify exhaustively in language but can be demonstrated naturally by video. We propose VIPER, a Visual In-Context Physics Reasoning framework for reference-guided image-to-video generation. Given a target image, a brief target prompt, and a reference video, VIPER treats the reference as a dense visual demonstration of the desired physical process rather than an appearance template. It uses a Multimodal Large Language Model (MLLM) to extract reference-derived physical cues and guide a pretrained image-to-video generator through a hierarchical training strategy, enabling physical behavior transfer while preserving the visual prior of the base generator. To support this setting, we construct VIPER-19K, a curated dataset with material, trajectory, and physical-impact annotations, together with filtered reference-target pairs. Experiments on an unseen validation set show that VIPER achieves stronger reference-video physical similarity and higher human preference than representative video generation and video-as-prompt baselines, while maintaining competitive general video quality. Qualitative results further demonstrate that VIPER can transfer reference-derived physical behavior to new target scenes without requiring carefully engineered prompts.
Reference graph
Works this paper leans on
-
[1]
Genesis Authors. 2024. Genesis: A Generative and Universal Physics Engine for Robotics and Beyond. https://github.com/Genesis-Embodied-AI/Genesis
2024
-
[2]
Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, Wenbin Ge, Zhifang Guo, Qidong Huang, Jie Huang, Fei Huang, Binyuan Hui, Shutong Jiang, Zhaohai Li, Mingsheng Li, Mei Li, Kaixin Li, Zicheng Lin, Junyang Lin, Xuejing Liu, Jiawei Liu, Chenglong Liu, Yang Liu, Dayiheng Liu, Shixuan ...
Pith/arXiv arXiv 2025
-
[3]
Yuxuan Bian, Xin Chen, Zenan Li, Tiancheng Zhi, Shen Sang, Linjie Luo, and Qiang Xu. 2025. Video- As-Prompt: Unified Semantic Control for Video Generation.arXiv preprint arXiv:2510.20888 (2025)
arXiv 2025
-
[4]
Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, Clarence Ng, Ricky Wang, and Aditya Ramesh
-
[5]
Haoxin Chen, Yong Zhang, Xiaodong Cun, Menghan Xia, Xintao Wang, Chao Weng, and Ying Shan. 2024. VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models. arXiv:2401.09047 [cs.CV]
Pith/arXiv arXiv 2024
-
[6]
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al. 2024. Scaling rectified flow transformers for high-resolution image synthesis. InICML
2024
-
[7]
Lin Geng Foo, Mark He Huang, Alexandros Lattas, Stylianos Moschoglou, Thabo Beeler, and Christian Theobalt. 2026. Physical Simulator In-the-Loop Video Generation.arXiv preprint arXiv:2603.06408 (2026)
arXiv 2026
-
[8]
Nate Gillman, Charles Herrmann, Michael Freeman, Daksh Aggarwal, Evan Luo, Deqing Sun, and Chen Sun. 2025. Force prompting: Video generation models can learn and generalize physics-based control signals. arXiv preprint arXiv:2505.19386 (2025)
arXiv 2025
-
[9]
Nate Gillman, Yinghua Zhou, Zitian Tang, Evan Luo, Arjan Chakravarthy, Daksh Aggarwal, Michael Freeman, Charles Herrmann, and Chen Sun. 2026. Goal Force: Teaching Video Models To Accomplish Physics-Conditioned Goals. arXiv preprint arXiv:2601.05848 (2026)
arXiv 2026
-
[10]
Yuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang, Yaohui Wang, Yu Qiao, Maneesh Agrawala, Dahua Lin, and Bo Dai. 2024. Animatediff: Animate your personalized text-to-image diffusion models without specific tuning. InICLR
2024
-
[11]
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-Rank Adaptation of Large Language Models. InInternational Conference on Learning Representations.https://openreview.net/forum?id=nZeVKeeFYf9
2022
-
[12]
Lianghua Huang, Wei Wang, Zhi-Fan Wu, Yupeng Shi, Huanzhang Dou, Chen Liang, Yutong Feng, Yu Liu, and Jingren Zhou. 2024. In-context lora for diffusion transformers.arXiv preprint arXiv:2410.23775 (2024)
Pith/arXiv arXiv 2024
-
[13]
Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianxing Wu, Qingyang Jin, Nattapol Chanpaisit, Yaohui Wang, Xinyuan Chen, Limin Wang, Dahua Lin, Yu Qiao, and Ziwei Liu. 2024. VBench: Comprehensive Benchmark Suite for Video Generative Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Re...
2024
-
[14]
Chenfanfu Jiang, Craig Schroeder, Joseph Teran, Alexey Stomakhin, and Andrew Selle. 2016. The material point method for simulating continuum materials. InAcm siggraph 2016 courses. 1–52
2016
-
[15]
Zeyinzi Jiang, Zhen Han, Chaojie Mao, Jingfeng Zhang, Yulin Pan, and Yu Liu. 2025. Vace: All-in- one video creation and editing. InProceedings of the IEEE/CVF International Conference on Computer Vision. 17191–17202
2025
-
[16]
Diederik P Kingma and Max Welling. 2013. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114 (2013)
Pith/arXiv arXiv 2013
-
[17]
Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, et al. 2024. HunyuanVideo: A Systematic Framework For Large Video Generative Models. arXiv preprint arXiv:2412.03603 (2024)
Pith/arXiv arXiv 2024
-
[18]
Zizhang Li, Hong-Xing Yu, Wei Liu, Yin Yang, Charles Herrmann, Gordon Wetzstein, and Jiajun Wu
-
[19]
Bin Lin, Yunyang Ge, Xinhua Cheng, Zongjian Li, Bin Zhu, Shaodong Wang, Xianyi He, Yang Ye, Shenghai Yuan, Liuhan Chen, et al. 2024. Open-Sora Plan: Open-Source Large Video Generation Model. arXiv preprint arXiv:2412.00131 (2024)
Pith/arXiv arXiv 2024
-
[20]
Yaron Lipman, Ricky TQ Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. 2022. Flow matching for generative modeling.arXiv preprint arXiv:2210.02747 (2022)
Pith/arXiv arXiv 2022
-
[21]
Wei Liu, Ziyu Chen, Zizhang Li, Yue Wang, Hong-Xing Yu, and Jiajun Wu. 2026. RealWonder: Real-Time Physical Action-Conditioned Video Generation.arXiv preprint arXiv:2603.05449 (2026)
arXiv 2026
-
[22]
Miles Macklin, Matthias Müller, and Nuttapong Chentanez. 2016. XPBD: position-based simulation of compliant constrained dynamics. InProceedings of the 9th International Conference on Motion in Games. 49–54
2016
-
[23]
William Peebles and Saining Xie. 2023. Scalable diffusion models with transformers. InICCV
2023
-
[24]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the Limits of Transfer Learning with a Unified Text- to-Text Transformer.Journal of Machine Learning Research 21, 140 (2020), 1–67. http://jmlr.org/ papers/v21/20-074.html
2020
-
[25]
Ying Shen, Jerry Xiong, Tianjiao Yu, and Ismini Lourentzou. 2026. Phantom: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics.arXiv preprint arXiv:2604.08503 (2026)
Pith/arXiv arXiv 2026
-
[26]
Zhenxiong Tan, Songhua Liu, Xingyi Yang, Qiaochu Xue, and Xinchao Wang. 2025. Ominicontrol: Minimal and universal control for diffusion transformer. InProceedings of the IEEE/CVF International Conference on Computer Vision. 14940–14950
2025
-
[27]
Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, et al. 2025. Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314 (2025)
Pith/arXiv arXiv 2025
-
[28]
Chen Wang, Chuhao Chen, Yiming Huang, Zhiyang Dou, Yuan Liu, Jiatao Gu, and Lingjie Liu. 2025. Physctrl: Generative physics for controllable and physics-grounded video generation.arXiv preprint arXiv:2509.20358 (2025)
arXiv 2025
-
[29]
Jing Wang, Ao Ma, Ke Cao, Jun Zheng, Zhanjie Zhang, Jiasong Feng, Shanyuan Liu, Yuhang Ma, Bo Cheng, Dawei Leng, et al. 2025. Wisa: World simulator assistant for physics-aware text-to-video generation. arXiv preprint arXiv:2503.08153 (2025). 15
Pith/arXiv arXiv 2025
-
[30]
Wenhao Wang and Yi Yang. 2024. VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to- Video Diffusion Models. (2024)
2024
-
[31]
Cong Wei, Quande Liu, Zixuan Ye, Qiulin Wang, Xintao Wang, Pengfei Wan, Kun Gai, and Wenhu Chen. 2025. Univideo: Unified understanding, generation, and editing for videos. arXiv preprint arXiv:2510.08377 (2025)
Pith/arXiv arXiv 2025
-
[32]
Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al. 2024. CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer. InICLR
2024
-
[33]
Zixuan Ye, Xuanhua He, Quande Liu, Qiulin Wang, Xintao Wang, Pengfei Wan, Di Zhang, Kun Gai, Qifeng Chen, and Wenhan Luo. 2025. Unic: Unified in-context video editing. arXiv preprint arXiv:2506.04216 (2025)
Pith/arXiv arXiv 2025
-
[34]
Yu Yuan, Xijun Wang, Tharindu Wickremasinghe, Zeeshan Nadir, Bole Ma, and Stanley H Chan
-
[35]
Haoze Zhang, Tianyu Huang, Zichen Wan, Xiaowei Jin, Hongzhi Zhang, Hui Li, and Wangmeng Zuo
-
[36]
Zechuan Zhang, Ji Xie, Yu Lu, Zongxin Yang, and Yi Yang. 2025. Enabling instructional image editing with in-context generation in large scale diffusion transformer. InThe Thirty-ninth Annual Conference on Neural Information Processing Systems
2025
-
[37]
arXiv preprint arXiv:2509.21309 (2025)
NewtonGen: Physics-Consistent and Controllable Text-to-Video Generation via Neural Newtonian Dynamics. arXiv preprint arXiv:2509.21309 (2025)
arXiv 2025
-
[39]
PhysChoreo: Physics-Controllable Video Generation with Part-Aware Semantic Grounding.arXiv preprint arXiv:2511.20562 (2025)
Pith/arXiv arXiv 2025
-
[41]
Dian Zheng, Ziqi Huang, Hongbo Liu, Kai Zou, Yinan He, Fan Zhang, Yuanhan Zhang, Jingwen He, Wei-Shi Zheng, Yu Qiao, and Ziwei Liu. 2025. VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness.arXiv preprint arXiv:2503.21755 (2025). 16 A Hyperparameter Settings Table 3 summarizes the hyperparameters used to train and sample VIPE...
Pith/arXiv arXiv 2025
-
[2024]
https://openai.com/research/ video-generation-models-as-world-simulators
Video generation models as world simulators. https://openai.com/research/ video-generation-models-as-world-simulators
-
[2025]
InProceedings of the IEEE/CVF International Conference on Computer Vision
Wonderplay: Dynamic 3d scene generation from a single image and actions. InProceedings of the IEEE/CVF International Conference on Computer Vision. 9080–9090
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.