Pith. sign in

REVIEW 5 major objections 7 minor 41 references

A reference video can specify physical behavior that text prompts cannot, and VIPER transfers that behavior to new scenes.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-30 21:24 UTC pith:OUD45OYC

load-bearing objection Solid systems paper on reference-as-physics for I2V; real packaging contribution, but the headline VLM-as-Judge gap is partly compromised by Qwen reuse across data, conditioning, and scoring. the 5 major comments →

arxiv 2607.23472 v1 pith:OUD45OYC submitted 2026-07-26 cs.CV

VIPER: Visual In-Context Physics Reasoning for Physically Plausible Video Generation

classification cs.CV
keywords video generationphysical plausibilityimage-to-videoin-context learningreference-guided generationmultimodal large language modelVIPER-19K
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Video generators already make smooth-looking clips, but they still struggle to obey real physical processes when the only controls are a short text prompt and a still image. The paper argues the bottleneck is not synthesis capacity so much as conditioning: material response, contact, deformation, and trajectories are continuous, relational cues that are hard to spell out in language and easy to show in video. VIPER treats a reference clip as a dense visual demonstration of the desired physics, uses a multimodal language model to pull compact physical cues from that clip, and steers a pretrained image-to-video generator so the target image follows the reference dynamics without copying its appearance. A curated dataset of about nineteen thousand annotated clips and filtered reference–target pairs supports this training. On held-out pairs, the method scores higher physical similarity to the reference and wins more human preference than strong generation and video-as-prompt baselines, while staying competitive on ordinary video-quality metrics.

Core claim

The paper claims that physically plausible image-to-video generation can be cast as visual in-context physics reasoning: given a target image, a brief prompt, and a reference video, an MLLM can extract transferable physical cues from the reference and, via hierarchical training into a DiT image-to-video backbone, produce a target video that inherits the reference’s material response, trajectory, and impact behavior while preserving the base generator’s visual prior—without exhaustive physics prompts.

What carries the argument

VIPER’s reference-video physical conditioning path: learnable physics query tokens read the reference through an MLLM, a connector projects those hidden states into DiT condition tokens, and a three-stage hierarchical training schedule first aligns self-reference conditioning, then trains cross-video physical transfer on VIPER-19K pairs, then lightly adapts the generator with LoRA so physics tokens matter without overwriting the pretrained visual prior.

Load-bearing premise

The method assumes that compact MLLM query tokens plus pair training isolate transferable physical dynamics rather than appearance shortcuts, semantic overlap, or artifacts of how the same model family helps build pairs, write baseline prompts, and judge physical similarity.

What would settle it

On truly held-out reference–target pairs with matched high-level labels but deliberately inverted dynamics (for example opposite spin direction, shatter versus bounce, or mismatched impact scale), measure whether VIPER’s physical-similarity and human preference advantages disappear while baselines catch up when given equally detailed text physics descriptions.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Reference video becomes a practical control interface for material response, contact, and trajectory without hand-written physics prompts.
  • Pretrained image-to-video models can be steered for physical behavior transfer while keeping their existing visual quality prior via staged alignment and light LoRA adaptation.
  • Datasets organized by material, trajectory, and physical-impact labels with filtered cross-appearance pairs become the natural supervision unit for this task.
  • Qualitative transfer of melting, deformation, and related processes to new target scenes is achievable from brief target prompts plus a demonstration clip.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If reference-as-physics-spec works at scale, interactive tools could let users pick a short demo clip instead of engineering force or material language for each shot.
  • Coarse taxonomy gaps called out in the limitations (for example spin without direction) suggest the next bottleneck is finer causal labels and compatibility checks, not only bigger generators.
  • Heavy reuse of one multimodal model family across pair filtering, baseline textification, and judging implies future benchmarks may need an independent physics scorer to separate method gains from judge alignment.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

5 major / 7 minor

Summary. The paper proposes VIPER, a reference-guided image-to-video framework in which a reference video serves as a dense visual demonstration of desired physical behavior (material response, trajectory, contact, impact). An MLLM (Qwen3-VL-4B) with learnable physics query tokens encodes the reference into compact condition tokens injected into a Wan2.2-I2V-14B DiT generator; a three-stage hierarchical training scheme (self-alignment → cross-video transfer → LoRA adaptation) is used to avoid pixel-copying and preserve the base prior. The authors also construct VIPER-19K, a dataset with material/trajectory/impact annotations and reference–target pairs filtered by Qwen3-VL-32B. On 75 sampled validation pairs, VIPER reports VLM-as-Judge physical similarity 3.42 vs ≤2.97 for baselines (Wan variants, VAP, VACE), competitive VBench scores, and 58.8% rank-1 human preference, plus ablations on queries, training stages, and query count.

Significance. If the results hold, the work makes two useful contributions: (1) a clearly formulated task — visual in-context physics transfer for I2V — that is distinct from prior video-as-prompt semantic control, and (2) VIPER-19K, a paired physical-transfer dataset whose annotation and pair-filtering prompts are fully disclosed in the appendix (Figs. 10–11), which is genuinely valuable for reuse and scrutiny. The paper also ships complete training/inference hyperparameters (Table 3), ablations over training stages and query-token count (Tables 2, 4), explicit failure-mode analysis with a known label-granularity limitation (App. F), and a qualitative demonstration that synthetic references can transfer (App. E). These are real strengths. However, the headline quantitative claim rests on a single lenient VLM judge over 75 pairs, and the judge–data relationship needs disentangling before the improvement over video-conditioned baselines (VAP-Wan: 2.87) can be considered established.

major comments (5)
  1. [§5.2, App. D, Fig. 12] The judge model is never identified, and this matters. §3.2 uses Qwen3-VL-32B to decide which pairs count as transferable physical behavior; Qwen3-VL-4B is both VIPER's physics encoder and the text extractor for baselines (§5.1). If the Fig. 12 judge is also Qwen-family, the evaluation risks measuring co-adaptation between the pair filter's notion of transferability and the judge's rubric — a bias that would credit VIPER specifically, since only VIPER trains on the retained pairs. Please (a) state the judge model explicitly, and (b) re-score the same 75 pairs with at least one judge from a different model family and report agreement. This is the load-bearing metric of the paper.
  2. [§5.1–5.3, Table 1, Fig. 12] The headline gap (3.42 vs 2.97/2.87) is reported on 75 pairs randomly sampled from a 500-pair validation set, with no confidence intervals or significance testing. The Fig. 12 prompt explicitly instructs leniency ('default expectation should be a positive score (3, 4, or 5)'), compressing scores upward so that a 0.45 gap on a 1–5 scale may not be robust. Please report paired statistics (bootstrap CIs or Wilcoxon signed-rank over per-pair score differences), justify the 75/500 subsample, and ideally evaluate on the full 500 pairs or show the result is stable across resamples.
  3. [§5.5, Table 2] The claim that 'both the learnable physics queries and the hierarchical training strategy are essential' is not supported at the reported precision: w/o Stage 1 scores 3.41 vs Full 3.42 on physical similarity — indistinguishable without error bars. Either soften the claim to what the data support, or add significance testing. Relatedly, the w/o-queries variant's aesthetic quality collapses to 34.17 (vs ~50 for all other rows, Table 2); this large, unexplained drop suggests instability in the no-query condition and deserves analysis rather than silence.
  4. [§5.1 (Baselines), Table 1] Non-video baselines receive the reference only as text extracted by Qwen3-VL-4B. Since the paper's motivating premise is that text is a lossy channel for physics, this comparison partly measures the premise rather than VIPER's transfer quality, and it leaves open whether carefully engineered prompts (Fig. 1b) close the gap. A control with best-effort human-written physical prompts on a subset would directly test the 'without carefully engineered prompts' claim. Note VAP and VACE natively accept video conditioning, so the text-conversion caveat applies mainly to the Wan rows — this should be stated precisely.
  5. [§5.2–5.3, Table 1, Fig. 6, App. D] The human study is the only evaluation leg independent of the Qwen pipeline, so its reporting needs strengthening. (i) Table 1 reports 58.8% rank-1 for VIPER while Fig. 6 reports per-criterion rates of 56.0/68.5/48.0% — the aggregation relating these numbers is never explained, and the 'Preference Rank↓' column mixes a mean rank with a parenthesized percentage without definition. (ii) §5.2 says 10–20 comparison groups per participant; App. D says 15–20 questions. (iii) No inter-rater agreement is reported. Please reconcile the numbers, define the column, and report agreement (e.g., Krippendorff's alpha or rank-1 vote concentration).
minor comments (7)
  1. [Fig. 12] Typo: 'You are a and pragmatic Physics Engine Evaluator' in the judge prompt.
  2. [§2.1, Fig. 3, App. B] Typos: 'Other methodss' (§2.1); 'materia' in the Fig. 3 caption; unresolved cross-reference 'Figure?? shows the collection and annotation prompt'.
  3. [§4.2, Eq. (1)] The system prompt p_sys in Eq. (1) that biases the MLLM toward physical behavior is never shown (Figs. 10–11 cover annotation and pair filtering, not extraction). Please include it in the appendix for reproducibility.
  4. [§4.2, App. C, Table 4] The number of learnable query tokens used in the main model is only discoverable via the App. C ablation; state it in §4.2 or §5.1. Also, Table 4 omits physical-similarity scores, so the insensitivity claim covers VBench only — say so explicitly in the main text.
  5. [§5.2, Table 1] VBench motion smoothness and temporal flickering are near-saturated across all methods (95.9–98.7) and do not discriminate physical plausibility. Consider adding a physics-oriented external benchmark or at least discussing the limited discriminative range of these metrics for this paper's claims.
  6. [§3, §6] Release plans for VIPER-19K, the trained weights, and code are not stated. Given that the dataset and prompts are a major part of the contribution, an explicit release statement would strengthen the paper.
  7. [§5.1 (Implementation details)] Stage-step budgets (3K/6K/6K) and per-stage data mixes are stated but never justified or ablated; a brief remark on sensitivity would help practitioners.

Circularity Check

1 steps flagged

Empirical systems paper with no by-construction derivation circularity; MLLM reuse is a validity concern, not a forced prediction.

specific steps
  1. other [Sec. 3.2 pair filtering; Sec. 5.1 baseline protocol; Sec. 5.2 VLM-as-Judge]
    "Second, we use Qwen3-VL-32B to assess whether the physical behavior in the reference video is transferable to the target video. ... for methods that do not natively support in-context video conditioning, we use Qwen3-VL-4B-Instruct to extract textual physical descriptions from the reference video and append these descriptions to the generation text prompt. ... we use a VLM-as-Judge protocol to assess physical similarity between the reference video and the generated video on a 1–5 scale."

    Not definitional or fit-as-prediction circularity: retained pairs and the physical-similarity rubric share an MLLM-family notion of transferable physics, and baselines are given only lossy text from the same family—so the VLM-as-Judge gap can be inflated by pipeline co-adaptation. The result is still not forced by construction (target-video supervision and human preference remain independent), so this is mild evaluation self-reinforcement rather than a circular derivation step.

full rationale

VIPER’s central claim is empirical (higher reference-video physical similarity and human preference than baselines on held-out pairs), not a first-principles derivation. The generator is supervised by target-video reconstruction under hierarchical training (Stages 1–3), not by the evaluation metric itself, so success is not definitionally guaranteed. Pair filtering with Qwen3-VL-32B (Sec. 3.2), textification of references for non-video baselines (Sec. 5.1), and VIPER’s own Qwen3-VL-4B physics encoder share a model family and create correlated conditioning/evaluation risk, but that is methodological co-adaptation rather than X-defined-as-Y, a fitted parameter renamed as prediction, or a self-cited uniqueness theorem forcing the result. Independent legs remain: VBench quality metrics, a human preference study (25 volunteers), and qualitative transfer cases. No equations equate a reported score to a fitted input; no author-overlapping uniqueness/ansatz citation load-bears the claim. Score 1 reflects only mild pipeline self-reinforcement, not circular derivation.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 3 invented entities

Load-bearing commitments are engineering and dataset assumptions, not mathematical axioms. The central empirical claim rests on pretrained generative priors containing usable physics, on MLLMs being able to factor physics from appearance into a small token bottleneck, on coarse material/trajectory/impact taxonomies plus MLLM pair filtering defining valid transfer supervision, and on frozen-backbone + LoRA adaptation preserving visual quality while accepting physics tokens. Free parameters are standard training/design choices (query count, stage lengths, LoRA rank, CFG, etc.). No new physical particles or forces are postulated.

free parameters (5)
  • Number of learnable physics query tokens = Not uniquely fixed in main results; appendix explores 64–512
    Capacity of the MLLM-to-DiT physics bottleneck; ablated at 64–512 with modest VBench differences; main setting not uniquely determined by theory.
  • Hierarchical stage step budgets and data mixes = 3K + 6K + 6K steps; lr 1e-5
    3K self-alignment / 6K transfer / 6K LoRA steps and stage-specific real vs synthetic mixes are hand-chosen schedules that enable reported stability and transfer.
  • LoRA rank/alpha and injected DiT layers = rank 64, alpha 32
    Controls how much the generator adapts to physics tokens vs retaining pretrained prior; chosen design knobs (rank 64, alpha 32, q/k/v/ffn layers).
  • Inference CFG scale and denoising steps = CFG 6.0, 50 steps
    Sampling hyperparameters that affect quality/physics tradeoff in all reported videos.
  • MLLM pair-filter keep threshold / few-shot decision boundary = Prompted keep true/false via Qwen3-VL-32B few-shot rubric
    Determines which reference–target pairs enter supervision; directly shapes what “transferable physics” means in VIPER-19K.
axioms (5)
  • domain assumption Large pretrained I2V generators already encode useful physical priors that can be elicited by richer conditioning rather than only by more text.
    Stated as the conditioning-bottleneck motivation in Sec. 1; without it, reference cues would not improve physical plausibility.
  • domain assumption An MLLM with system prompts and learnable queries can extract generation-relevant physical dynamics while suppressing object identity and scene appearance.
    Core of Sec. 4.2 reference-video physical conditioning; method fails if queries entangle appearance with physics.
  • ad hoc to paper Material, trajectory, and physical-impact label compatibility plus MLLM filtering yields pairs whose shared structure is physical transfer rather than semantic/appearance shortcut.
    Sec. 3 pair construction; Limitations admit coarse labels (e.g., spin direction) still cause errors.
  • ad hoc to paper Three-stage training (self-alignment → cross-video transfer → LoRA) is necessary/sufficient to avoid target-image overfitting and preserve base visual priors.
    Sec. 4.3; supported by ablation Table 2 but remains a design hypothesis.
  • domain assumption VBench + VLM physical-similarity + small human preference study are adequate proxies for “physically plausible” reference-guided generation.
    Sec. 5 evaluation protocol; no simulator ground truth.
invented entities (3)
  • Learnable physics query tokens (q_learnable) as compact visual-physics interface no independent evidence
    purpose: Pool reference-video dynamics into a small set of tokens projected into DiT conditioning without passing full MLLM hidden states.
    Architectural postulate of Sec. 4.2; standard query-token pattern applied to physics conditioning. Independent evidence only via ablations/metrics inside this paper.
  • VIPER-19K reference–target pair dataset with material/trajectory/impact taxonomy no independent evidence
    purpose: Provide supervision for visual in-context physics transfer across appearance changes.
    New curated resource (Sec. 3); existence is operational, but the taxonomy and pair notion are paper-defined constructs without external standardized benchmark status yet.
  • Visual in-context physics reasoning task formulation no independent evidence
    purpose: Frame reference-guided I2V as inferring and applying physical process from a demo video rather than appearance cloning or exhaustive text.
    Conceptual framing in abstract/Sec. 1; useful task name, not a physical entity. Evidence is the empirical system built around it.

pith-pipeline@v1.2.0-grok45-kimik3 · 19303 in / 4253 out tokens · 88187 ms · 2026-07-30T21:24:54.685679+00:00 · methodology

0 comments
read the original abstract

Modern video generation models can synthesize visually compelling and temporally coherent clips, yet controlling their physical behavior remains difficult with standard text and image conditions. The core challenge is a conditioning bottleneck: material response, contact interaction, deformation, and motion trajectory are continuous and relational physical cues that are hard to specify exhaustively in language but can be demonstrated naturally by video. We propose VIPER, a Visual In-Context Physics Reasoning framework for reference-guided image-to-video generation. Given a target image, a brief target prompt, and a reference video, VIPER treats the reference as a dense visual demonstration of the desired physical process rather than an appearance template. It uses a Multimodal Large Language Model (MLLM) to extract reference-derived physical cues and guide a pretrained image-to-video generator through a hierarchical training strategy, enabling physical behavior transfer while preserving the visual prior of the base generator. To support this setting, we construct VIPER-19K, a curated dataset with material, trajectory, and physical-impact annotations, together with filtered reference-target pairs. Experiments on an unseen validation set show that VIPER achieves stronger reference-video physical similarity and higher human preference than representative video generation and video-as-prompt baselines, while maintaining competitive general video quality. Qualitative results further demonstrate that VIPER can transfer reference-derived physical behavior to new target scenes without requiring carefully engineered prompts.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

41 extracted references · 14 linked inside Pith

  1. [1]

    Genesis Authors. 2024. Genesis: A Generative and Universal Physics Engine for Robotics and Beyond. https://github.com/Genesis-Embodied-AI/Genesis

  2. [2]

    Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, Wenbin Ge, Zhifang Guo, Qidong Huang, Jie Huang, Fei Huang, Binyuan Hui, Shutong Jiang, Zhaohai Li, Mingsheng Li, Mei Li, Kaixin Li, Zicheng Lin, Junyang Lin, Xuejing Liu, Jiawei Liu, Chenglong Liu, Yang Liu, Dayiheng Liu, Shixuan ...

  3. [3]

    Yuxuan Bian, Xin Chen, Zenan Li, Tiancheng Zhi, Shen Sang, Linjie Luo, and Qiang Xu. 2025. Video- As-Prompt: Unified Semantic Control for Video Generation.arXiv preprint arXiv:2510.20888 (2025)

  4. [4]

    Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, Clarence Ng, Ricky Wang, and Aditya Ramesh

  5. [5]

    Haoxin Chen, Yong Zhang, Xiaodong Cun, Menghan Xia, Xintao Wang, Chao Weng, and Ying Shan. 2024. VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models. arXiv:2401.09047 [cs.CV]

  6. [6]

    Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al. 2024. Scaling rectified flow transformers for high-resolution image synthesis. InICML

  7. [7]

    Lin Geng Foo, Mark He Huang, Alexandros Lattas, Stylianos Moschoglou, Thabo Beeler, and Christian Theobalt. 2026. Physical Simulator In-the-Loop Video Generation.arXiv preprint arXiv:2603.06408 (2026)

  8. [8]

    Nate Gillman, Charles Herrmann, Michael Freeman, Daksh Aggarwal, Evan Luo, Deqing Sun, and Chen Sun. 2025. Force prompting: Video generation models can learn and generalize physics-based control signals. arXiv preprint arXiv:2505.19386 (2025)

  9. [9]

    Nate Gillman, Yinghua Zhou, Zitian Tang, Evan Luo, Arjan Chakravarthy, Daksh Aggarwal, Michael Freeman, Charles Herrmann, and Chen Sun. 2026. Goal Force: Teaching Video Models To Accomplish Physics-Conditioned Goals. arXiv preprint arXiv:2601.05848 (2026)

  10. [10]

    Yuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang, Yaohui Wang, Yu Qiao, Maneesh Agrawala, Dahua Lin, and Bo Dai. 2024. Animatediff: Animate your personalized text-to-image diffusion models without specific tuning. InICLR

  11. [11]

    Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-Rank Adaptation of Large Language Models. InInternational Conference on Learning Representations.https://openreview.net/forum?id=nZeVKeeFYf9

  12. [12]

    Lianghua Huang, Wei Wang, Zhi-Fan Wu, Yupeng Shi, Huanzhang Dou, Chen Liang, Yutong Feng, Yu Liu, and Jingren Zhou. 2024. In-context lora for diffusion transformers.arXiv preprint arXiv:2410.23775 (2024)

  13. [13]

    Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianxing Wu, Qingyang Jin, Nattapol Chanpaisit, Yaohui Wang, Xinyuan Chen, Limin Wang, Dahua Lin, Yu Qiao, and Ziwei Liu. 2024. VBench: Comprehensive Benchmark Suite for Video Generative Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Re...

  14. [14]

    Chenfanfu Jiang, Craig Schroeder, Joseph Teran, Alexey Stomakhin, and Andrew Selle. 2016. The material point method for simulating continuum materials. InAcm siggraph 2016 courses. 1–52

  15. [15]

    Zeyinzi Jiang, Zhen Han, Chaojie Mao, Jingfeng Zhang, Yulin Pan, and Yu Liu. 2025. Vace: All-in- one video creation and editing. InProceedings of the IEEE/CVF International Conference on Computer Vision. 17191–17202

  16. [16]

    Diederik P Kingma and Max Welling. 2013. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114 (2013)

  17. [17]

    Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, et al. 2024. HunyuanVideo: A Systematic Framework For Large Video Generative Models. arXiv preprint arXiv:2412.03603 (2024)

  18. [18]

    Zizhang Li, Hong-Xing Yu, Wei Liu, Yin Yang, Charles Herrmann, Gordon Wetzstein, and Jiajun Wu

  19. [19]

    Bin Lin, Yunyang Ge, Xinhua Cheng, Zongjian Li, Bin Zhu, Shaodong Wang, Xianyi He, Yang Ye, Shenghai Yuan, Liuhan Chen, et al. 2024. Open-Sora Plan: Open-Source Large Video Generation Model. arXiv preprint arXiv:2412.00131 (2024)

  20. [20]

    Yaron Lipman, Ricky TQ Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. 2022. Flow matching for generative modeling.arXiv preprint arXiv:2210.02747 (2022)

  21. [21]

    Wei Liu, Ziyu Chen, Zizhang Li, Yue Wang, Hong-Xing Yu, and Jiajun Wu. 2026. RealWonder: Real-Time Physical Action-Conditioned Video Generation.arXiv preprint arXiv:2603.05449 (2026)

  22. [22]

    Miles Macklin, Matthias Müller, and Nuttapong Chentanez. 2016. XPBD: position-based simulation of compliant constrained dynamics. InProceedings of the 9th International Conference on Motion in Games. 49–54

  23. [23]

    William Peebles and Saining Xie. 2023. Scalable diffusion models with transformers. InICCV

  24. [24]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the Limits of Transfer Learning with a Unified Text- to-Text Transformer.Journal of Machine Learning Research 21, 140 (2020), 1–67. http://jmlr.org/ papers/v21/20-074.html

  25. [25]

    Ying Shen, Jerry Xiong, Tianjiao Yu, and Ismini Lourentzou. 2026. Phantom: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics.arXiv preprint arXiv:2604.08503 (2026)

  26. [26]

    Zhenxiong Tan, Songhua Liu, Xingyi Yang, Qiaochu Xue, and Xinchao Wang. 2025. Ominicontrol: Minimal and universal control for diffusion transformer. InProceedings of the IEEE/CVF International Conference on Computer Vision. 14940–14950

  27. [27]

    Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, et al. 2025. Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314 (2025)

  28. [28]

    Chen Wang, Chuhao Chen, Yiming Huang, Zhiyang Dou, Yuan Liu, Jiatao Gu, and Lingjie Liu. 2025. Physctrl: Generative physics for controllable and physics-grounded video generation.arXiv preprint arXiv:2509.20358 (2025)

  29. [29]

    Jing Wang, Ao Ma, Ke Cao, Jun Zheng, Zhanjie Zhang, Jiasong Feng, Shanyuan Liu, Yuhang Ma, Bo Cheng, Dawei Leng, et al. 2025. Wisa: World simulator assistant for physics-aware text-to-video generation. arXiv preprint arXiv:2503.08153 (2025). 15

  30. [30]

    Wenhao Wang and Yi Yang. 2024. VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to- Video Diffusion Models. (2024)

  31. [31]

    Cong Wei, Quande Liu, Zixuan Ye, Qiulin Wang, Xintao Wang, Pengfei Wan, Kun Gai, and Wenhu Chen. 2025. Univideo: Unified understanding, generation, and editing for videos. arXiv preprint arXiv:2510.08377 (2025)

  32. [32]

    Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al. 2024. CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer. InICLR

  33. [33]

    Zixuan Ye, Xuanhua He, Quande Liu, Qiulin Wang, Xintao Wang, Pengfei Wan, Di Zhang, Kun Gai, Qifeng Chen, and Wenhan Luo. 2025. Unic: Unified in-context video editing. arXiv preprint arXiv:2506.04216 (2025)

  34. [34]

    Yu Yuan, Xijun Wang, Tharindu Wickremasinghe, Zeeshan Nadir, Bole Ma, and Stanley H Chan

  35. [35]

    Haoze Zhang, Tianyu Huang, Zichen Wan, Xiaowei Jin, Hongzhi Zhang, Hui Li, and Wangmeng Zuo

  36. [36]

    Zechuan Zhang, Ji Xie, Yu Lu, Zongxin Yang, and Yi Yang. 2025. Enabling instructional image editing with in-context generation in large scale diffusion transformer. InThe Thirty-ninth Annual Conference on Neural Information Processing Systems

  37. [37]

    arXiv preprint arXiv:2509.21309 (2025)

    NewtonGen: Physics-Consistent and Controllable Text-to-Video Generation via Neural Newtonian Dynamics. arXiv preprint arXiv:2509.21309 (2025)

  38. [39]

    PhysChoreo: Physics-Controllable Video Generation with Part-Aware Semantic Grounding.arXiv preprint arXiv:2511.20562 (2025)

  39. [41]

    Rigid", NOT

    Dian Zheng, Ziqi Huang, Hongbo Liu, Kai Zou, Yinan He, Fan Zhang, Yuanhan Zhang, Jingwen He, Wei-Shi Zheng, Yu Qiao, and Ziwei Liu. 2025. VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness.arXiv preprint arXiv:2503.21755 (2025). 16 A Hyperparameter Settings Table 3 summarizes the hyperparameters used to train and sample VIPE...

  40. [2024]

    https://openai.com/research/ video-generation-models-as-world-simulators

    Video generation models as world simulators. https://openai.com/research/ video-generation-models-as-world-simulators

  41. [2025]

    InProceedings of the IEEE/CVF International Conference on Computer Vision

    Wonderplay: Dynamic 3d scene generation from a single image and actions. InProceedings of the IEEE/CVF International Conference on Computer Vision. 9080–9090