REVIEW 4 major objections 5 minor 1 cited by
Video-GPT via Next Clip Diffusion
T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read The paper claims that next clip diffusion pretraining — denoising each new video clip only from the clean clips before it — produces state-of-the-art video prediction and a general video foundation model.
desk verdict Genuinely new clip-level pretraining paradigm with a broad empirical program, but the SOTA benchmark claims are not yet clean enough to verify—needs dedup, a fixed evaluation protocol, and baseline normalization. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the noise-clean interleaved clip sequence under a hierarchical causal mask. Each video is split into clips; some clips are noised by flow matching, and the input alternates noisy clip, clean clip, noisy clip, clean clip in temporal order, with boundary tokens (<diff>, <img>) and the noise weight alpha feeding the timestep information. The clip-level mask lets the k-th noisy clip attend to itself and to all earlier clean clips but not to earlier noisy clips, which forces denoising to be conditioned on the correct history; frame-level and patch-level masks then decide bidirectional versus causal attention inside frames and patches. This is what lets diffusion operate inside a clip while GPT-style autoregression operates across clips, and it is what the paper credits for combining short-term generation quality with long-range prediction.
What would settle it
Fix the inference clip size and history-conditioned guidance scale without consulting test scores, or run the same pretrained weights on a held-out physics benchmark; if the margin over VideoPoet narrows to noise, the claim that next clip diffusion causes the improvement fails. Alternatively, pretrain an identical-architecture baseline with full-sequence diffusion or next-frame autoregression on the same data: matching Physics-IQ would also falsify the paradigm's specific role.
Extended reading notes
Core claim
The central claim is that next clip diffusion — autoregressively denoising each new clip against the already-clean clips in its past — is a sufficient pretraining objective for a general video model. The paper argues that earlier diffusion video models generate within the whole clip or sequence but struggle with long-term prediction, while pure autoregressive token predictors lag behind diffusion in synthesis quality; interleaving noisy and clean clips in temporal order lets the same transformer do both. Concretely, the model is trained with an L2 loss that predicts the clean clip directly, uses flow matching noise schedules, and applies a hierarchical attention mask at clip, frame, and patch levels so that noisy clips see only past clean context while patch tokens inside a frame attend fully. The authors report that this pretraining, on 70 million unlabeled videos, outperforms prior video prediction systems on deterministic physical futures and on uncertain human motion, and that the same weights adapt to class-to-video, text-to-video, image animation, video classification, retrieval, and object segmentation.
Load-bearing premise
The headline results assume the Physics-IQ and Kinetics-600 comparisons measure the method fairly, so the reported gains come from next clip diffusion and not from inference or pretraining choices tuned on those same benchmarks.
Editorial extensions
If this is right
- One unlabeled-video objective can serve as foundation pretraining for both generative and understanding tasks, so video-only data becomes a viable scaling axis.
- Longer pretraining context and larger unlabeled corpora directly improve physical prediction, suggesting the paradigm follows a data-scaling trend like language models.
- The model can generate arbitrarily long futures by iteratively treating each denoised clip as clean history, which is exactly the GPT analogy carried to inference.
- A plain vanilla transformer without U-Net or DiT designs reaches competitive FVD on Kinetics-600, so architectural complexity is not what drives the result.
- Downstream tasks transfer with small fine-tuning sets — fewer than one hundred videos for image animation — indicating the pretrained representation is broadly reusable.
Reading between the lines
- If the Physics-IQ margin is genuine, next clip diffusion may be a better prior for learning intuitive physics than full-sequence diffusion, because causal clean-history conditioning matches how physical dynamics unfold; a direct comparison holding architecture and data fixed would test that.
- The same interleaved noising recipe could transfer to other continuous temporal modalities such as audio or sensor streams, where clean history is also available and next-step prediction defines the task.
- Because the inference clip size and guidance scale were selected on the evaluation benchmark itself, part of the reported margin could come from benchmark alignment rather than the paradigm; an independent held-out physics suite would separate the two.
- If it extends, combining next clip tokens with language tokens in one transformer could yield a true multimodal world model that narrates and predicts video jointly.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Video-GPT, a 3.8B-parameter transformer for video generation and understanding, trained with a proposed 'next clip diffusion' objective. A video is partitioned into clips; each clip is noised via flow matching, and the model must denoise each noisy clip conditioned only on the previous clean clips through hierarchical causal attention masks. The pretrained model is evaluated zero-shot on Physics-IQ (34.97) and Kinetics-600 (FVD 315.40) for deterministic and stochastic video prediction, and is fine-tuned on six tasks: class-to-video, text-to-video, image animation, video classification, video retrieval, and video object segmentation. The paper reports state-of-the-art performance on Physics-IQ and Kinetics-600 and favorable results on the downstream tasks.
Significance. If the claims hold, the main contribution is a single self-supervised objective on unlabeled video that unifies autoregressive and diffusion modeling at the clip level and transfers to both generation and understanding tasks, without requiring text annotations for pretraining. The method is clearly described, the progressive training strategy and ablations are useful, and the qualitative results are plausible. The central risk is not internal inconsistency but the benchmark evidence: the headline scores are produced after tuning inference hyperparameters on the same benchmark, and the comparisons lack protocol-equivalence and deduplication controls, so the SOTA claim is not independently verifiable as presented.
major comments (4)
- [Sec. 4.3, Table 4] The history-conditioned CFG scale c=3.0 and the inference clip size Nk are selected by maximizing the Physics-IQ score in Table 4, and the same selection is then reported as the headline result (34.97) in Table 2 and the abstract. This means the reported SOTA number is a test-set-tuned result, not an independent evaluation. Please report a validation-split or fixed-protocol evaluation, or at least show the score without this tuning, and clearly state the tuning procedure in the main text.
- [Sec. 4.2, Table 2] Table 2 mixes image-to-video (I2V) and video-to-video (V2V) baselines, models with and without text conditioning, and methods whose condition-frame length and prediction length are not stated to match the Physics-IQ protocol (3 s condition, 5 s prediction). Since several numbers are taken from external papers while only LVM, Open-Sora-Plan, and Seine are re-tested by the authors, the comparison does not establish that all models received the same input conditions and were evaluated on the same output horizon. Without protocol equivalence, the 5-point gap over VideoPoet cannot be attributed to the next-clip-diffusion paradigm.
- [Sec. 4.1, Sec. 4.2, Table 3] Pretraining is performed on Panda-70M, and evaluation on Kinetics-600 (and Physics-IQ) with no reported deduplication. Panda-70M aggregates large-scale internet video collections and may contain Kinetics-600 or Physics-IQ clips or near-duplicates; if so, the reported FVD 315.40 and Physics-IQ 34.97 partly reflect memorization rather than generalization. The Limitations (Sec. C) mention only model scale, not this risk. Please perform and report a hash-based or feature-based overlap analysis between the pretraining set and both evaluation sets, and discuss the impact on the numbers.
- [Sec. 4.2, Table 3] The Kinetics-600 FVD comparison does not state whether baseline FVDs use the same protocol as Video-GPT: first 3 frames given, 13 frames predicted, 500-video subset, same resolution/frame interval, and same FVD computation (e.g., I3D features and number of real videos). If the LVM, Seine, and Open-Sora-Plan numbers come from different setups, Table 3 does not support the claim of best FVD among vanilla-transformers. Please specify the protocol for every row and, if needed, recompute baselines under identical conditions.
minor comments (5)
- [Sec. 3.2, Training Target] The statement that Video-GPT predicts the video clip 'directly instead of noise or velocity' is potentially confusing, since predicting the clean sample under an L2 loss is mathematically equivalent to predicting velocity up to an affine scaling in flow matching; please add a sentence clarifying the parameterization.
- [Sec. 4.1, Image Animation paragraph] The sentence 'we colloct three datasets from the internet' contains a typo; it should read 'collect'.
- [Sec. 4.3, Tables 4-5] The ablation tables do not state the fixed values of the other hyperparameters (e.g., the CFG scale used in the clip-size rows of Table 4 is presumably 3.0, and the Table 5 rows presumably use Nk=16). Please state these explicitly so the reader can reconstruct the settings.
- [Sec. 3.3, Eq. (5)] The symbol DNS(k+1,:) denotes both the noisy input clip and the denoised output clip; please use different notation, such as NS_in and NS_out, to avoid ambiguity.
- [Sec. 4.4, Table 6] The UCF-101 comparison mixes resolutions (128x128 to 240x320) and architectures, and 'state-of-the-art at high resolution' is not a controlled apples-to-apples claim; please state the number of frames and resolution for each baseline or temper the claim.
Circularity Check
No significant circularity: the next clip diffusion objective is an independent training loss, and benchmark hyperparameter selection is a fairness issue, not a definitional reduction.
full rationale
Video-GPT's central derivation is the next clip diffusion pretraining objective: Eq. (1) is a standard flow-matching interpolation between a clean latent and noise, Eq. (4) arranges noisy and clean clips causally, and the hierarchical attention mask makes each noisy clip depend on previous clean clips while the training target is an L2 reconstruction of the clean clip. These are design choices with an independent per-clip reconstruction signal, not a definitional restatement of the benchmark claim. The reported Physics-IQ 34.97 and Kinetics-600 FVD 315.40 are empirical measurements of a trained 3.8B model, not quantities obtained by fitting the model to the evaluation metric. The inference ablations in Tables 4 and 5 select the CFG scale c=3.0 and inference clip size by maximizing Physics-IQ scores, which is a test-set selection concern that can inflate reported numbers; however, it is not a circular reduction because the score still depends on the pretrained model, the data, and the forward process, and no equation in the paper makes the benchmark value equal to the training loss by construction. The Limitations section (Sec. C) mentions only model scale and does not address possible pretraining/evaluation overlap, but that is a missing benchmark control rather than a circular argument. The paper contains several self-citations, including Seine [14], WeGen [30], V-stylist [93], and GetIn-1M [97], but none of these carries the central argument: the core next clip diffusion objective does not rest on a uniqueness theorem, an imported ansatz, or a fitted parameter from those works. Under the strict circularity standard requiring an explicit reduction of the claimed result to its inputs, the derivation is self-contained.
Assumptions & free parameters
free parameters (5)
- History-conditioned classifier-free guidance scale c =
3.0
- Inference frames per clip =
16 for most final results; ablation ranges 1 to 32
- Clean-clip noise retention beta =
0.9
- Pretraining frames sampled =
80 for final Stage 4
- Pretraining dataset scale =
70M videos (Panda-70M)
assumptions (4)
- standard math Flow matching with direct clean-clip regression under an L2 loss is a valid generative training objective.
- domain assumption The SDXL VAE latent space preserves the spatiotemporal information needed for video prediction.
- ad hoc to paper The hierarchical attention masks, especially bidirectional attention among noisy frames in a clip, do not leak future clean information.
- domain assumption Uncurated Panda-70M raw videos without annotation are sufficient self-supervised pretraining data.
Cite this review
Pith. "Pith review of Video-GPT via Next Clip Diffusion." pith.science (2026). https://pith.science/paper/SBFVWLWB
@misc{pith2026250512489,
author = {Pith},
title = {Pith review of: Video-GPT via Next Clip Diffusion},
year = {2026},
howpublished = {\url{https://pith.science/paper/SBFVWLWB}},
note = {Machine review of arXiv:2505.12489}
}
read the original abstract
GPT has shown its remarkable success in natural language processing. However, the language sequence is not sufficient to describe spatial-temporal details in the visual world. Alternatively, the video sequence is good at capturing such details. Motivated by this fact, we propose a concise Video-GPT in this paper by treating video as new language for visual world modeling. By analogy to next token prediction in GPT, we introduce a novel next clip diffusion paradigm for pretraining Video-GPT. Different from the previous works, this distinct paradigm allows Video-GPT to tackle both short-term generation and long-term prediction, by autoregressively denoising the noisy clip according to the clean clips in the history. Extensive experiments show our Video-GPT achieves the state-of-the-art performance on video prediction, which is the key factor towards world modeling (Physics-IQ Benchmark: Video-GPT 34.97 vs. Kling 23.64 vs. Wan 20.89). Moreover, it can be well adapted on 6 mainstream video tasks in both video generation and understanding, showing its great generalization capacity in downstream. The project page is at https://zhuangshaobin.github.io/Video-GPT.github.io/.
Figures
Figures from the paper (8 more)
Forward citations
Cited by 1 Pith paper
-
LoViC: Efficient Long Video Generation with Context Compression
LoViC uses FlexFormer, a single-query-token Q-Former with interpolated rotary positional encoding, to compress long video-text context for efficient long-video generation.
Reference graph
Works this paper leans on
-
[1]
M. Abdin, S. A. Jacobs, A. A. Awan, J. Aneja, A. Awadallah, H. H. Awadalla, N. Bach, A. Bahree, A. Bakhtiari, H. S. Behl, A. Benhaim, M. Bilenko, J. Bjorck, S. Bubeck, M. Cai, C. C. T. Mendes, W. Chen, V . Chaudhary, P. Chopra, A. D. Giorno, G. de Rosa, M. Dixon, R. Eldan, D. Iter, A. Goswami, S. Gunasekar, E. Haider, J. Hao, R. J. Hewett, J. Huynh, M. Ja...
arXiv 2024
-
[2]
O. J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Al- tenschmidt, S. Altman, S. Anadkat, R. Avila, I. Babuschkin, S. Balaji, V . Balcom, P. Baltescu, H. ing Bao, M. Bavarian, J. Belgum, I. Bello, J. Berdine, G. Bernadett-Shapiro, C. Berner, L. Bogdonoff, O. Boiko, M. laine Boyd, A.-L. Brakman, G. Brockman, T. Brooks, M....
2023
-
[3]
Y . Bai, X. Geng, K. Mangalam, A. Bar, A. L. Yuille, T. Darrell, J. Malik, and A. A. Efros. Sequential modeling enables scalable learning for large vision models. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 22861–22872, 2023
2024
-
[4]
M. Bain, A. Nagrani, G. Varol, and A. Zisserman. Frozen in time: A joint video and image encoder for end-to-end retrieval. 2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 1708–1718, 2021
2021
-
[5]
Bar-Tal, H
O. Bar-Tal, H. Chefer, O. Tov, C. Herrmann, R. Paiss, S. Zada, A. Ephrat, J. Hur, Y . Li, T. Michaeli, O. Wang, D. Sun, T. Dekel, and I. Mosseri. Lumiere: A space-time diffusion model for video generation. In ACM SIGGRAPH Conference and Exhibition on Computer Graphics and Interactive Techniques in Asia, 2024
2024
-
[6]
Betker, G
J. Betker, G. Goh, L. Jing, TimBrooks, J. Wang, L. Li, LongOuyang, JuntangZhuang, JoyceLee, YufeiGuo, WesamManassra, PrafullaDhariwal, CaseyChu, YunxinJiao, and A. Ramesh. Im- proving image generation with better captions. URL https://api.semanticscholar.org/ CorpusID:264403242
-
[7]
A. Blattmann, T. Dockhorn, S. Kulal, D. Mendelevitch, M. Kilian, and D. Lorenz. Stable video diffusion: Scaling latent video diffusion models to large datasets. ArXiv, abs/2311.15127, 2023
arXiv 2023
-
[8]
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-V oss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. teusz Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and ...
arXiv 2005
Show all 97 references
-
[9]
Bruce, M
J. Bruce, M. Dennis, A. Edwards, J. Parker-Holder, Y . Shi, E. Hughes, M. Lai, A. Mavalankar, R. Steigerwald, C. Apps, Y . Aytar, S. Bechtle, F. M. P. Behbahani, S. Chan, N. M. O. Heess, L. Gonzalez, S. Osindero, S. Ozair, S. Reed, J. Zhang, K. Zolna, J. Clune, N. de Freitas, ...
2024 arXiv
-
[10]
Carreira, E
J. Carreira, E. Noland, A. Banki-Horvath, C. Hillier, and A. Zisserman. A short note about kinetics-600. ArXiv, abs/1808.01340, 2018
2018 arXiv
-
[11]
B. Chen, D. M. Monso, Y . Du, M. Simchowitz, R. Tedrake, and V . Sitzmann. Diffusion forcing: Next-token prediction meets full-sequence diffusion. ArXiv, abs/2407.01392, 2024
2024 arXiv
-
[12]
J. Chen, J. Yu, C. Ge, L. Yao, E. Xie, Y . Wu, Z. Wang, J. T. Kwok, P. Luo, H. Lu, and Z. Li. Pixart-α: Fast training of diffusion transformer for photorealistic text-to-image synthesis. ArXiv, abs/2310.00426, 2023
2023 arXiv
-
[13]
T.-S. Chen, A. Siarohin, W. Menapace, E. Deyneka, H. wei Chao, B. E. Jeon, Y . Fang, H.-Y . Lee, J. Ren, M.-H. Yang, and S. Tulyakov. Panda-70m: Captioning 70m videos with multiple cross- modality teachers. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (C...
2024
-
[14]
X. Chen, Y . Wang, L. Zhang, S. Zhuang, X. Ma, J. Yu, Y . Wang, D. Lin, Y . Qiao, and Z. Liu. Seine: Short-to-long video diffusion model for generative transition and prediction. ArXiv, abs/2310.20700, 2023
2023 arXiv
-
[15]
Clark, J
A. Clark, J. Donahue, and K. Simonyan. Adversarial video generation on complex datasets. arXiv: Computer Vision and Pattern Recognition, 2019. 11
2019
-
[16]
C. Deng, D. Zhu, K. Li, S. Guang, and H. Fan. Causal diffusion transformers for generative modeling. ArXiv, abs/2412.12095, 2024
2024 arXiv
-
[17]
Dhariwal and A
P. Dhariwal and A. Nichol. Diffusion models beat gans on image synthesis. ArXiv, abs/2105.05233, 2021
2021 arXiv
-
[18]
Dubey, A
A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Yang, A. Fan, A. Goyal, A. S. Hartshorn, A. Yang, A. Mitra, A. Sravankumar, A. Korenev, A. Hinsvark, A. Rao, A. Zhang, A. Rodriguez, A. Gregerson, A. Spataru, B. tiste Rozière, B. Bir...
2024 arXiv
-
[19]
Esser, S
P. Esser, S. Kulal, A. Blattmann, R. Entezari, J. Muller, H. Saini, Y . Levi, D. Lorenz, A. Sauer, F. Boesel, D. Podell, T. Dockhorn, Z. English, K. Lacey, A. Goodwin, Y . Marek, and R. Rombach. Scaling rectified flow transformers for high-resolution image synthesis. ArXiv, ab...
2024 arXiv
-
[20]
S. Ge, T. Hayes, H. Yang, X. Yin, G. Pang, D. Jacobs, J.-B. Huang, and D. Parikh. Long video generation with time-agnostic vqgan and time-sensitive transformer. In European Conference on Computer Vision, 2022
2022
-
[21]
Goyal, P
P. Goyal, P. Dollár, R. B. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y . Jia, and K. He. Accurate, large minibatch sgd: Training imagenet in 1 hour. ArXiv, abs/1706.02677, 2017
2017 arXiv
-
[22]
Y . Gu, W. Mao, and M. Z. Shou. Long-context autoregressive video modeling with next-frame prediction. ArXiv, abs/2503.19325, 2025
2025 arXiv
-
[23]
T. Han, W. Xie, and A. Zisserman. Memory-augmented dense predictive coding for video representation learning. In European Conference on Computer Vision, 2020
2020
-
[24]
J. Ho. Classifier-free diffusion guidance. ArXiv, abs/2207.12598, 2022
2022 arXiv
-
[25]
J. Ho, A. Jain, and P. Abbeel. Denoising diffusion probabilistic models. In Conference and Workshop on Neural Information Processing Systems, 2020
2020
-
[26]
J. Ho, T. Salimans, A. Gritsenko, W. Chan, M. Norouzi, and D. J. Fleet. Video diffusion models. ArXiv, abs/2204.03458, 2022
2022 arXiv
-
[27]
W. Hong, M. Ding, W. Zheng, X. Liu, and J. Tang. Cogvideo: Large-scale pretraining for text-to-video generation via transformers. ArXiv, abs/2205.15868, 2022
2022 arXiv
-
[28]
J. Hu, S. Hu, Y . Song, Y . Huang, M. Wang, H. Zhou, Z. Liu, W.-Y . Ma, and M. Sun. Acdit: Inter- polating autoregressive conditional modeling and diffusion transformer. ArXiv, abs/2412.07720, 2024
2024
-
[29]
Hu and D
Z. Hu and D. Xu. Videocontrolnet: A motion-guided video-to-video translation framework by using diffusion model with controlnet. ArXiv, abs/2307.14073, 2023
2023 arXiv
-
[30]
Huang, S
Z. Huang, S. Zhuang, C. Fu, B. Yang, Y . Zhang, C. Sun, Z. Zhang, Y . Wang, C. Li, and Z.-J. Zha. Wegen: A unified model for interactive multimodal generation as we chat. ArXiv, abs/2503.01115, 2025
2025 arXiv
-
[31]
Kahembwe and S
E. Kahembwe and S. Ramamoorthy. Lower dimensional kernels for video discriminators. Neural networks : the official journal of the International Neural Network Society, 132:506–520, 2019
2019
-
[32]
Kaplan, S
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei. Scaling laws for neural language models. ArXiv, abs/2001.08361, 2020
2001 arXiv
-
[33]
T. Kim, G. Song, S. Lee, S. Kim, Y . Seo, S. Lee, S. H. Kim, H. Lee, and K. Bae. L-verse: Bidirectional generation between image and text. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 16505–16515, 2021. 13
2022
-
[34]
Kondratyuk, L
D. Kondratyuk, L. Yu, X. Gu, J. Lezama, J. Huang, R. Hornung, H. Adam, H. Akbari, Y . Alon, V . Birodkar, Y . Cheng, M.-C. Chiu, J. Dillon, I. Essa, A. Gupta, M. Hahn, A. Hauth, D. Hendon, A. Martinez, D. C. Minnen, D. A. Ross, G. Schindler, M. Sirotenko, K. Sohn, K. Somandepa...
2023 arXiv
-
[35]
W. Kong, Q. Tian, Z. Zhang, R. Min, Z. Dai, J. Zhou, J.-L. Xiong, X. Li, B. Wu, J. Zhang, K. Wu, Q. Lin, J. Yuan, Y . Long, A. Wang, A. Wang, C. Li, D. Huang, F. Yang, H. Tan, H. Wang, J. Song, J. Bai, J. Wu, J. Xue, J. Wang, K. Wang, M. Liu, P.-Y . Li, S. Li, W. Wang, W. Yu, ...
2024 arXiv
-
[36]
Kling ai, 2024
Kuaishou. Kling ai, 2024. URL https://klingai.kuaishou.com/
2024
-
[37]
B. F. Labs. Flux, 2024. URL https://blackforestlabs.ai/
2024
-
[38]
C. Li, D. Huang, Z. Lu, Y . Xiao, Q. Pei, and L. Bai. A survey on long video generation: Challenges, methods, and prospects. ArXiv, abs/2403.16407, 2024
2024 arXiv
-
[39]
C. Liao, L. Liu, X. Wang, Z. Luo, X. Zhang, W. Zhao, J. Wu, L. Li, Z. Tian, and W. Huang. Mo- gao: An omni foundation model for interleaved multi-modal generation. ArXiv, abs/2505.05472, 2025
2025 arXiv
-
[40]
B. Lin, Y . Ge, X. Cheng, Z. Li, B. Zhu, S. Wang, X. He, Y . Ye, S. Yuan, L. Chen, T. Jia, J. Zhang, Z. Tang, Y . Pang, B. She, C. Yan, Z. Hu, X. wen Dong, L. Chen, Z. Pan, X. Zhou, S. Dong, Y . Tian, and L. Yuan. Open-sora plan: Open-source large video generation model. ArXiv...
2024 arXiv
-
[41]
Lipman, R
Y . Lipman, R. T. Q. Chen, H. Ben-Hamu, M. Nickel, and M. Le. Flow matching for generative modeling. ArXiv, abs/2210.02747, 2022
2022 arXiv
-
[42]
H. Liu, W. Yan, M. Zaharia, and P. Abbeel. World model on million-length video and language with blockwise ringattention. ArXiv, abs/2402.08268, 2024
2024 arXiv
-
[43]
Loshchilov and F
I. Loshchilov and F. Hutter. Decoupled weight decay regularization. InInternational Conference on Learning Representations, 2017
2017
-
[44]
Z. Luo, D. Chen, Y . Zhang, Y . Huang, L. Wang, Y . Shen, D. Zhao, J. Zhou, and T.-P. Tan. Videofusion: Decomposed diffusion models for high-quality video generation. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10209–10218, 2023
2023
-
[45]
X. Ma, Y . Wang, G. Jia, X. Chen, Z. Liu, Y .-F. Li, C. Chen, and Y . Qiao. Latte: Latent diffusion transformer for video generation. ArXiv, abs/2401.03048, 2024
2024 arXiv
-
[46]
Menapace, A
W. Menapace, A. Siarohin, I. Skorokhodov, E. Deyneka, T.-S. Chen, A. Kag, Y . Fang, A. Stoliar, E. Ricci, J. Ren, and S. Tulyakov. Snap video: Scaled spatiotemporal transformers for text- to-video synthesis. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (...
2024
-
[47]
Motamed, L
S. Motamed, L. Culp, K. Swersky, P. Jaini, and R. Geirhos. Do generative video models understand physical principles? ArXiv, abs/2501.09038, 2025
2025 arXiv
-
[48]
K. Nan, R. Xie, P. Zhou, T. Fan, Z. Yang, Z. Chen, X. Li, J. Yang, and Y . Tai. Openvid-1m: A large-scale high-quality dataset for text-to-video generation. ArXiv, abs/2407.02371, 2024
2024 arXiv
-
[49]
Video generation models as world simulators, 2024
OpenAI. Video generation models as world simulators, 2024. URL https://openai.com/ index/video-generation-models-as-world-simulators/
2024
-
[50]
Ouyang, J
Y . Ouyang, J. Yuan, H. Zhao, G. Wang, and B. Zhao. Flexifilm: Long video generation with flexible conditions. ArXiv, abs/2404.18620, 2024
2024 arXiv
-
[51]
Parmar, A
N. Parmar, A. Vaswani, J. Uszkoreit, L. Kaiser, N. M. Shazeer, A. Ku, and D. Tran. Image transformer. In International Conference on Machine Learning, 2018. 14
2018
-
[52]
Patrick, P.-Y
M. Patrick, P.-Y . B. Huang, Y . M. Asano, F. Metze, A. Hauptmann, J. F. Henriques, and A. Vedaldi. Support-set bottlenecks for video-text representation learning. ArXiv, abs/2010.02824, 2020
2010 arXiv
-
[53]
W. S. Peebles and S. Xie. Scalable diffusion models with transformers. 2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 4172–4182, 2022
2023
-
[54]
Pika 1.5, 2024
PikaLabs. Pika 1.5, 2024. URL https://pika.art/
2024
-
[55]
Podell, Z
D. Podell, Z. English, K. Lacey, A. Blattmann, T. Dockhorn, J. Muller, J. Penna, and R. Rom- bach. Sdxl: Improving latent diffusion models for high-resolution image synthesis. ArXiv, abs/2307.01952, 2023
2023 arXiv
-
[56]
Polyak, A
A. Polyak, A. Zohar, A. Brown, A. Tjandra, A. Sinha, A. Lee, A. Vyas, B. Shi, C.-Y . Ma, C.-Y . Chuang, D. Yan, D. Choudhary, D. Wang, G. Sethi, G. Pang, H. Ma, I. Misra, J. Hou, J. Wang, K. ran Jagadeesh, K. Li, L. Zhang, M. Singh, M. Williamson, M. Le, M. Yu, M. K. Singh, P....
-
[57]
Radford and K
A. Radford and K. Narasimhan. Improving language understanding by generative pre-training. 2018
2018
-
[58]
Radford, J
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever. Language models are unsupervised multitask learners. 2019
2019
-
[59]
Radford, J
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, 2021
2021
-
[60]
Ramesh, P
A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen. Hierarchical text-conditional image generation with clip latents. ArXiv, abs/2204.06125, 2022
2022 arXiv
-
[61]
Z. Ren, Y . Wei, X. Guo, Y . Zhao, B. Kang, J. Feng, and X. Jin. Videoworld: Exploring knowledge learning from unlabeled videos. ArXiv, abs/2501.09781, 2025
2025 arXiv
-
[62]
Rombach, A
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer. High-resolution image synthesis with latent diffusion models. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10674–10685, 2021
2022
-
[63]
Ronneberger, P
O. Ronneberger, P. Fischer, and T. Brox. U-net: Convolutional networks for biomedical image segmentation. ArXiv, abs/1505.04597, 2015
2015 arXiv
-
[64]
Gen-3, 2024
Runway. Gen-3, 2024. URL https://runwayml.com/
2024
-
[65]
Saharia, W
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, S. K. S. Ghasemipour, B. K. Ayan, S. S. Mahdavi, R. G. Lopes, T. Salimans, J. Ho, D. J. Fleet, and M. Norouzi. Photorealistic text-to-image diffusion models with deep language understanding. ArXiv, abs/2205.11487, 2022
2022 arXiv
-
[66]
Saito, S
M. Saito, S. Saito, M. Koyama, and S. Kobayashi. Train sparsely, generate densely: Memory- efficient unsupervised training of high-resolution temporal gan. International Journal of Computer Vision, 128:2586 – 2606, 2018
2018
-
[67]
Segalis, D
E. Segalis, D. Valevski, D. Lumen, Y . Matias, and Y . Leviathan. A picture is worth a thousand words: Principled recaptioning improves image generation. ArXiv, abs/2310.16656, 2023. 15
2023 arXiv
-
[68]
Singer, A
U. Singer, A. Polyak, T. Hayes, X. Yin, J. An, S. Zhang, Q. Hu, H. Yang, O. Ashual, O. Gafni, D. Parikh, S. Gupta, and Y . Taigman. Make-a-video: Text-to-video generation without text-video data. ArXiv, abs/2209.14792, 2022
2022 arXiv
-
[69]
Y . Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole. Score-based generative modeling through stochastic differential equations. In International Conference on Machine Learning, 2021
2021
-
[70]
Soomro, A
K. Soomro, A. Zamir, and M. Shah. Ucf101: A dataset of 101 human actions classes from videos in the wild. ArXiv, abs/1212.0402, 2012
2012 arXiv
-
[71]
P. Sun, Y . Jiang, S. Chen, S. Zhang, B. Peng, P. Luo, and Z. Yuan. Autoregressive model beats diffusion: Llama for scalable image generation. ArXiv, abs/2406.06525, 2024
2024 arXiv
-
[72]
C. Team. Chameleon: Mixed-modal early-fusion foundation models. ArXiv, abs/2405.09818, 2024
2024 arXiv
-
[73]
Unterthiner, S
T. Unterthiner, S. van Steenkiste, K. Kurach, R. Marinier, M. Michalski, and S. Gelly. Towards accurate generative models of video: A new metric & challenges. ArXiv, abs/1812.01717, 2018
2018 arXiv
-
[74]
A. Wang, B. Ai, B. Wen, C. Mao, C.-W. Xie, D. Chen, F. Yu, H. Zhao, J. Yang, J. Zeng, J. Wang, J. Zhang, J. Zhou, J. Wang, J. Chen, K. Zhu, K. Zhao, K. Yan, L. Huang, X. Meng, N. Zhang, P. Li, P. Wu, R. Chu, R. Feng, S. Zhang, S. Sun, T. Fang, T. Wang, T. Gui, T. Weng, T. Shen...
2025 arXiv
-
[75]
H. Wang, S. Suri, Y . Ren, H. Chen, and A. Shrivastava. Larp: Tokenizing videos with a learned autoregressive generative prior. ArXiv, abs/2410.21264, 2024
2024 arXiv
-
[76]
J. Wang, Y . Jiang, Z. Yuan, B. Peng, Z. Wu, and Y .-G. Jiang. Omnitokenizer: A joint image- video tokenizer for visual generation. ArXiv, abs/2406.09399, 2024
2024 arXiv
-
[77]
L. Wang, B. Huang, Z. Zhao, Z. Tong, Y . He, Y . Wang, Y . Wang, and Y . Qiao. Videomae v2: Scaling video masked autoencoders with dual masking. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14549–14560, 2023
2023
-
[78]
Q. Wang, Y . Shi, J. Ou, R. Chen, K. Lin, J. Wang, B. Jiang, H. Yang, M. Zheng, X. Tao, F. Yang, P. Wan, and D. Zhang. Koala-36m: A large-scale video dataset improving consistency between fine-grained conditions and video content. ArXiv, abs/2410.08260, 2024
-
[79]
X. Wang, X. Zhang, Z. Luo, Q. Sun, Y . Cui, J. Wang, F. Zhang, Y . Wang, Z. Li, Q. Yu, Y . Zhao, Y . Ao, X. Min, T. Li, B. Wu, B. Zhao, B. Zhang, L. zi Wang, G. Liu, Z. He, X. Yang, J. Liu, Y . Lin, T. Huang, and Z. Wang. Emu3: Next-token prediction is all you need. ArXiv, abs...
2024 arXiv
-
[80]
Y . Wang, X. Chen, X. Ma, S. Zhou, Z. Huang, Y . Wang, C. Yang, Y . He, J. Yu, P. der Yang, Y . Guo, T. Wu, C. Si, Y . Jiang, C. Chen, C. C. Loy, B. Dai, D. Lin, Y . Qiao, and Z. Liu. Lavie: High-quality video generation with cascaded latent diffusion models. ArXiv, abs/2309.1...
2023 arXiv
-
[81]
Y . Wang, Y . He, Y . Li, K. Li, J. Yu, X. J. Ma, X. Chen, Y . Wang, P. Luo, Z. Liu, Y . Wang, L. Wang, and Y . Qiao. Internvid: A large-scale video-text dataset for multimodal understanding and generation. ArXiv, abs/2307.06942, 2023
2023 arXiv
-
[82]
Q. Wu, H. Ye, Y . Gu, H. Zhang, L. Wang, and D. He. Denoising masked autoencoders are certifiable robust vision learners. ArXiv, abs/2210.06983, 2022
2022 arXiv
-
[83]
S. Xiao, Y . Wang, J. Zhou, H. Yuan, X. Xing, R. Yan, S. Wang, T. Huang, and Z. Liu. Omnigen: Unified image generation. ArXiv, abs/2409.11340, 2024
2024 arXiv
-
[84]
H. Xu, G. Ghosh, P.-Y . B. Huang, D. Okhonko, A. Aghajanyan, and F. M. L. Z. C. Feichtenhofer. Videoclip: Contrastive pre-training for zero-shot video-text understanding. In Conference on Empirical Methods in Natural Language Processing, 2021. 16
2021
-
[85]
J. Xu, T. Mei, T. Yao, and Y . Rui. Msr-vtt: A large video description dataset for bridging video and language. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 5288–5296, 2016
2016
-
[86]
L. Xue, N. Constant, A. Roberts, M. Kale, R. Al-Rfou, A. Siddhant, A. Barua, and C. Raffel. mt5: A massively multilingual pre-trained text-to-text transformer. In North American Chapter of the Association for Computational Linguistics, 2020
2020
-
[87]
A. Yang, B. Yang, B. Hui, B. Zheng, B. Yu, C. Zhou, C. Li, C. Li, D. Liu, F. Huang, G. Dong, H. Wei, H. Lin, J. Tang, J. Wang, J. Yang, J. Tu, J. Zhang, J. Ma, J. Xu, J. Zhou, J. Bai, J. He, J. Lin, K. Dang, K. Lu, K.-Y . Chen, K. Yang, M. Li, M. Xue, N. Ni, P. Zhang, P. Wang,...
2024 arXiv
-
[88]
S. Yang, J. Walker, J. Parker-Holder, Y . Du, J. Bruce, A. Barreto, P. Abbeel, and D. Schuurmans. Video as the new language for real-world decision making. ArXiv, abs/2402.17139, 2024
2024 arXiv
-
[89]
Z. Yang, J. Teng, W. Zheng, M. Ding, S. Huang, J. Xu, Y . Yang, W. Hong, X. Zhang, G. Feng, D. Yin, X. Gu, Y . Zhang, W. Wang, Y . Cheng, T. Liu, B. Xu, Y . Dong, and J. Tang. Cogvideox: Text-to-video diffusion models with an expert transformer. ArXiv, abs/2408.06072, 2024
2024 arXiv
-
[90]
H. Yi, S. Shao, T. Ye, J. Zhao, Q. Yin, M. Lingelbach, L. Yuan, Y . Tian, E. Xie, and D. Zhou. Magic 1-for-1: Generating one minute video clips within one minute. ArXiv, abs/2502.07701, 2025
2025 arXiv
-
[91]
T. Yin, Q. Zhang, R. Zhang, W. T. Freeman, F. Durand, E. Shechtman, and X. Huang. From slow bidirectional to fast causal video generators. ArXiv, abs/2412.07772, 2024
2024
-
[92]
L. Yu, J. Lezama, N. B. Gundavarapu, L. Versari, K. Sohn, D. Minnen, Y . Cheng, A. Gupta, X. Gu, A. G. Hauptmann, B. Gong, M.-H. Yang, I. Essa, D. A. Ross, and L. Jiang. Language model beats diffusion - tokenizer is key to visual generation. In International Conference on Lear...
2024
-
[93]
Z. Yue, S. Zhuang, K. Li, Y . Ding, and Y . Wang. V-stylist: Video stylization via collaboration and reflection of mllm agents. ArXiv, abs/2503.12077, 2025
2025 arXiv
-
[94]
C. Zhou, L. Yu, A. Babu, K. Tirumala, M. Yasunaga, L. Shamis, J. Kahn, X. Ma, L. Zettlemoyer, and O. Levy. Transfusion: Predict the next token and diffuse images with one multi-modal model. ArXiv, abs/2408.11039, 2024
2024 arXiv
-
[95]
D. Zhou, Q. Sun, Y . Peng, K. Yan, R. Dong, D. Wang, Z. Ge, N. Duan, X. Zhang, L. M. Ni, and H. yeung Shum. Taming teacher forcing for masked autoregressive video generation. ArXiv, abs/2501.12389, 2025
2025 arXiv
-
[96]
Zhuang, K
S. Zhuang, K. Li, X. Chen, Y . Wang, Z. Liu, Y . Qiao, and Y . Wang. Vlogger: Make your dream a vlog. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 8806–8817, 2024
2024
-
[97]
noisy_img3<diff>𝛼
S. Zhuang, Z. Huang, B. Yang, Y . Zhang, F. Wang, C. Fu, C. Sun, Z.-J. Zha, C. Li, and Y . Wang. Get in video: Add anything you want to the video. ArXiv, abs/2503.06268, 2025. 17 A More Implementation Details Table 9: Pretraining Stage 1 setting. config Panda-70M optimizer Ada...
2025 arXiv
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.