Pith. sign in

REVIEW 4 major objections 5 minor 1 cited by

Video-GPT via Next Clip Diffusion

T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read The paper claims that next clip diffusion pretraining — denoising each new video clip only from the clean clips before it — produces state-of-the-art video prediction and a general video foundation model.

desk verdict Genuinely new clip-level pretraining paradigm with a broad empirical program, but the SOTA benchmark claims are not yet clean enough to verify—needs dedup, a fixed evaluation protocol, and baseline normalization. read the letter →

arxiv 2505.12489 v2 pith:SBFVWLWB submitted 2025-05-18 cs.CV cs.AI

classification cs.CVcs.AI
keywords videogenerationpredictionnextclipdiffusionautoregressivepretrainingworldmodelsself-supervisedlearningfoundation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that video can be treated as a language for world modeling, with each short clip playing the role of a word. Its proposal, next clip diffusion, pretrains a vanilla transformer by adding noise to some clips of a video and asking the model to reconstruct each noisy clip while attending only to the clean clips that came before — the video analogue of GPT's next-token prediction. If correct, a single self-supervised objective on unlabeled video supports both short-term generation and long-term prediction without any text annotation. The paper reports state-of-the-art results on Physics-IQ (34.97 versus 29.50 for VideoPoet) and Kinetics-600 (FVD 315.40), and shows that fine-tuning the same checkpoint transfers to six downstream video generation and understanding tasks.

What carries the argument

The load-bearing object is the noise-clean interleaved clip sequence under a hierarchical causal mask. Each video is split into clips; some clips are noised by flow matching, and the input alternates noisy clip, clean clip, noisy clip, clean clip in temporal order, with boundary tokens (<diff>, <img>) and the noise weight alpha feeding the timestep information. The clip-level mask lets the k-th noisy clip attend to itself and to all earlier clean clips but not to earlier noisy clips, which forces denoising to be conditioned on the correct history; frame-level and patch-level masks then decide bidirectional versus causal attention inside frames and patches. This is what lets diffusion operate inside a clip while GPT-style autoregression operates across clips, and it is what the paper credits for combining short-term generation quality with long-range prediction.

What would settle it

Fix the inference clip size and history-conditioned guidance scale without consulting test scores, or run the same pretrained weights on a held-out physics benchmark; if the margin over VideoPoet narrows to noise, the claim that next clip diffusion causes the improvement fails. Alternatively, pretrain an identical-architecture baseline with full-sequence diffusion or next-frame autoregression on the same data: matching Physics-IQ would also falsify the paradigm's specific role.

Watch

Extended reading notes

Core claim

The central claim is that next clip diffusion — autoregressively denoising each new clip against the already-clean clips in its past — is a sufficient pretraining objective for a general video model. The paper argues that earlier diffusion video models generate within the whole clip or sequence but struggle with long-term prediction, while pure autoregressive token predictors lag behind diffusion in synthesis quality; interleaving noisy and clean clips in temporal order lets the same transformer do both. Concretely, the model is trained with an L2 loss that predicts the clean clip directly, uses flow matching noise schedules, and applies a hierarchical attention mask at clip, frame, and patch levels so that noisy clips see only past clean context while patch tokens inside a frame attend fully. The authors report that this pretraining, on 70 million unlabeled videos, outperforms prior video prediction systems on deterministic physical futures and on uncertain human motion, and that the same weights adapt to class-to-video, text-to-video, image animation, video classification, retrieval, and object segmentation.

Load-bearing premise

The headline results assume the Physics-IQ and Kinetics-600 comparisons measure the method fairly, so the reported gains come from next clip diffusion and not from inference or pretraining choices tuned on those same benchmarks.

Editorial extensions

If this is right

  • One unlabeled-video objective can serve as foundation pretraining for both generative and understanding tasks, so video-only data becomes a viable scaling axis.
  • Longer pretraining context and larger unlabeled corpora directly improve physical prediction, suggesting the paradigm follows a data-scaling trend like language models.
  • The model can generate arbitrarily long futures by iteratively treating each denoised clip as clean history, which is exactly the GPT analogy carried to inference.
  • A plain vanilla transformer without U-Net or DiT designs reaches competitive FVD on Kinetics-600, so architectural complexity is not what drives the result.
  • Downstream tasks transfer with small fine-tuning sets — fewer than one hundred videos for image animation — indicating the pretrained representation is broadly reusable.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the Physics-IQ margin is genuine, next clip diffusion may be a better prior for learning intuitive physics than full-sequence diffusion, because causal clean-history conditioning matches how physical dynamics unfold; a direct comparison holding architecture and data fixed would test that.
  • The same interleaved noising recipe could transfer to other continuous temporal modalities such as audio or sensor streams, where clean history is also available and next-step prediction defines the task.
  • Because the inference clip size and guidance scale were selected on the evaluation benchmark itself, part of the reported margin could come from benchmark alignment rather than the paradigm; an independent held-out physics suite would separate the two.
  • If it extends, combining next clip tokens with language tokens in one transformer could yield a true multimodal world model that narrates and predicts video jointly.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper introduces Video-GPT, a 3.8B-parameter transformer for video generation and understanding, trained with a proposed 'next clip diffusion' objective. A video is partitioned into clips; each clip is noised via flow matching, and the model must denoise each noisy clip conditioned only on the previous clean clips through hierarchical causal attention masks. The pretrained model is evaluated zero-shot on Physics-IQ (34.97) and Kinetics-600 (FVD 315.40) for deterministic and stochastic video prediction, and is fine-tuned on six tasks: class-to-video, text-to-video, image animation, video classification, video retrieval, and video object segmentation. The paper reports state-of-the-art performance on Physics-IQ and Kinetics-600 and favorable results on the downstream tasks.

Significance. If the claims hold, the main contribution is a single self-supervised objective on unlabeled video that unifies autoregressive and diffusion modeling at the clip level and transfers to both generation and understanding tasks, without requiring text annotations for pretraining. The method is clearly described, the progressive training strategy and ablations are useful, and the qualitative results are plausible. The central risk is not internal inconsistency but the benchmark evidence: the headline scores are produced after tuning inference hyperparameters on the same benchmark, and the comparisons lack protocol-equivalence and deduplication controls, so the SOTA claim is not independently verifiable as presented.

major comments (4)
  1. [Sec. 4.3, Table 4] The history-conditioned CFG scale c=3.0 and the inference clip size Nk are selected by maximizing the Physics-IQ score in Table 4, and the same selection is then reported as the headline result (34.97) in Table 2 and the abstract. This means the reported SOTA number is a test-set-tuned result, not an independent evaluation. Please report a validation-split or fixed-protocol evaluation, or at least show the score without this tuning, and clearly state the tuning procedure in the main text.
  2. [Sec. 4.2, Table 2] Table 2 mixes image-to-video (I2V) and video-to-video (V2V) baselines, models with and without text conditioning, and methods whose condition-frame length and prediction length are not stated to match the Physics-IQ protocol (3 s condition, 5 s prediction). Since several numbers are taken from external papers while only LVM, Open-Sora-Plan, and Seine are re-tested by the authors, the comparison does not establish that all models received the same input conditions and were evaluated on the same output horizon. Without protocol equivalence, the 5-point gap over VideoPoet cannot be attributed to the next-clip-diffusion paradigm.
  3. [Sec. 4.1, Sec. 4.2, Table 3] Pretraining is performed on Panda-70M, and evaluation on Kinetics-600 (and Physics-IQ) with no reported deduplication. Panda-70M aggregates large-scale internet video collections and may contain Kinetics-600 or Physics-IQ clips or near-duplicates; if so, the reported FVD 315.40 and Physics-IQ 34.97 partly reflect memorization rather than generalization. The Limitations (Sec. C) mention only model scale, not this risk. Please perform and report a hash-based or feature-based overlap analysis between the pretraining set and both evaluation sets, and discuss the impact on the numbers.
  4. [Sec. 4.2, Table 3] The Kinetics-600 FVD comparison does not state whether baseline FVDs use the same protocol as Video-GPT: first 3 frames given, 13 frames predicted, 500-video subset, same resolution/frame interval, and same FVD computation (e.g., I3D features and number of real videos). If the LVM, Seine, and Open-Sora-Plan numbers come from different setups, Table 3 does not support the claim of best FVD among vanilla-transformers. Please specify the protocol for every row and, if needed, recompute baselines under identical conditions.
minor comments (5)
  1. [Sec. 3.2, Training Target] The statement that Video-GPT predicts the video clip 'directly instead of noise or velocity' is potentially confusing, since predicting the clean sample under an L2 loss is mathematically equivalent to predicting velocity up to an affine scaling in flow matching; please add a sentence clarifying the parameterization.
  2. [Sec. 4.1, Image Animation paragraph] The sentence 'we colloct three datasets from the internet' contains a typo; it should read 'collect'.
  3. [Sec. 4.3, Tables 4-5] The ablation tables do not state the fixed values of the other hyperparameters (e.g., the CFG scale used in the clip-size rows of Table 4 is presumably 3.0, and the Table 5 rows presumably use Nk=16). Please state these explicitly so the reader can reconstruct the settings.
  4. [Sec. 3.3, Eq. (5)] The symbol DNS(k+1,:) denotes both the noisy input clip and the denoised output clip; please use different notation, such as NS_in and NS_out, to avoid ambiguity.
  5. [Sec. 4.4, Table 6] The UCF-101 comparison mixes resolutions (128x128 to 240x320) and architectures, and 'state-of-the-art at high resolution' is not a controlled apples-to-apples claim; please state the number of frames and resolution for each baseline or temper the claim.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the next clip diffusion objective is an independent training loss, and benchmark hyperparameter selection is a fairness issue, not a definitional reduction.

full rationale

Video-GPT's central derivation is the next clip diffusion pretraining objective: Eq. (1) is a standard flow-matching interpolation between a clean latent and noise, Eq. (4) arranges noisy and clean clips causally, and the hierarchical attention mask makes each noisy clip depend on previous clean clips while the training target is an L2 reconstruction of the clean clip. These are design choices with an independent per-clip reconstruction signal, not a definitional restatement of the benchmark claim. The reported Physics-IQ 34.97 and Kinetics-600 FVD 315.40 are empirical measurements of a trained 3.8B model, not quantities obtained by fitting the model to the evaluation metric. The inference ablations in Tables 4 and 5 select the CFG scale c=3.0 and inference clip size by maximizing Physics-IQ scores, which is a test-set selection concern that can inflate reported numbers; however, it is not a circular reduction because the score still depends on the pretrained model, the data, and the forward process, and no equation in the paper makes the benchmark value equal to the training loss by construction. The Limitations section (Sec. C) mentions only model scale and does not address possible pretraining/evaluation overlap, but that is a missing benchmark control rather than a circular argument. The paper contains several self-citations, including Seine [14], WeGen [30], V-stylist [93], and GetIn-1M [97], but none of these carries the central argument: the core next clip diffusion objective does not rest on a uniqueness theorem, an imported ansatz, or a fitted parameter from those works. Under the strict circularity standard requiring an explicit reduction of the claimed result to its inputs, the derivation is self-contained.

Assumptions & free parameters 5 free parameters · 4 assumptions · 0 invented entities

No invented physical entities are introduced; the only new artifacts are model tokens (<diff>, <img>) and masking rules. Free parameters are mostly inference and training settings selected on the Physics-IQ benchmark. The central method is empirical, so the axiomatic assumptions listed are design and data assumptions rather than mathematical premises.

free parameters (5)
  • History-conditioned classifier-free guidance scale c = 3.0
    Selected by maximizing Physics-IQ in Sec. 4.3 (Table 4) and used for the final reported Physics-IQ and downstream numbers.
  • Inference frames per clip = 16 for most final results; ablation ranges 1 to 32
    Ablated on Physics-IQ (Table 4); larger clips improve score, and the final Stage 4 model uses a larger clip configuration for the 34.97 result.
  • Clean-clip noise retention beta = 0.9
    Eq. (6) hyperparameter for adding slight noise to clean clips during pretraining; chosen by ablation on Physics-IQ (Table 5).
  • Pretraining frames sampled = 80 for final Stage 4
    Progressive training schedule goes from 16 to 80 frames; Table 5 shows 80 frames gives 34.94 versus 33.09 at 48 frames.
  • Pretraining dataset scale = 70M videos (Panda-70M)
    Ablation of 1M versus 70M videos changes Physics-IQ from 23.16 to 33.09 (Stage 3 model), showing the final claim depends on this scale choice.
assumptions (4)
  • standard math Flow matching with direct clean-clip regression under an L2 loss is a valid generative training objective.
    Sec. 3.2 uses Eq. (1) plus an L2 loss that regresses the clean clip directly rather than noise or velocity; this relies on standard denoising and flow-matching convergence assumptions.
  • domain assumption The SDXL VAE latent space preserves the spatiotemporal information needed for video prediction.
    Sec. 4.1 uses the SDXL VAE; the paper provides no analysis of temporal consistency loss or reconstruction error in the latent space.
  • ad hoc to paper The hierarchical attention masks, especially bidirectional attention among noisy frames in a clip, do not leak future clean information.
    The mask in Fig. 2b allows all frames in a noisy clip to attend to each other while conditioning on past clean clips; this design choice is validated only indirectly by benchmark results.
  • domain assumption Uncurated Panda-70M raw videos without annotation are sufficient self-supervised pretraining data.
    Sec. 4.1 pretrains on 70M videos; no filtering or data-quality analysis is given, and the paper's own scale ablation shows a strong dependence on dataset size.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Video-GPT via Next Clip Diffusion." pith.science (2026). https://pith.science/paper/SBFVWLWB

@misc{pith2026250512489,
  author       = {Pith},
  title        = {Pith review of: Video-GPT via Next Clip Diffusion},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/SBFVWLWB}},
  note         = {Machine review of arXiv:2505.12489}
}
read the original abstract

GPT has shown its remarkable success in natural language processing. However, the language sequence is not sufficient to describe spatial-temporal details in the visual world. Alternatively, the video sequence is good at capturing such details. Motivated by this fact, we propose a concise Video-GPT in this paper by treating video as new language for visual world modeling. By analogy to next token prediction in GPT, we introduce a novel next clip diffusion paradigm for pretraining Video-GPT. Different from the previous works, this distinct paradigm allows Video-GPT to tackle both short-term generation and long-term prediction, by autoregressively denoising the noisy clip according to the clean clips in the history. Extensive experiments show our Video-GPT achieves the state-of-the-art performance on video prediction, which is the key factor towards world modeling (Physics-IQ Benchmark: Video-GPT 34.97 vs. Kling 23.64 vs. Wan 20.89). Moreover, it can be well adapted on 6 mainstream video tasks in both video generation and understanding, showing its great generalization capacity in downstream. The project page is at https://zhuangshaobin.github.io/Video-GPT.github.io/.

Figures

Figures reproduced from arXiv: 2505.12489 by the authors.

Figure 1
Figure 1. Next clip diffusion. We draw an analogy with GPT’s next token prediction and model each video clip as a visual word by denoising the next noisy clip, conditioning on the previous video. Although such a paradigm has achieved significant progress [5, 35, 36, 74], it often suffers from difficulty in long-term future prediction that is a critical factor of world modeling [50, 38]. To address this problem, the autoregres… view at source ↗
Figure 2
Figure 2. Video-GPT pretraining framework. The full attention mask is shown in [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Video-GPT inference framework. We iteratively denoise the 2nd noisy clip NS(2, :) to its clean version CL(2, :), and use it along with the 1st clean clip CL(1, :) to condition the prediction of the 3rd noisy clip NS(3, :). The number of frames in each clip can also vary during inference. where Φ(k, i) is the latent feature of the i-th frame in the k-th clip, and Ψ(k, i) is the noisy feature of this frame. The weight… view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: Fine-tuning Video-GPT on downstream tasks. (b), for the i-th frame in the k-th clean clip, it depends on itself, the (i − 1) frames in this clean clip, and all the frames in the previous (k − 1) clean clips. 2) Noisy Frame Mask. Alternatively, for the i-th frame in the…
Figure 5
Figure 5. Figure 5: Qualitative results on Physics-IQ Benchmark. The videos predicted by our Video-GPT based on condition frames are more consistent with physical laws than other methods [14, 3] [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 6
Figure 6. Figure 6: Qualitative results of Video-GPT on class-to-video generation on UCF-101 [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 7
Figure 7. Figure 7: Qualitative results of Video-GPT on text-to video-generation. <lie down> <transform> <fly> [PITH_FULL_IMAGE:figures/full_fig_p009_7.png]
Figure 9
Figure 9. Figure 9: Qualitative results of Video-GPT on video object segmentation. The frames in blue box are condition frames. The frame in red box is the 1-st frame of the condition with object mask, and we circle the segmented object in red. Frames in green box are the generated segmen…
Figure 10
Figure 10. Figure 10: The full attention mask of [PITH_FULL_IMAGE:figures/full_fig_p020_10.png]
Figure 11
Figure 11. Figure 11: The long video zero-shot prediction result generated by Video-GPT. The frames in blue box are condition frames. Frames in green box are the generated prediction result. 21 [PITH_FULL_IMAGE:figures/full_fig_p021_11.png]
Figure 12
Figure 12. Figure 12: The long video zero-shot prediction result generated by Video-GPT. The frames in blue box are condition frames. Frames in green box are the generated prediction result. 22 [PITH_FULL_IMAGE:figures/full_fig_p022_12.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LoViC: Efficient Long Video Generation with Context Compression

    cs.CV 2025-07 conditional novelty 6.0 of 10

    LoViC uses FlexFormer, a single-query-token Q-Former with interpolated rotary positional encoding, to compress long video-text context for efficient long-video generation.

Reference graph

Works this paper leans on

97 extracted references · 17 canonical work pages · cited by 1 Pith paper

  1. [1]

    Abdin, S

    M. Abdin, S. A. Jacobs, A. A. Awan, J. Aneja, A. Awadallah, H. H. Awadalla, N. Bach, A. Bahree, A. Bakhtiari, H. S. Behl, A. Benhaim, M. Bilenko, J. Bjorck, S. Bubeck, M. Cai, C. C. T. Mendes, W. Chen, V . Chaudhary, P. Chopra, A. D. Giorno, G. de Rosa, M. Dixon, R. Eldan, D. Iter, A. Goswami, S. Gunasekar, E. Haider, J. Hao, R. J. Hewett, J. Huynh, M. Ja...

  2. [2]

    O. J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Al- tenschmidt, S. Altman, S. Anadkat, R. Avila, I. Babuschkin, S. Balaji, V . Balcom, P. Baltescu, H. ing Bao, M. Bavarian, J. Belgum, I. Bello, J. Berdine, G. Bernadett-Shapiro, C. Berner, L. Bogdonoff, O. Boiko, M. laine Boyd, A.-L. Brakman, G. Brockman, T. Brooks, M....

  3. [3]

    Y . Bai, X. Geng, K. Mangalam, A. Bar, A. L. Yuille, T. Darrell, J. Malik, and A. A. Efros. Sequential modeling enables scalable learning for large vision models. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 22861–22872, 2023

  4. [4]

    M. Bain, A. Nagrani, G. Varol, and A. Zisserman. Frozen in time: A joint video and image encoder for end-to-end retrieval. 2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 1708–1718, 2021

  5. [5]

    Bar-Tal, H

    O. Bar-Tal, H. Chefer, O. Tov, C. Herrmann, R. Paiss, S. Zada, A. Ephrat, J. Hur, Y . Li, T. Michaeli, O. Wang, D. Sun, T. Dekel, and I. Mosseri. Lumiere: A space-time diffusion model for video generation. In ACM SIGGRAPH Conference and Exhibition on Computer Graphics and Interactive Techniques in Asia, 2024

  6. [6]

    Betker, G

    J. Betker, G. Goh, L. Jing, TimBrooks, J. Wang, L. Li, LongOuyang, JuntangZhuang, JoyceLee, YufeiGuo, WesamManassra, PrafullaDhariwal, CaseyChu, YunxinJiao, and A. Ramesh. Im- proving image generation with better captions. URL https://api.semanticscholar.org/ CorpusID:264403242

  7. [7]

    Blattmann, T

    A. Blattmann, T. Dockhorn, S. Kulal, D. Mendelevitch, M. Kilian, and D. Lorenz. Stable video diffusion: Scaling latent video diffusion models to large datasets. ArXiv, abs/2311.15127, 2023

  8. [8]

    T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-V oss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. teusz Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and ...

Show all 97 references
  1. [9]

    Bruce, M

    J. Bruce, M. Dennis, A. Edwards, J. Parker-Holder, Y . Shi, E. Hughes, M. Lai, A. Mavalankar, R. Steigerwald, C. Apps, Y . Aytar, S. Bechtle, F. M. P. Behbahani, S. Chan, N. M. O. Heess, L. Gonzalez, S. Osindero, S. Ozair, S. Reed, J. Zhang, K. Zolna, J. Clune, N. de Freitas, ...

  2. [10]

    Carreira, E

    J. Carreira, E. Noland, A. Banki-Horvath, C. Hillier, and A. Zisserman. A short note about kinetics-600. ArXiv, abs/1808.01340, 2018

  3. [11]

    B. Chen, D. M. Monso, Y . Du, M. Simchowitz, R. Tedrake, and V . Sitzmann. Diffusion forcing: Next-token prediction meets full-sequence diffusion. ArXiv, abs/2407.01392, 2024

  4. [12]

    J. Chen, J. Yu, C. Ge, L. Yao, E. Xie, Y . Wu, Z. Wang, J. T. Kwok, P. Luo, H. Lu, and Z. Li. Pixart-α: Fast training of diffusion transformer for photorealistic text-to-image synthesis. ArXiv, abs/2310.00426, 2023

  5. [13]

    T.-S. Chen, A. Siarohin, W. Menapace, E. Deyneka, H. wei Chao, B. E. Jeon, Y . Fang, H.-Y . Lee, J. Ren, M.-H. Yang, and S. Tulyakov. Panda-70m: Captioning 70m videos with multiple cross- modality teachers. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (C...

  6. [14]

    X. Chen, Y . Wang, L. Zhang, S. Zhuang, X. Ma, J. Yu, Y . Wang, D. Lin, Y . Qiao, and Z. Liu. Seine: Short-to-long video diffusion model for generative transition and prediction. ArXiv, abs/2310.20700, 2023

  7. [15]

    Clark, J

    A. Clark, J. Donahue, and K. Simonyan. Adversarial video generation on complex datasets. arXiv: Computer Vision and Pattern Recognition, 2019. 11

  8. [16]

    C. Deng, D. Zhu, K. Li, S. Guang, and H. Fan. Causal diffusion transformers for generative modeling. ArXiv, abs/2412.12095, 2024

  9. [17]

    Dhariwal and A

    P. Dhariwal and A. Nichol. Diffusion models beat gans on image synthesis. ArXiv, abs/2105.05233, 2021

  10. [18]

    Dubey, A

    A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Yang, A. Fan, A. Goyal, A. S. Hartshorn, A. Yang, A. Mitra, A. Sravankumar, A. Korenev, A. Hinsvark, A. Rao, A. Zhang, A. Rodriguez, A. Gregerson, A. Spataru, B. tiste Rozière, B. Bir...

  11. [19]

    Esser, S

    P. Esser, S. Kulal, A. Blattmann, R. Entezari, J. Muller, H. Saini, Y . Levi, D. Lorenz, A. Sauer, F. Boesel, D. Podell, T. Dockhorn, Z. English, K. Lacey, A. Goodwin, Y . Marek, and R. Rombach. Scaling rectified flow transformers for high-resolution image synthesis. ArXiv, ab...

  12. [20]

    S. Ge, T. Hayes, H. Yang, X. Yin, G. Pang, D. Jacobs, J.-B. Huang, and D. Parikh. Long video generation with time-agnostic vqgan and time-sensitive transformer. In European Conference on Computer Vision, 2022

  13. [21]

    Goyal, P

    P. Goyal, P. Dollár, R. B. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y . Jia, and K. He. Accurate, large minibatch sgd: Training imagenet in 1 hour. ArXiv, abs/1706.02677, 2017

  14. [22]

    Y . Gu, W. Mao, and M. Z. Shou. Long-context autoregressive video modeling with next-frame prediction. ArXiv, abs/2503.19325, 2025

  15. [23]

    T. Han, W. Xie, and A. Zisserman. Memory-augmented dense predictive coding for video representation learning. In European Conference on Computer Vision, 2020

  16. [24]

    J. Ho. Classifier-free diffusion guidance. ArXiv, abs/2207.12598, 2022

  17. [25]

    J. Ho, A. Jain, and P. Abbeel. Denoising diffusion probabilistic models. In Conference and Workshop on Neural Information Processing Systems, 2020

  18. [26]

    J. Ho, T. Salimans, A. Gritsenko, W. Chan, M. Norouzi, and D. J. Fleet. Video diffusion models. ArXiv, abs/2204.03458, 2022

  19. [27]

    W. Hong, M. Ding, W. Zheng, X. Liu, and J. Tang. Cogvideo: Large-scale pretraining for text-to-video generation via transformers. ArXiv, abs/2205.15868, 2022

  20. [28]

    J. Hu, S. Hu, Y . Song, Y . Huang, M. Wang, H. Zhou, Z. Liu, W.-Y . Ma, and M. Sun. Acdit: Inter- polating autoregressive conditional modeling and diffusion transformer. ArXiv, abs/2412.07720, 2024

  21. [29]

    Hu and D

    Z. Hu and D. Xu. Videocontrolnet: A motion-guided video-to-video translation framework by using diffusion model with controlnet. ArXiv, abs/2307.14073, 2023

  22. [30]

    Huang, S

    Z. Huang, S. Zhuang, C. Fu, B. Yang, Y . Zhang, C. Sun, Z. Zhang, Y . Wang, C. Li, and Z.-J. Zha. Wegen: A unified model for interactive multimodal generation as we chat. ArXiv, abs/2503.01115, 2025

  23. [31]

    Kahembwe and S

    E. Kahembwe and S. Ramamoorthy. Lower dimensional kernels for video discriminators. Neural networks : the official journal of the International Neural Network Society, 132:506–520, 2019

  24. [32]

    Kaplan, S

    J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei. Scaling laws for neural language models. ArXiv, abs/2001.08361, 2020

  25. [33]

    T. Kim, G. Song, S. Lee, S. Kim, Y . Seo, S. Lee, S. H. Kim, H. Lee, and K. Bae. L-verse: Bidirectional generation between image and text. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 16505–16515, 2021. 13

  26. [34]

    Kondratyuk, L

    D. Kondratyuk, L. Yu, X. Gu, J. Lezama, J. Huang, R. Hornung, H. Adam, H. Akbari, Y . Alon, V . Birodkar, Y . Cheng, M.-C. Chiu, J. Dillon, I. Essa, A. Gupta, M. Hahn, A. Hauth, D. Hendon, A. Martinez, D. C. Minnen, D. A. Ross, G. Schindler, M. Sirotenko, K. Sohn, K. Somandepa...

  27. [35]

    W. Kong, Q. Tian, Z. Zhang, R. Min, Z. Dai, J. Zhou, J.-L. Xiong, X. Li, B. Wu, J. Zhang, K. Wu, Q. Lin, J. Yuan, Y . Long, A. Wang, A. Wang, C. Li, D. Huang, F. Yang, H. Tan, H. Wang, J. Song, J. Bai, J. Wu, J. Xue, J. Wang, K. Wang, M. Liu, P.-Y . Li, S. Li, W. Wang, W. Yu, ...

  28. [36]

    Kling ai, 2024

    Kuaishou. Kling ai, 2024. URL https://klingai.kuaishou.com/

  29. [37]

    B. F. Labs. Flux, 2024. URL https://blackforestlabs.ai/

  30. [38]

    C. Li, D. Huang, Z. Lu, Y . Xiao, Q. Pei, and L. Bai. A survey on long video generation: Challenges, methods, and prospects. ArXiv, abs/2403.16407, 2024

  31. [39]

    C. Liao, L. Liu, X. Wang, Z. Luo, X. Zhang, W. Zhao, J. Wu, L. Li, Z. Tian, and W. Huang. Mo- gao: An omni foundation model for interleaved multi-modal generation. ArXiv, abs/2505.05472, 2025

  32. [40]

    B. Lin, Y . Ge, X. Cheng, Z. Li, B. Zhu, S. Wang, X. He, Y . Ye, S. Yuan, L. Chen, T. Jia, J. Zhang, Z. Tang, Y . Pang, B. She, C. Yan, Z. Hu, X. wen Dong, L. Chen, Z. Pan, X. Zhou, S. Dong, Y . Tian, and L. Yuan. Open-sora plan: Open-source large video generation model. ArXiv...

  33. [41]

    Lipman, R

    Y . Lipman, R. T. Q. Chen, H. Ben-Hamu, M. Nickel, and M. Le. Flow matching for generative modeling. ArXiv, abs/2210.02747, 2022

  34. [42]

    H. Liu, W. Yan, M. Zaharia, and P. Abbeel. World model on million-length video and language with blockwise ringattention. ArXiv, abs/2402.08268, 2024

  35. [43]

    Loshchilov and F

    I. Loshchilov and F. Hutter. Decoupled weight decay regularization. InInternational Conference on Learning Representations, 2017

  36. [44]

    Z. Luo, D. Chen, Y . Zhang, Y . Huang, L. Wang, Y . Shen, D. Zhao, J. Zhou, and T.-P. Tan. Videofusion: Decomposed diffusion models for high-quality video generation. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10209–10218, 2023

  37. [45]

    X. Ma, Y . Wang, G. Jia, X. Chen, Z. Liu, Y .-F. Li, C. Chen, and Y . Qiao. Latte: Latent diffusion transformer for video generation. ArXiv, abs/2401.03048, 2024

  38. [46]

    Menapace, A

    W. Menapace, A. Siarohin, I. Skorokhodov, E. Deyneka, T.-S. Chen, A. Kag, Y . Fang, A. Stoliar, E. Ricci, J. Ren, and S. Tulyakov. Snap video: Scaled spatiotemporal transformers for text- to-video synthesis. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (...

  39. [47]

    Motamed, L

    S. Motamed, L. Culp, K. Swersky, P. Jaini, and R. Geirhos. Do generative video models understand physical principles? ArXiv, abs/2501.09038, 2025

  40. [48]

    K. Nan, R. Xie, P. Zhou, T. Fan, Z. Yang, Z. Chen, X. Li, J. Yang, and Y . Tai. Openvid-1m: A large-scale high-quality dataset for text-to-video generation. ArXiv, abs/2407.02371, 2024

  41. [49]

    Video generation models as world simulators, 2024

    OpenAI. Video generation models as world simulators, 2024. URL https://openai.com/ index/video-generation-models-as-world-simulators/

  42. [50]

    Ouyang, J

    Y . Ouyang, J. Yuan, H. Zhao, G. Wang, and B. Zhao. Flexifilm: Long video generation with flexible conditions. ArXiv, abs/2404.18620, 2024

  43. [51]

    Parmar, A

    N. Parmar, A. Vaswani, J. Uszkoreit, L. Kaiser, N. M. Shazeer, A. Ku, and D. Tran. Image transformer. In International Conference on Machine Learning, 2018. 14

  44. [52]

    Patrick, P.-Y

    M. Patrick, P.-Y . B. Huang, Y . M. Asano, F. Metze, A. Hauptmann, J. F. Henriques, and A. Vedaldi. Support-set bottlenecks for video-text representation learning. ArXiv, abs/2010.02824, 2020

  45. [53]

    W. S. Peebles and S. Xie. Scalable diffusion models with transformers. 2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 4172–4182, 2022

  46. [54]

    Pika 1.5, 2024

    PikaLabs. Pika 1.5, 2024. URL https://pika.art/

  47. [55]

    Podell, Z

    D. Podell, Z. English, K. Lacey, A. Blattmann, T. Dockhorn, J. Muller, J. Penna, and R. Rom- bach. Sdxl: Improving latent diffusion models for high-resolution image synthesis. ArXiv, abs/2307.01952, 2023

  48. [56]

    Polyak, A

    A. Polyak, A. Zohar, A. Brown, A. Tjandra, A. Sinha, A. Lee, A. Vyas, B. Shi, C.-Y . Ma, C.-Y . Chuang, D. Yan, D. Choudhary, D. Wang, G. Sethi, G. Pang, H. Ma, I. Misra, J. Hou, J. Wang, K. ran Jagadeesh, K. Li, L. Zhang, M. Singh, M. Williamson, M. Le, M. Yu, M. K. Singh, P....

  49. [57]

    Radford and K

    A. Radford and K. Narasimhan. Improving language understanding by generative pre-training. 2018

  50. [58]

    Radford, J

    A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever. Language models are unsupervised multitask learners. 2019

  51. [59]

    Radford, J

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, 2021

  52. [60]

    Ramesh, P

    A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen. Hierarchical text-conditional image generation with clip latents. ArXiv, abs/2204.06125, 2022

  53. [61]

    Z. Ren, Y . Wei, X. Guo, Y . Zhao, B. Kang, J. Feng, and X. Jin. Videoworld: Exploring knowledge learning from unlabeled videos. ArXiv, abs/2501.09781, 2025

  54. [62]

    Rombach, A

    R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer. High-resolution image synthesis with latent diffusion models. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10674–10685, 2021

  55. [63]

    Ronneberger, P

    O. Ronneberger, P. Fischer, and T. Brox. U-net: Convolutional networks for biomedical image segmentation. ArXiv, abs/1505.04597, 2015

  56. [64]

    Gen-3, 2024

    Runway. Gen-3, 2024. URL https://runwayml.com/

  57. [65]

    Saharia, W

    C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, S. K. S. Ghasemipour, B. K. Ayan, S. S. Mahdavi, R. G. Lopes, T. Salimans, J. Ho, D. J. Fleet, and M. Norouzi. Photorealistic text-to-image diffusion models with deep language understanding. ArXiv, abs/2205.11487, 2022

  58. [66]

    Saito, S

    M. Saito, S. Saito, M. Koyama, and S. Kobayashi. Train sparsely, generate densely: Memory- efficient unsupervised training of high-resolution temporal gan. International Journal of Computer Vision, 128:2586 – 2606, 2018

  59. [67]

    Segalis, D

    E. Segalis, D. Valevski, D. Lumen, Y . Matias, and Y . Leviathan. A picture is worth a thousand words: Principled recaptioning improves image generation. ArXiv, abs/2310.16656, 2023. 15

  60. [68]

    Singer, A

    U. Singer, A. Polyak, T. Hayes, X. Yin, J. An, S. Zhang, Q. Hu, H. Yang, O. Ashual, O. Gafni, D. Parikh, S. Gupta, and Y . Taigman. Make-a-video: Text-to-video generation without text-video data. ArXiv, abs/2209.14792, 2022

  61. [69]

    Y . Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole. Score-based generative modeling through stochastic differential equations. In International Conference on Machine Learning, 2021

  62. [70]

    Soomro, A

    K. Soomro, A. Zamir, and M. Shah. Ucf101: A dataset of 101 human actions classes from videos in the wild. ArXiv, abs/1212.0402, 2012

  63. [71]

    P. Sun, Y . Jiang, S. Chen, S. Zhang, B. Peng, P. Luo, and Z. Yuan. Autoregressive model beats diffusion: Llama for scalable image generation. ArXiv, abs/2406.06525, 2024

  64. [72]

    C. Team. Chameleon: Mixed-modal early-fusion foundation models. ArXiv, abs/2405.09818, 2024

  65. [73]

    Unterthiner, S

    T. Unterthiner, S. van Steenkiste, K. Kurach, R. Marinier, M. Michalski, and S. Gelly. Towards accurate generative models of video: A new metric & challenges. ArXiv, abs/1812.01717, 2018

  66. [74]

    A. Wang, B. Ai, B. Wen, C. Mao, C.-W. Xie, D. Chen, F. Yu, H. Zhao, J. Yang, J. Zeng, J. Wang, J. Zhang, J. Zhou, J. Wang, J. Chen, K. Zhu, K. Zhao, K. Yan, L. Huang, X. Meng, N. Zhang, P. Li, P. Wu, R. Chu, R. Feng, S. Zhang, S. Sun, T. Fang, T. Wang, T. Gui, T. Weng, T. Shen...

  67. [75]

    H. Wang, S. Suri, Y . Ren, H. Chen, and A. Shrivastava. Larp: Tokenizing videos with a learned autoregressive generative prior. ArXiv, abs/2410.21264, 2024

  68. [76]

    J. Wang, Y . Jiang, Z. Yuan, B. Peng, Z. Wu, and Y .-G. Jiang. Omnitokenizer: A joint image- video tokenizer for visual generation. ArXiv, abs/2406.09399, 2024

  69. [77]

    L. Wang, B. Huang, Z. Zhao, Z. Tong, Y . He, Y . Wang, Y . Wang, and Y . Qiao. Videomae v2: Scaling video masked autoencoders with dual masking. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14549–14560, 2023

  70. [78]

    Q. Wang, Y . Shi, J. Ou, R. Chen, K. Lin, J. Wang, B. Jiang, H. Yang, M. Zheng, X. Tao, F. Yang, P. Wan, and D. Zhang. Koala-36m: A large-scale video dataset improving consistency between fine-grained conditions and video content. ArXiv, abs/2410.08260, 2024

  71. [79]

    X. Wang, X. Zhang, Z. Luo, Q. Sun, Y . Cui, J. Wang, F. Zhang, Y . Wang, Z. Li, Q. Yu, Y . Zhao, Y . Ao, X. Min, T. Li, B. Wu, B. Zhao, B. Zhang, L. zi Wang, G. Liu, Z. He, X. Yang, J. Liu, Y . Lin, T. Huang, and Z. Wang. Emu3: Next-token prediction is all you need. ArXiv, abs...

  72. [80]

    Y . Wang, X. Chen, X. Ma, S. Zhou, Z. Huang, Y . Wang, C. Yang, Y . He, J. Yu, P. der Yang, Y . Guo, T. Wu, C. Si, Y . Jiang, C. Chen, C. C. Loy, B. Dai, D. Lin, Y . Qiao, and Z. Liu. Lavie: High-quality video generation with cascaded latent diffusion models. ArXiv, abs/2309.1...

  73. [81]

    Y . Wang, Y . He, Y . Li, K. Li, J. Yu, X. J. Ma, X. Chen, Y . Wang, P. Luo, Z. Liu, Y . Wang, L. Wang, and Y . Qiao. Internvid: A large-scale video-text dataset for multimodal understanding and generation. ArXiv, abs/2307.06942, 2023

  74. [82]

    Q. Wu, H. Ye, Y . Gu, H. Zhang, L. Wang, and D. He. Denoising masked autoencoders are certifiable robust vision learners. ArXiv, abs/2210.06983, 2022

  75. [83]

    S. Xiao, Y . Wang, J. Zhou, H. Yuan, X. Xing, R. Yan, S. Wang, T. Huang, and Z. Liu. Omnigen: Unified image generation. ArXiv, abs/2409.11340, 2024

  76. [84]

    H. Xu, G. Ghosh, P.-Y . B. Huang, D. Okhonko, A. Aghajanyan, and F. M. L. Z. C. Feichtenhofer. Videoclip: Contrastive pre-training for zero-shot video-text understanding. In Conference on Empirical Methods in Natural Language Processing, 2021. 16

  77. [85]

    J. Xu, T. Mei, T. Yao, and Y . Rui. Msr-vtt: A large video description dataset for bridging video and language. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 5288–5296, 2016

  78. [86]

    L. Xue, N. Constant, A. Roberts, M. Kale, R. Al-Rfou, A. Siddhant, A. Barua, and C. Raffel. mt5: A massively multilingual pre-trained text-to-text transformer. In North American Chapter of the Association for Computational Linguistics, 2020

  79. [87]

    A. Yang, B. Yang, B. Hui, B. Zheng, B. Yu, C. Zhou, C. Li, C. Li, D. Liu, F. Huang, G. Dong, H. Wei, H. Lin, J. Tang, J. Wang, J. Yang, J. Tu, J. Zhang, J. Ma, J. Xu, J. Zhou, J. Bai, J. He, J. Lin, K. Dang, K. Lu, K.-Y . Chen, K. Yang, M. Li, M. Xue, N. Ni, P. Zhang, P. Wang,...

  80. [88]

    S. Yang, J. Walker, J. Parker-Holder, Y . Du, J. Bruce, A. Barreto, P. Abbeel, and D. Schuurmans. Video as the new language for real-world decision making. ArXiv, abs/2402.17139, 2024

  81. [89]

    Z. Yang, J. Teng, W. Zheng, M. Ding, S. Huang, J. Xu, Y . Yang, W. Hong, X. Zhang, G. Feng, D. Yin, X. Gu, Y . Zhang, W. Wang, Y . Cheng, T. Liu, B. Xu, Y . Dong, and J. Tang. Cogvideox: Text-to-video diffusion models with an expert transformer. ArXiv, abs/2408.06072, 2024

  82. [90]

    H. Yi, S. Shao, T. Ye, J. Zhao, Q. Yin, M. Lingelbach, L. Yuan, Y . Tian, E. Xie, and D. Zhou. Magic 1-for-1: Generating one minute video clips within one minute. ArXiv, abs/2502.07701, 2025

  83. [91]

    T. Yin, Q. Zhang, R. Zhang, W. T. Freeman, F. Durand, E. Shechtman, and X. Huang. From slow bidirectional to fast causal video generators. ArXiv, abs/2412.07772, 2024

  84. [92]

    L. Yu, J. Lezama, N. B. Gundavarapu, L. Versari, K. Sohn, D. Minnen, Y . Cheng, A. Gupta, X. Gu, A. G. Hauptmann, B. Gong, M.-H. Yang, I. Essa, D. A. Ross, and L. Jiang. Language model beats diffusion - tokenizer is key to visual generation. In International Conference on Lear...

  85. [93]

    Z. Yue, S. Zhuang, K. Li, Y . Ding, and Y . Wang. V-stylist: Video stylization via collaboration and reflection of mllm agents. ArXiv, abs/2503.12077, 2025

  86. [94]

    C. Zhou, L. Yu, A. Babu, K. Tirumala, M. Yasunaga, L. Shamis, J. Kahn, X. Ma, L. Zettlemoyer, and O. Levy. Transfusion: Predict the next token and diffuse images with one multi-modal model. ArXiv, abs/2408.11039, 2024

  87. [95]

    D. Zhou, Q. Sun, Y . Peng, K. Yan, R. Dong, D. Wang, Z. Ge, N. Duan, X. Zhang, L. M. Ni, and H. yeung Shum. Taming teacher forcing for masked autoregressive video generation. ArXiv, abs/2501.12389, 2025

  88. [96]

    Zhuang, K

    S. Zhuang, K. Li, X. Chen, Y . Wang, Z. Liu, Y . Qiao, and Y . Wang. Vlogger: Make your dream a vlog. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 8806–8817, 2024

  89. [97]

    noisy_img3<diff>𝛼

    S. Zhuang, Z. Huang, B. Yang, Y . Zhang, F. Wang, C. Fu, C. Sun, Z.-J. Zha, C. Li, and Y . Wang. Get in video: Add anything you want to the video. ArXiv, abs/2503.06268, 2025. 17 A More Implementation Details Table 9: Pretraining Stage 1 setting. config Panda-70M optimizer Ada...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.