Pith. sign in

REVIEW 2 major objections 4 minor 84 references

FlashDecoder, a pure-Transformer video decoder, matches convolutional decoders' reconstruction quality while decoding 3.6–4.7× faster with up to 11× less GPU memory.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 00:45 UTC pith:NLIVV4QA

load-bearing objection A genuinely useful streaming decoder with honest numbers, but the abstract overclaims quality: per-frame fidelity matches, temporal realism doesn't, and the paper's own limitations section says so. the 2 major comments →

arxiv 2607.14898 v1 pith:NLIVV4QA submitted 2026-07-16 cs.CV cs.AI

FlashDecoder: Real-Time Latent-to-Pixel Streaming Decoder with Transformers

classification cs.CV cs.AI
keywords video decodinglatent diffusiontransformer decoderstreamingKV cachetemporal upsamplingreconstruction qualityreal-time inference
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

FlashDecoder is a pure-Transformer video decoder that turns latent video frames into pixels one frame at a time, attending only to the previous frame through a rolling cache. The paper's central claim is that this architecture removes the two barriers that kept Transformer decoders out of production video: explicit causal masks, which made training at 1080p infeasible, and unbounded temporal context, which made streaming slow. On the Wan2.1 and Wan2.2 latent spaces it matches each convolutional decoder's reconstruction quality (for example, 41.55 vs. 41.49 dB PSNR at 1080p) while decoding 3.6–4.7× faster and using up to 11× less GPU memory; with inference optimizations the gap widens to 12×. A sympathetic reader would care because VAE decoding is now the main bottleneck in real-time video generation, and a decoder with constant per-frame latency and flat memory makes streaming at high resolution practical.

Core claim

The discovery is that causality can be enforced by processing order instead of attention masks, and that this single change unlocks high-resolution training for a streaming Transformer decoder. Instead of loading all frames and masking out the future, FlashDecoder feeds one latent frame at a time through the same code path at train and test time; a rolling key-value cache holds at most two latent frames, so each step attends to the current frame and the immediately preceding one. This keeps memory and per-frame latency constant regardless of video length, allows training at up to 1080p on an 80 GB GPU, and — with the larger model and adversarial fine-tuning — closes the reconstruction-qualit

What carries the argument

The rolling key-value cache with a fixed temporal window (Wfrm = 2) is the central mechanism: each decoded frame's attention is limited to the current latent frame plus one past frame, so cost per frame and cache size are bounded by a constant rather than video length. Combined with training that uses the identical streaming protocol, it removes the need for causal attention masks entirely, which is what lets the model train at high resolution. The temporal-first upsampling — channel expansion to multiply frame count, two refinement Transformer layers, then MLP+PixelShuffle for spatial upsampling — keeps the compute scaling manageable.

Load-bearing premise

The load-bearing premise is that reconstruction quality measured on encodings of real UltraVideo clips transfers to latents actually produced by the Wan2.2 diffusion model — the paper never runs an end-to-end generation experiment (Table 2 is reconstruction-only), and its limitations section (D) does not list this gap; a secondary premise is that a window of two latent frames carries enough temporal context for all motion content, a claim tested only on reconstruction clips,

What would settle it

Run FlashDecoder on latents sampled from the actual diffusion model over long generations (say 400+ frames) and measure PSNR/rFVD against the same model's training distribution; if quality collapses or the fixed window causes temporal drift on generated motion, the transfer claim fails. Alternatively, training a same-budget convolutional decoder at 1080p and finding it matches FlashDecoder on both quality and speed would undercut the speed-at-equal-quality claim.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Video generation pipelines can keep decoding time and memory flat as generated videos get longer; decode cost no longer grows with the number of frames.
  • High-resolution (1080p) Transformer decoding becomes trainable on a single 8-GPU node, so decoder quality no longer trails convolutional decoders.
  • The same decoder can be dropped into different latent spaces (8× and 16× spatial compression were tested) with only a PixelUnshuffle preprocessing step, making it a portable replacement decoder.
  • Streaming causality is obtained for free at inference, so no padding, blending, or chunking is needed for long-form output.
  • With FP8 quantization and CUDA-graph execution, the optimized variant reaches about 151 FPS at 720p and 43 FPS at 1080p on one H100, pointing toward interactive generation rates.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If decoder-only quality transfers to generated latents, the natural next step — pairing a streaming Transformer encoder with this decoder and training the full autoencoder from scratch — would likely produce latent spaces shaped for Transformer generation rather than convolutional spatial locality. The paper names this as future work; the unstated payoff is removing the current resolution mismatch
  • The window-size ablation (quality nearly flat for 2, 3, or 4 frames) suggests that in these high-compression latents, almost all temporal context needed for reconstruction lives in the immediately preceding frame. If that holds broadly, even cheaper 1-frame-window streaming decoders may be viable, and it hints that temporal redundancy in latent video is low after compression.
  • Because the optimizations (torch.compile, CUDA graphs, precomputed RoPE, FP8) stack multiplicatively, similar tricks should transfer to other streaming transformer decoders; the paper only demonstrates them for the 16× compression variant.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper proposes FlashDecoder, a pure-Transformer video decoder that maps latent frames to pixels in a streaming, frame-by-frame manner. Causality is enforced by processing order rather than attention masks, and a fixed-size rolling KV cache (W_frm=2 throughout most experiments) keeps per-frame cost and memory bounded. The decoder uses temporal-first upsampling with Transformer refinement and MLP+PixelShuffle spatial upsampling, trained with L1, LPIPS, and adversarial losses. Experiments on UltraVideo evaluate reconstruction on real-video latents encoded by Wan2.1/Wan2.2, reporting PSNR/LPIPS close to the convolutional baselines, 3.6–4.7x throughput gains (up to 12x with FP8/CUDA-graph optimizations), and up to 11x lower peak memory. The paper claims that FlashDecoder matches convolutional decoder reconstruction quality while enabling real-time streaming.

Significance. The architectural insight is valuable: aligning training and inference through the same streaming protocol removes the need for causal-mask materialization, which is a concrete obstacle to high-resolution training of causal Transformer decoders. The complexity analysis and ablations are clear, and the speed/memory numbers, if reproduced, support the practical relevance of the design. The main caveat is that the paper's central 'matches quality' claim is too broad relative to its own temporal metric, and the evaluation does not include end-to-end decoding of latents produced by a diffusion model. With appropriate qualification and additional validation, this would be a useful contribution to streaming video-decoder design.

major comments (2)
  1. [Abstract; §4.3, Table 2; §D] The unqualified claim that FlashDecoder matches convolutional decoders in reconstruction quality is contradicted by the paper's own temporal metric. Section 4.1 introduces rCD-FVD as 'a more faithful measure of temporal coherence.' In Table 2 (4x16x16 group), FlashDecoder-XL's rFVD is 10.77 vs. 7.97 at 480p, 12.75 vs. 10.39 at 720p, and 12.08 vs. 8.16 at 1080p — roughly 30–50% worse than Wan2.2. Section D explicitly concedes this gap. The abstract and Section 4.3 currently report PSNR/LPIPS matching as if it were holistic quality. Please either narrow the quality claim to per-frame fidelity (PSNR/LPIPS) or provide evidence that the rFVD gap is not perceptually significant (e.g., a user study or additional temporal-realism analysis). The speculation that production decoders used more compute/data is not a substitute for such evidence.
  2. [§4.1, Table 2; §D] The evaluation is reconstruction-only on latents of real UltraVideo clips. There is no end-to-end experiment in which latents are produced by the Wan2.2 diffusion model and then decoded by FlashDecoder. Since the decoder is intended for deployment after a generative model, the distribution of generated latents may differ from encoder latents, and the W_frm=2 window's sufficiency for generated motion content is untested. This is load-bearing for the 'real-time video generation' framing. Please add at least one end-to-end comparison (e.g., decode Wan2.2-sampled latents and report rFVD/FVD, or a user study), or explicitly state this as a limitation and provide a distribution-shift analysis. Section D's limitation list does not currently mention this omission.
minor comments (4)
  1. [§4.1] The evaluation-data paragraph contains the odd string 'clips short 1.zipsplit'; this appears to be a formatting artifact and should be corrected.
  2. [§3.3, Eq. (2)] The head dimension D_h is used in the cache-shape equation and complexity analysis but is never defined; please define it as D/N (or the per-head dimension after GQA partitioning).
  3. [§4.6, Table 2] The text states that FP8 quantization increases rFVD by up to 0.94, but in Table 2's 4x16x16 group at 720p, FlashDecoder-XL-Opt rFVD (12.22) is lower than FlashDecoder-XL (12.75). Please report the direction and magnitude of the FP8 effect consistently, or state the exceptions.
  4. [Table 1] The baseline row shows '331.4→16.6' in the FPS column. This is explained in the text, but the table would be clearer if the arrow were defined in the caption or split into two rows.

Circularity Check

0 steps flagged

No circular derivation chain; the rFVD gap is an overclaim issue, not a circular step.

full rationale

FlashDecoder's central claims are evaluated against external baselines (Wan2.1/Wan2.2, HunyuanVideo, AToken, MAGI-1, OmniTokenizer) on the external UltraVideo benchmark. The training objective in Eq. (4) uses standard L1, LPIPS, and adversarial losses; no parameter is fitted to the headline metrics, and the reconstruction-quality numbers in Table 2 are measured comparisons rather than outputs derived from the paper's own assumptions. The speed and memory claims follow from the fixed rolling-KV window (Eq. (2), Algorithm 1) and are directly measured on an H100 GPU. The only author self-citation appears as [23] inside a survey list about few-step distillation, and it is not load-bearing for the decoder architecture or for any result. The paper itself discloses the rFVD shortfall in Section D ('FlashDecoder-XL falls short of Wan2.2 and HunyuanVideo in rFVD, despite comparable PSNR and LPIPS'); this contradicts the abstract's unqualified 'matches ... reconstruction quality' wording and is a correctness/overclaim risk, but it is not a circular reduction. Nothing in the derivation chain equates an output to an input by construction.

Axiom & Free-Parameter Ledger

4 free parameters · 6 axioms · 0 invented entities

FlashDecoder introduces no new physical or latent entity; it is a neural architecture. The central result depends on training choices (window size, loss weights, model scale), on the fixed Wan2.x encoders, and on an untested transfer assumption from real-video latents to generated latents.

free parameters (4)
  • Temporal window size W_frm = 2
    Fixed at 2 for all main experiments; ablation Table 4 shows W_frm in {2,3,4} are similar with 2 best. Chosen from evaluation data, not derived.
  • Loss weights (lambda_L1, lambda_LPIPS, lambda_adv, R1) = 1.0, 0.1/0.25, 1e-4, 0.1024
    Hand-set training hyperparameters in Supplementary Table A; affects quality-vs-sharpness tradeoff but not stated as optimized.
  • Model scale (depth D, heads N, KV groups G) = 20 blocks, D=1536, N=24, G=3, 769.3M params
    Selected via scaling study Table 5; an engineering choice within compute budget, not a quantity fitted to a theory.
  • FP8 static-calibrated quantization = not specified
    Calibration protocol not given; causes PSNR drop 0.06–0.71 dB and rFVD increase up to 0.94 depending on resolution.
axioms (6)
  • domain assumption The pretrained Wan2.1/Wan2.2 encoder produces latent frames that the decoder is trained to invert; encoder is treated as fixed and correct.
    Section 3.1; the whole claim is defined on these specific latent spaces.
  • domain assumption Relative-position RoPE encoding with positions assigned inside the current window generalizes to arbitrarily long videos.
    Section 3.3 and 4.3; supports the infinite-length streaming claim, tested only up to 400 frames in Figure 4.
  • ad hoc to paper A sliding temporal window of two latent frames is sufficient context for high-quality reconstruction.
    Section 3.3 and Table 4; W_frm=2 was chosen after an ablation, not guaranteed for fast motion or large scene changes.
  • domain assumption Reconstruction accuracy on real-video encodings (UltraVideo) is representative of decoding diffusion-generated latents.
    Section 4.1; no end-to-end text-to-video experiment is run, so distribution shift is unmeasured.
  • domain assumption PSNR, LPIPS, and rFVD on 25-frame clips capture reconstruction and temporal quality for streaming video.
    Section 4.1; these are standard metrics, but rFVD is computed without confidence intervals and clips are short relative to streaming use.
  • standard math FlashAttention/FlexAttention compute exact attention; memory claims rely on no hidden approximation from these kernels.
    Section 3.3; the OOM and memory comparisons treat these implementations as exact attention with predictable memory.

pith-pipeline@v1.3.0-alltime-deepseek · 18836 in / 13355 out tokens · 123806 ms · 2026-08-02T00:45:05.771545+00:00 · methodology

0 comments
read the original abstract

Real-time video generation demands fast decoding as much as fast denoising, yet current latent video diffusion models rely on 3D convolutional decoders that are slow and memory-intensive at high resolutions or for long video. We introduce FlashDecoder, a fast, memory-efficient pure-Transformer video decoder that decodes latents to pixels frame by frame. At each step, the current frame attends only to a fixed-size window of past frames through a rolling KV cache. The fixed temporal window keeps decoding fast and memory bounded regardless of video length, enabling constant-latency streaming. Because frames are processed sequentially, temporal causality is enforced without explicit attention masks, enabling training at resolutions up to 1080p and matching the reconstruction quality of convolutional decoders. On the Wan2.1 and Wan2.2 latent spaces, FlashDecoder matches each convolutional decoder in reconstruction quality (e.g., 41.55dB vs. 41.49dB PSNR at 1080p) while decoding 3.6x-4.7x faster with up to 11x less memory on a single H100 GPU. With architecture-aware inference optimizations, the speedup widens to 12x.

Figures

Figures reproduced from arXiv: 2607.14898 by Minguk Kang, Suha Kwak.

Figure 1
Figure 1. Figure 1: VAE decoding is a major bottleneck for real-time video generation. Measured with our MotionStream [49] imple￾mentation at 720p. The Wan2.2 [65] decoder consumes 64.6% of total inference time, limiting generation to 10.4 FPS. FlashDe￾coder reduces this share to 16.4%, more than doubling end-to-end throughput to 24.8 FPS. 71, 83], higher-compression VAEs [1, 5, 16, 65, 72], and few-step distillation [6, 23, … view at source ↗
Figure 2
Figure 2. Figure 2: Qualitative comparison of 720p reconstruction results. We compare reconstructed frames from video decoders with 4× temporal and 16× spatial compression: (a) Wan2.2-TAEHV [3], (b) AToken [37], (c) Wan2.2 [65], (d) our FlashDecoder-XL-Opt, and (e) ground truth. (a) fails to synthesize fine details such as wall textures, while (b) produces blurry reconstructions. (c) and (d) yield visually comparable outputs,… view at source ↗
Figure 3
Figure 3. Figure 3: FlashDecoder pipeline. FlashDecoder is a pure-Transformer decoder that converts video latents to pixels in a frame-by-frame manner. Each latent frame zt is linearly projected, processed by a Transformer backbone with a fixed-size rolling KV cache that stores the most recent Wfrm frames (temporal window size), temporally upsampled by factor rt via channel expansion and refinement layers, and spatially upsam… view at source ↗
Figure 4
Figure 4. Figure 4: Per-frame PSNR on long videos at 720p. Averaged over 40 videos (>400 frames each) from UltraVideo. FlashDe￾coder maintains stable quality with constant memory regardless of video length. 4.4. Window Size Ablation [PITH_FULL_IMAGE:figures/full_fig_p008_4.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

84 extracted references · 25 linked inside Pith

  1. [1]

    Cosmos world foun- dation model platform for physical ai.arXiv preprint arXiv:2501.03575, 2025

    Niket Agarwal, Arslan Ali, Maciej Bala, Yogesh Balaji, Erik Barker, Tiffany Cai, Prithvijit Chattopadhyay, Yongxin Chen, Yin Cui, Yifan Ding, et al. Cosmos world foun- dation model platform for physical ai.arXiv preprint arXiv:2501.03575, 2025. 1, 2

  2. [2]

    Gqa: Training generalized multi-query transformer models from multi-head checkpoints.arXiv preprint arXiv:2305.13245,

    Joshua Ainslie, James Lee-Thorp, Michiel De Jong, Yury Zemlyanskiy, Federico Lebr ´on, and Sumit Sanghai. Gqa: Training generalized multi-query transformer models from multi-head checkpoints.arXiv preprint arXiv:2305.13245,

  3. [3]

    Taehv: Tiny autoencoder for hun- yuan video.https://github.com/madebyollin/ taehv, 2025

    Ollin Boer Bohan. Taehv: Tiny autoencoder for hun- yuan video.https://github.com/madebyollin/ taehv, 2025. 2, 3, 7, 13, 16, 17, 18

  4. [4]

    A short note about kinetics- 600.arXiv preprint arXiv:1808.01340, 2018

    Joao Carreira, Eric Noland, Andras Banki-Horvath, Chloe Hillier, and Andrew Zisserman. A short note about kinetics- 600.arXiv preprint arXiv:1808.01340, 2018. 6, 13

  5. [5]

    Deep compres- sion autoencoder for efficient high-resolution diffusion mod- els

    Junyu Chen, Han Cai, Junsong Chen, Enze Xie, Shang Yang, Haotian Tang, Muyang Li, and Song Han. Deep compres- sion autoencoder for efficient high-resolution diffusion mod- els. InInternational Conference on Learning Representa- tions (ICLR), 2025. 1, 2

  6. [6]

    Sana-sprint: One-step diffusion with continuous-time con- sistency distillation.arXiv preprint arXiv:2503.09641, 2025

    Junsong Chen, Shuchen Xue, Yuyang Zhao, Jincheng Yu, Sayak Paul, Junyu Chen, Han Cai, Song Han, and Enze Xie. Sana-sprint: One-step diffusion with continuous-time con- sistency distillation.arXiv preprint arXiv:2503.09641, 2025. 1

  7. [7]

    Fu, Stefano Ermon, Atri Rudra, and Christopher R´e

    Tri Dao, Daniel Y . Fu, Stefano Ermon, Atri Rudra, and Christopher R´e. FlashAttention: Fast and memory-efficient exact attention with IO-awareness. InConference on Neural Information Processing Systems (NeurIPS), 2022. 2, 5

  8. [8]

    Veo: a text-to-video generation system.https: //storage.googleapis.com/deepmind-media/ veo/Veo-3-Tech-Report.pdf, 2024

    DeepMind. Veo: a text-to-video generation system.https: //storage.googleapis.com/deepmind-media/ veo/Veo-3-Tech-Report.pdf, 2024. 1

  9. [9]

    Tam- ing transformers for high-resolution image synthesis

    Patrick Esser, Robin Rombach, and Bjorn Ommer. Tam- ing transformers for high-resolution image synthesis. In IEEE Conference on Computer Vision and Pattern Recog- nition (CVPR), 2021. 5, 14

  10. [10]

    Taming transformers for high-resolution image synthe- sis.https : / / github

    Patrick Esser, Robin Rombach, and Bjorn Ommer. Taming transformers for high-resolution image synthe- sis.https : / / github . com / CompVis / taming - transformers, 2021. 13

  11. [11]

    Scaling recti- fied flow transformers for high-resolution image synthesis

    Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas M ¨uller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al. Scaling recti- fied flow transformers for high-resolution image synthesis. InInternational Conference on Machine Learning (ICML),

  12. [12]

    Dat- acomp: In search of the next generation of multimodal datasets

    Samir Yitzhak Gadre, Gabriel Ilharco, Alex Fang, Jonathan Hayase, Georgios Smyrnis, Thao Nguyen, Ryan Marten, Mitchell Wortsman, Dhruba Ghosh, Jieyu Zhang, et al. Dat- acomp: In search of the next generation of multimodal datasets. InConference on Neural Information Processing Systems (NeurIPS), 2023. 6, 13

  13. [13]

    Seedance 1.0: Exploring the boundaries of video generation models.arXiv preprint arXiv:2506.09113,

    Yu Gao, Haoyuan Guo, Tuyen Hoang, Weilin Huang, Lu Jiang, Fangyuan Kong, Huixia Li, Jiashi Li, Liang Li, Xi- aojie Li, et al. Seedance 1.0: Exploring the boundaries of video generation models.arXiv preprint arXiv:2506.09113,

  14. [14]

    On the content bias in fr ´echet video distance

    Songwei Ge, Aniruddha Mahapatra, Gaurav Parmar, Jun- Yan Zhu, and Jia-Bin Huang. On the content bias in fr ´echet video distance. InIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2024. 6, 7, 15

  15. [15]

    Generative Adversarial Nets

    Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative Adversarial Nets. InConference on Neural Information Processing Systems (NeurIPS), 2014. 5, 6, 13

  16. [16]

    Ltx-video: Realtime video latent diffusion.arXiv preprint arXiv:2501.00103,

    Yoav HaCohen, Nisan Chiprut, Benny Brazowski, Daniel Shalem, Dudu Moshe, Eitan Richardson, Eran Levin, Guy Shiran, Nir Zabari, Ori Gordon, et al. Ltx-video: Realtime video latent diffusion.arXiv preprint arXiv:2501.00103,

  17. [17]

    Learnings from scaling visual tokenizers for reconstruction and generation

    Philippe Hansen-Estruch, David Yan, Ching-Yao Chuang, Orr Zohar, Jialiang Wang, Tingbo Hou, Tao Xu, Sriram Vishwanath, Peter Vajda, and Xinlei Chen. Learnings from scaling visual tokenizers for reconstruction and generation. InInternational Conference on Machine Learning (ICML),

  18. [18]

    Unified latents (ul): How to train your latents

    Jonathan Heek, Emiel Hoogeboom, Thomas Mensink, and Tim Salimans. Unified latents (ul): How to train your latents. arXiv preprint arXiv:2602.17270, 2026. 15

  19. [19]

    Reducing the dimensionality of data with neural networks.science,

    Geoffrey E Hinton and Ruslan R Salakhutdinov. Reducing the dimensionality of data with neural networks.science,

  20. [20]

    Denoising dif- fusion probabilistic models

    Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising dif- fusion probabilistic models. InConference on Neural Infor- mation Processing Systems (NeurIPS), 2020. 1

  21. [21]

    Self forcing: Bridging the train- test gap in autoregressive video diffusion.arXiv preprint arXiv:2506.08009, 2025

    Xun Huang, Zhengqi Li, Guande He, Mingyuan Zhou, and Eli Shechtman. Self forcing: Bridging the train- test gap in autoregressive video diffusion.arXiv preprint arXiv:2506.08009, 2025. 1

  22. [22]

    Image-to-image translation with conditional adver- sarial networks

    Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros. Image-to-image translation with conditional adver- sarial networks. InIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017. 5, 13, 14

  23. [23]

    Distilling diffusion models into condi- tional gans

    Minguk Kang, Richard Zhang, Connelly Barnes, Sylvain Paris, Suha Kwak, Jaesik Park, Eli Shechtman, Jun-Yan Zhu, and Taesung Park. Distilling diffusion models into condi- tional gans. InEuropean Conference on Computer Vision (ECCV), 2024. 1

  24. [24]

    R. Keys. Cubic convolution interpolation for digital image processing.IEEE Transactions on Acoustics, Speech, and Signal Processing, 1981. 6, 13

  25. [25]

    Consistency trajectory mod- els: Learning probability flow ode trajectory of diffusion

    Dongjun Kim, Chieh-Hsin Lai, Wei-Hsiang Liao, Naoki Mu- rata, Yuhta Takida, Toshimitsu Uesaka, Yutong He, Yuki Mitsufuji, and Stefano Ermon. Consistency trajectory mod- els: Learning probability flow ode trajectory of diffusion. arXiv preprint arXiv:2310.02279, 2023. 1

  26. [26]

    Auto-encoding varia- tional bayes.arXiv preprint arXiv:1312.6114, 2013

    Diederik P Kingma and Max Welling. Auto-encoding varia- tional bayes.arXiv preprint arXiv:1312.6114, 2013. 1

  27. [27]

    Hunyuanvideo: A systematic framework for large video generative models.arXiv preprint arXiv:2412.03603, 2024

    Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, et al. Hunyuanvideo: A systematic framework for large video generative models.arXiv preprint arXiv:2412.03603, 2024. 1, 2, 7, 15

  28. [28]

    Black Forest Labs, Stephen Batifol, Andreas Blattmann, Frederic Boesel, Saksham Consul, Cyril Diagne, Tim Dock- horn, Jack English, Zion English, Patrick Esser, et al. Flux. 1 kontext: Flow matching for in-context image generation and editing in latent space.arXiv preprint arXiv:2506.15742,

  29. [29]

    Flex- attention for efficient high-resolution vision-language mod- els

    Junyan Li, Delin Chen, Tianle Cai, Peihao Chen, Yining Hong, Zhenfang Chen, Yikang Shen, and Chuang Gan. Flex- attention for efficient high-resolution vision-language mod- els. InEuropean Conference on Computer Vision (ECCV),

  30. [30]

    Wf-vae: Enhancing video vae by wavelet-driven energy flow for latent video diffusion model

    Zongjian Li, Bin Lin, Yang Ye, Liuhan Chen, Xinhua Cheng, Shenghai Yuan, and Li Yuan. Wf-vae: Enhancing video vae by wavelet-driven energy flow for latent video diffusion model. InIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2025. 2

  31. [31]

    Open-sora plan: Open-source large video generation model.arXiv preprint arXiv:2412.00131, 2024

    Bin Lin, Yunyang Ge, Xinhua Cheng, Zongjian Li, Bin Zhu, Shaodong Wang, Xianyi He, Yang Ye, Shenghai Yuan, Li- uhan Chen, et al. Open-sora plan: Open-source large video generation model.arXiv preprint arXiv:2412.00131, 2024. 1, 2

  32. [32]

    Diffusion adversarial post-training for one-step video generation

    Shanchuan Lin, Xin Xia, Yuxi Ren, Ceyuan Yang, Xuefeng Xiao, and Lu Jiang. Diffusion adversarial post-training for one-step video generation. InInternational Conference on Machine Learning (ICML), 2025. 1

  33. [33]

    Autoregressive adversarial post-training for real-time inter- active video generation.arXiv preprint arXiv:2506.09350,

    Shanchuan Lin, Ceyuan Yang, Hao He, Jianwen Jiang, Yuxi Ren, Xin Xia, Yang Zhao, Xuefeng Xiao, and Lu Jiang. Autoregressive adversarial post-training for real-time inter- active video generation.arXiv preprint arXiv:2506.09350,

  34. [34]

    Rolling forcing: Autoregressive long video diffusion in real time.arXiv preprint arXiv:2509.25161, 2025

    Kunhao Liu, Wenbo Hu, Jiale Xu, Ying Shan, and Shijian Lu. Rolling forcing: Autoregressive long video diffusion in real time.arXiv preprint arXiv:2509.25161, 2025. 1

  35. [35]

    Im- proving reconstruction of representation autoencoder.arXiv preprint arXiv:2602.08620, 2026

    Siyu Liu, Chujie Qin, Hubery Yin, Qixin Yan, Zheng-Peng Duan, Chen Li, Jing Lyu, Chun-Le Guo, and Chongyi Li. Im- proving reconstruction of representation autoencoder.arXiv preprint arXiv:2602.08620, 2026. 2

  36. [36]

    Decoupled Weight De- cay Regularization

    Ilya Loshchilov and Frank Hutter. Decoupled Weight De- cay Regularization. InInternational Conference on Learning Representations (ICLR), 2019. 14

  37. [37]

    Atoken: A unified tokenizer for vision.arXiv preprint arXiv:2509.14476, 2025

    Jiasen Lu, Liangchen Song, Mingze Xu, Byeongjoo Ahn, Yanjun Wang, Chen Chen, Afshin Dehghan, and Yinfei Yang. Atoken: A unified tokenizer for vision.arXiv preprint arXiv:2509.14476, 2025. 2, 3, 7, 13, 16, 17, 18

  38. [38]

    On distillation of guided diffusion models

    Chenlin Meng, Robin Rombach, Ruiqi Gao, Diederik Kingma, Stefano Ermon, Jonathan Ho, and Tim Salimans. On distillation of guided diffusion models. InIEEE Confer- ence on Computer Vision and Pattern Recognition (CVPR),

  39. [39]

    Which Training Methods for GANs do actually Converge? InInternational Conference on Machine Learning (ICML),

    Lars Mescheder, Sebastian Nowozin, and Andreas Geiger. Which Training Methods for GANs do actually Converge? InInternational Conference on Machine Learning (ICML),

  40. [40]

    Improved denoising dif- fusion probabilistic models

    Alex Nichol and Prafulla Dhariwal. Improved denoising dif- fusion probabilistic models. InInternational Conference on Machine Learning (ICML), 2021. 1

  41. [41]

    Video generation models as world simula- tors.https://openai.com/research/video- generation - models - as - world - simulators,

    OpenAI. Video generation models as world simula- tors.https://openai.com/research/video- generation - models - as - world - simulators,

  42. [42]

    Dinov2: Learning robust visual features without supervision

    Maxime Oquab, Timoth ´ee Darcet, Th ´eo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023. 2, 15

  43. [43]

    Scalable diffusion mod- els with transformers

    William Peebles and Saining Xie. Scalable diffusion mod- els with transformers. InIEEE International Conference on Computer Vision (ICCV), 2023. 1

  44. [44]

    Hierarchical text-conditional image gen- eration with clip latents.arXiv preprint arXiv:2204.06125,

    Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image gen- eration with clip latents.arXiv preprint arXiv:2204.06125,

  45. [45]

    High-resolution image syn- thesis with latent diffusion models

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj¨orn Ommer. High-resolution image syn- thesis with latent diffusion models. InIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 1, 2

  46. [46]

    Progressive Distillation for Fast Sampling of Diffusion Models

    Tim Salimans and Jonathan Ho. Progressive Distillation for Fast Sampling of Diffusion Models. InInternational Con- ference on Learning Representations (ICLR), 2022. 1

  47. [47]

    Fast high- resolution image synthesis with latent adversarial diffusion distillation

    Axel Sauer, Frederic Boesel, Tim Dockhorn, Andreas Blattmann, Patrick Esser, and Robin Rombach. Fast high- resolution image synthesis with latent adversarial diffusion distillation. InSIGGRAPH Asia 2024 Conference Papers,

  48. [48]

    Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network

    Wenzhe Shi, Jose Caballero, Ferenc Husz ´ar, Johannes Totz, Andrew P Aitken, Rob Bishop, Daniel Rueckert, and Zehan Wang. Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network. In IEEE Conference on Computer Vision and Pattern Recogni- tion (CVPR), 2016. 5

  49. [49]

    Motion- stream: Real-time video generation with interactive motion controls.arXiv preprint arXiv:2511.01266, 2025

    Joonghyuk Shin, Zhengqi Li, Richard Zhang, Jun-Yan Zhu, Jaesik Park, Eli Schechtman, and Xun Huang. Motion- stream: Real-time video generation with interactive motion controls.arXiv preprint arXiv:2511.01266, 2025. 1

  50. [50]

    Score-based generative modeling through stochastic differential equa- tions.arXiv preprint arXiv:2011.13456, 2020

    Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Ab- hishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equa- tions.arXiv preprint arXiv:2011.13456, 2020. 1

  51. [51]

    Consistency Models

    Yang Song, Prafulla Dhariwal, Mark Chen, and Ilya Sutskever. Consistency Models. InInternational Conference on Machine Learning (ICML), 2023. 1

  52. [52]

    Ltx-2: The complete ai creative engine for video production.https://ltx.studio/blog/ltx- 2- the- complete- ai- creative- engine- for- video-production, 2025

    LTX Studio. Ltx-2: The complete ai creative engine for video production.https://ltx.studio/blog/ltx- 2- the- complete- ai- creative- engine- for- video-production, 2025. 1, 2

  53. [53]

    Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 2024

    Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 2024. 4

  54. [54]

    Vidtok: A versatile and open-source video tokenizer.arXiv preprint arXiv:2412.13061, 2024

    Anni Tang, Tianyu He, Junliang Guo, Xinle Cheng, Li Song, and Jiang Bian. Vidtok: A versatile and open-source video tokenizer.arXiv preprint arXiv:2412.13061, 2024. 2

  55. [55]

    Mochi 1.https :/ /github .com/ genmoai/models, 2024

    Genmo Team. Mochi 1.https :/ /github .com/ genmoai/models, 2024. 2

  56. [56]

    Gemma: Open models based on gemini research and tech- nology.arXiv preprint arXiv:2403.08295, 2024

    Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivi `ere, Mihir Sanjay Kale, Juliette Love, et al. Gemma: Open models based on gemini research and tech- nology.arXiv preprint arXiv:2403.08295, 2024. 4

  57. [57]

    Krea realtime 14b: Real-time, long-form ai video generation

    Krea Team. Krea realtime 14b: Real-time, long-form ai video generation. Blog post, Krea AI, 2025. 1

  58. [58]

    Magi-1: Autoregressive video genera- tion at scale.arXiv preprint arXiv:2505.13211, 2025

    Hansi Teng, Hongyu Jia, Lei Sun, Lingzhi Li, Maolin Li, Mingqiu Tang, Shuai Han, Tianning Zhang, WQ Zhang, Weifeng Luo, et al. Magi-1: Autoregressive video genera- tion at scale.arXiv preprint arXiv:2505.13211, 2025. 2, 7

  59. [59]

    Scaling text-to-image dif- fusion transformers with representation autoencoders.arXiv preprint arXiv:2601.16208, 2026

    Shengbang Tong, Boyang Zheng, Ziteng Wang, Bingda Tang, Nanye Ma, Ellis Brown, Jihan Yang, Rob Fergus, Yann LeCun, and Saining Xie. Scaling text-to-image dif- fusion transformers with representation autoencoders.arXiv preprint arXiv:2601.16208, 2026. 2

  60. [60]

    Siglip 2: Multilingual vision-language en- coders with improved semantic understanding, localization, and dense features.arXiv preprint arXiv:2502.14786, 2025

    Michael Tschannen, Alexey Gritsenko, Xiao Wang, Muham- mad Ferjad Naeem, Ibrahim Alabdulmohsin, Nikhil Parthasarathy, Talfan Evans, Lucas Beyer, Ye Xia, Basil Mustafa, et al. Siglip 2: Multilingual vision-language en- coders with improved semantic understanding, localization, and dense features.arXiv preprint arXiv:2502.14786, 2025. 2, 15

  61. [61]

    Image processing in python.CSI Communica- tions, 23(2), 2012

    P Umesh. Image processing in python.CSI Communica- tions, 23(2), 2012. 6, 13

  62. [62]

    Fvd: A new metric for video generation

    Thomas Unterthiner, Sjoerd van Steenkiste, Karol Kurach, Rapha¨el Marinier, Marcin Michalski, and Sylvain Gelly. Fvd: A new metric for video generation. InDGS@ICLR,

  63. [63]

    Attention is all you need

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. InConference on Neural Information Processing Systems (NeurIPS), 2017. 2, 4

  64. [64]

    Phenaki: Variable length video generation from open domain textual descriptions

    Ruben Villegas, Mohammad Babaeizadeh, Pieter-Jan Kin- dermans, Hernan Moraldo, Han Zhang, Mohammad Taghi Saffar, Santiago Castro, Julius Kunze, and Dumitru Erhan. Phenaki: Variable length video generation from open domain textual descriptions. InInternational Conference on Learn- ing Representations (ICLR), 2023. 2

  65. [65]

    Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314, 2025

    Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianx- iao Yang, Jianyuan Zeng, Jiayu Wang, Jingfeng Zhang, Jin- gren Zhou, Jinkai Wang, Jixuan Chen, Kai Zhu, Kang Zhao, Keyu Yan, Lianghua Huang, Mengyang Feng, Ningyi Zhang, Pandeng Li, Pingyu Wu, Ruihang Chu, Ruili Feng, Shiwei Zhang, Siyang Sun, Tao Fan...

  66. [66]

    Omnitokenizer: A joint image- video tokenizer for visual generation

    Junke Wang, Yi Jiang, Zehuan Yuan, Bingyue Peng, Zux- uan Wu, and Yu-Gang Jiang. Omnitokenizer: A joint image- video tokenizer for visual generation. InConference on Neu- ral Information Processing Systems (NeurIPS), 2024. 2, 7

  67. [67]

    Sana: Efficient high-resolution image syn- thesis with linear diffusion transformers.arXiv preprint arXiv:2410.10629, 2024

    Enze Xie, Junsong Chen, Junyu Chen, Han Cai, Haotian Tang, Yujun Lin, Zhekai Zhang, Muyang Li, Ligeng Zhu, Yao Lu, et al. Sana: Efficient high-resolution image syn- thesis with linear diffusion transformers.arXiv preprint arXiv:2410.10629, 2024. 1

  68. [68]

    Videovae+: Large motion video autoencoding with cross-modal video vae

    Yazhou Xing, Yang Fei, Yingqing He, Jingye Chen, Jiaxin Xie, Xiaowei Chi, and Qifeng Chen. Videovae+: Large motion video autoencoding with cross-modal video vae. In IEEE International Conference on Computer Vision (ICCV),

  69. [69]

    Ultravideo: High-quality uhd video dataset with comprehensive captions.arXiv preprint arXiv:2506.13691, 2025

    Zhucun Xue, Jiangning Zhang, Teng Hu, Haoyang He, Yinan Chen, Yuxuan Cai, Yabiao Wang, Chengjie Wang, Yong Liu, Xiangtai Li, and Dacheng Tao. Ultravideo: High-quality uhd video dataset with comprehensive captions.arXiv preprint arXiv:2506.13691, 2025. 6, 7

  70. [70]

    Cogvideox: Text-to-video diffusion models with an expert transformer

    Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiao- han Zhang, Guanyu Feng, Da Yin, Yuxuan.Zhang, Weihan Wang, Yean Cheng, Bin Xu, Xiaotao Gu, Yuxiao Dong, and Jie Tang. Cogvideox: Text-to-video diffusion models with an expert transformer. InInternational Conference on Learning Representations (ICLR...

  71. [71]

    Fasterdit: Towards faster diffusion transformers train- ing without architecture modification

    Jingfeng Yao, Cheng Wang, Wenyu Liu, and Xinggang Wang. Fasterdit: Towards faster diffusion transformers train- ing without architecture modification. InConference on Neu- ral Information Processing Systems (NeurIPS), 2024. 1

  72. [72]

    Reconstruc- tion vs

    Jingfeng Yao, Bin Yang, and Xinggang Wang. Reconstruc- tion vs. generation: Taming optimization dilemma in latent diffusion models. InIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2025. 1

  73. [73]

    Improved distribution matching distillation for fast image synthesis

    Tianwei Yin, Micha ¨el Gharbi, Taesung Park, Richard Zhang, Eli Shechtman, Fredo Durand, and Bill Freeman. Improved distribution matching distillation for fast image synthesis. InConference on Neural Information Processing Systems (NeurIPS), 2024. 1

  74. [74]

    From slow bidirectional to fast autoregressive video diffusion mod- els

    Tianwei Yin, Qiang Zhang, Richard Zhang, William T Free- man, Fredo Durand, Eli Shechtman, and Xun Huang. From slow bidirectional to fast autoregressive video diffusion mod- els. InIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2025. 1

  75. [75]

    Vector-quantized image modeling with im- proved VQGAN

    Jiahui Yu, Xin Li, Jing Yu Koh, Han Zhang, Ruoming Pang, James Qin, Alexander Ku, Yuanzhong Xu, Jason Baldridge, and Yonghui Wu. Vector-quantized image modeling with im- proved VQGAN. InInternational Conference on Learning Representations (ICLR), 2022. 2

  76. [76]

    Language model beats diffusion - tokenizer is key to visual generation

    Lijun Yu, Jose Lezama, Nitesh Bharadwaj Gundavarapu, Luca Versari, Kihyuk Sohn, David Minnen, Yong Cheng, Agrim Gupta, Xiuye Gu, Alexander G Hauptmann, Boqing Gong, Ming-Hsuan Yang, Irfan Essa, David A Ross, and Lu Jiang. Language model beats diffusion - tokenizer is key to visual generation. InInternational Conference on Learning Representations (ICLR), 2024. 2

  77. [77]

    Root mean square layer nor- malization

    Biao Zhang and Rico Sennrich. Root mean square layer nor- malization. InConference on Neural Information Processing Systems (NeurIPS), 2019. 4

  78. [78]

    The unreasonable effectiveness of deep features as a perceptual metric

    Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shecht- man, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. InIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018. 5, 6, 7, 13, 14

  79. [79]

    Waver: Wave your way to lifelike video genera- tion.arXiv preprint arXiv:2508.15761, 2025

    Yifu Zhang, Hao Yang, Yuqi Zhang, Yifei Hu, Fengda Zhu, Chuang Lin, Xiaofeng Mei, Yi Jiang, Bingyue Peng, and Ze- huan Yuan. Waver: Wave your way to lifelike video genera- tion.arXiv preprint arXiv:2508.15761, 2025. 1, 2

  80. [80]

    Cv- vae: A compatible video vae for latent generative video mod- els

    Sijie Zhao, Yong Zhang, Xiaodong Cun, Shaoshu Yang, Muyao Niu, Xiaoyu Li, Wenbo Hu, and Ying Shan. Cv- vae: A compatible video vae for latent generative video mod- els. InConference on Neural Information Processing Sys- tems (NeurIPS), 2024. 2

Showing first 80 references.