REVIEW 2 major objections 4 minor 84 references
FlashDecoder, a pure-Transformer video decoder, matches convolutional decoders' reconstruction quality while decoding 3.6–4.7× faster with up to 11× less GPU memory.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 00:45 UTC pith:NLIVV4QA
load-bearing objection A genuinely useful streaming decoder with honest numbers, but the abstract overclaims quality: per-frame fidelity matches, temporal realism doesn't, and the paper's own limitations section says so. the 2 major comments →
FlashDecoder: Real-Time Latent-to-Pixel Streaming Decoder with Transformers
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The discovery is that causality can be enforced by processing order instead of attention masks, and that this single change unlocks high-resolution training for a streaming Transformer decoder. Instead of loading all frames and masking out the future, FlashDecoder feeds one latent frame at a time through the same code path at train and test time; a rolling key-value cache holds at most two latent frames, so each step attends to the current frame and the immediately preceding one. This keeps memory and per-frame latency constant regardless of video length, allows training at up to 1080p on an 80 GB GPU, and — with the larger model and adversarial fine-tuning — closes the reconstruction-qualit
What carries the argument
The rolling key-value cache with a fixed temporal window (Wfrm = 2) is the central mechanism: each decoded frame's attention is limited to the current latent frame plus one past frame, so cost per frame and cache size are bounded by a constant rather than video length. Combined with training that uses the identical streaming protocol, it removes the need for causal attention masks entirely, which is what lets the model train at high resolution. The temporal-first upsampling — channel expansion to multiply frame count, two refinement Transformer layers, then MLP+PixelShuffle for spatial upsampling — keeps the compute scaling manageable.
Load-bearing premise
The load-bearing premise is that reconstruction quality measured on encodings of real UltraVideo clips transfers to latents actually produced by the Wan2.2 diffusion model — the paper never runs an end-to-end generation experiment (Table 2 is reconstruction-only), and its limitations section (D) does not list this gap; a secondary premise is that a window of two latent frames carries enough temporal context for all motion content, a claim tested only on reconstruction clips,
What would settle it
Run FlashDecoder on latents sampled from the actual diffusion model over long generations (say 400+ frames) and measure PSNR/rFVD against the same model's training distribution; if quality collapses or the fixed window causes temporal drift on generated motion, the transfer claim fails. Alternatively, training a same-budget convolutional decoder at 1080p and finding it matches FlashDecoder on both quality and speed would undercut the speed-at-equal-quality claim.
If this is right
- Video generation pipelines can keep decoding time and memory flat as generated videos get longer; decode cost no longer grows with the number of frames.
- High-resolution (1080p) Transformer decoding becomes trainable on a single 8-GPU node, so decoder quality no longer trails convolutional decoders.
- The same decoder can be dropped into different latent spaces (8× and 16× spatial compression were tested) with only a PixelUnshuffle preprocessing step, making it a portable replacement decoder.
- Streaming causality is obtained for free at inference, so no padding, blending, or chunking is needed for long-form output.
- With FP8 quantization and CUDA-graph execution, the optimized variant reaches about 151 FPS at 720p and 43 FPS at 1080p on one H100, pointing toward interactive generation rates.
Where Pith is reading between the lines
- If decoder-only quality transfers to generated latents, the natural next step — pairing a streaming Transformer encoder with this decoder and training the full autoencoder from scratch — would likely produce latent spaces shaped for Transformer generation rather than convolutional spatial locality. The paper names this as future work; the unstated payoff is removing the current resolution mismatch
- The window-size ablation (quality nearly flat for 2, 3, or 4 frames) suggests that in these high-compression latents, almost all temporal context needed for reconstruction lives in the immediately preceding frame. If that holds broadly, even cheaper 1-frame-window streaming decoders may be viable, and it hints that temporal redundancy in latent video is low after compression.
- Because the optimizations (torch.compile, CUDA graphs, precomputed RoPE, FP8) stack multiplicatively, similar tricks should transfer to other streaming transformer decoders; the paper only demonstrates them for the 16× compression variant.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes FlashDecoder, a pure-Transformer video decoder that maps latent frames to pixels in a streaming, frame-by-frame manner. Causality is enforced by processing order rather than attention masks, and a fixed-size rolling KV cache (W_frm=2 throughout most experiments) keeps per-frame cost and memory bounded. The decoder uses temporal-first upsampling with Transformer refinement and MLP+PixelShuffle spatial upsampling, trained with L1, LPIPS, and adversarial losses. Experiments on UltraVideo evaluate reconstruction on real-video latents encoded by Wan2.1/Wan2.2, reporting PSNR/LPIPS close to the convolutional baselines, 3.6–4.7x throughput gains (up to 12x with FP8/CUDA-graph optimizations), and up to 11x lower peak memory. The paper claims that FlashDecoder matches convolutional decoder reconstruction quality while enabling real-time streaming.
Significance. The architectural insight is valuable: aligning training and inference through the same streaming protocol removes the need for causal-mask materialization, which is a concrete obstacle to high-resolution training of causal Transformer decoders. The complexity analysis and ablations are clear, and the speed/memory numbers, if reproduced, support the practical relevance of the design. The main caveat is that the paper's central 'matches quality' claim is too broad relative to its own temporal metric, and the evaluation does not include end-to-end decoding of latents produced by a diffusion model. With appropriate qualification and additional validation, this would be a useful contribution to streaming video-decoder design.
major comments (2)
- [Abstract; §4.3, Table 2; §D] The unqualified claim that FlashDecoder matches convolutional decoders in reconstruction quality is contradicted by the paper's own temporal metric. Section 4.1 introduces rCD-FVD as 'a more faithful measure of temporal coherence.' In Table 2 (4x16x16 group), FlashDecoder-XL's rFVD is 10.77 vs. 7.97 at 480p, 12.75 vs. 10.39 at 720p, and 12.08 vs. 8.16 at 1080p — roughly 30–50% worse than Wan2.2. Section D explicitly concedes this gap. The abstract and Section 4.3 currently report PSNR/LPIPS matching as if it were holistic quality. Please either narrow the quality claim to per-frame fidelity (PSNR/LPIPS) or provide evidence that the rFVD gap is not perceptually significant (e.g., a user study or additional temporal-realism analysis). The speculation that production decoders used more compute/data is not a substitute for such evidence.
- [§4.1, Table 2; §D] The evaluation is reconstruction-only on latents of real UltraVideo clips. There is no end-to-end experiment in which latents are produced by the Wan2.2 diffusion model and then decoded by FlashDecoder. Since the decoder is intended for deployment after a generative model, the distribution of generated latents may differ from encoder latents, and the W_frm=2 window's sufficiency for generated motion content is untested. This is load-bearing for the 'real-time video generation' framing. Please add at least one end-to-end comparison (e.g., decode Wan2.2-sampled latents and report rFVD/FVD, or a user study), or explicitly state this as a limitation and provide a distribution-shift analysis. Section D's limitation list does not currently mention this omission.
minor comments (4)
- [§4.1] The evaluation-data paragraph contains the odd string 'clips short 1.zipsplit'; this appears to be a formatting artifact and should be corrected.
- [§3.3, Eq. (2)] The head dimension D_h is used in the cache-shape equation and complexity analysis but is never defined; please define it as D/N (or the per-head dimension after GQA partitioning).
- [§4.6, Table 2] The text states that FP8 quantization increases rFVD by up to 0.94, but in Table 2's 4x16x16 group at 720p, FlashDecoder-XL-Opt rFVD (12.22) is lower than FlashDecoder-XL (12.75). Please report the direction and magnitude of the FP8 effect consistently, or state the exceptions.
- [Table 1] The baseline row shows '331.4→16.6' in the FPS column. This is explained in the text, but the table would be clearer if the arrow were defined in the caption or split into two rows.
Circularity Check
No circular derivation chain; the rFVD gap is an overclaim issue, not a circular step.
full rationale
FlashDecoder's central claims are evaluated against external baselines (Wan2.1/Wan2.2, HunyuanVideo, AToken, MAGI-1, OmniTokenizer) on the external UltraVideo benchmark. The training objective in Eq. (4) uses standard L1, LPIPS, and adversarial losses; no parameter is fitted to the headline metrics, and the reconstruction-quality numbers in Table 2 are measured comparisons rather than outputs derived from the paper's own assumptions. The speed and memory claims follow from the fixed rolling-KV window (Eq. (2), Algorithm 1) and are directly measured on an H100 GPU. The only author self-citation appears as [23] inside a survey list about few-step distillation, and it is not load-bearing for the decoder architecture or for any result. The paper itself discloses the rFVD shortfall in Section D ('FlashDecoder-XL falls short of Wan2.2 and HunyuanVideo in rFVD, despite comparable PSNR and LPIPS'); this contradicts the abstract's unqualified 'matches ... reconstruction quality' wording and is a correctness/overclaim risk, but it is not a circular reduction. Nothing in the derivation chain equates an output to an input by construction.
Axiom & Free-Parameter Ledger
free parameters (4)
- Temporal window size W_frm =
2
- Loss weights (lambda_L1, lambda_LPIPS, lambda_adv, R1) =
1.0, 0.1/0.25, 1e-4, 0.1024
- Model scale (depth D, heads N, KV groups G) =
20 blocks, D=1536, N=24, G=3, 769.3M params
- FP8 static-calibrated quantization =
not specified
axioms (6)
- domain assumption The pretrained Wan2.1/Wan2.2 encoder produces latent frames that the decoder is trained to invert; encoder is treated as fixed and correct.
- domain assumption Relative-position RoPE encoding with positions assigned inside the current window generalizes to arbitrarily long videos.
- ad hoc to paper A sliding temporal window of two latent frames is sufficient context for high-quality reconstruction.
- domain assumption Reconstruction accuracy on real-video encodings (UltraVideo) is representative of decoding diffusion-generated latents.
- domain assumption PSNR, LPIPS, and rFVD on 25-frame clips capture reconstruction and temporal quality for streaming video.
- standard math FlashAttention/FlexAttention compute exact attention; memory claims rely on no hidden approximation from these kernels.
read the original abstract
Real-time video generation demands fast decoding as much as fast denoising, yet current latent video diffusion models rely on 3D convolutional decoders that are slow and memory-intensive at high resolutions or for long video. We introduce FlashDecoder, a fast, memory-efficient pure-Transformer video decoder that decodes latents to pixels frame by frame. At each step, the current frame attends only to a fixed-size window of past frames through a rolling KV cache. The fixed temporal window keeps decoding fast and memory bounded regardless of video length, enabling constant-latency streaming. Because frames are processed sequentially, temporal causality is enforced without explicit attention masks, enabling training at resolutions up to 1080p and matching the reconstruction quality of convolutional decoders. On the Wan2.1 and Wan2.2 latent spaces, FlashDecoder matches each convolutional decoder in reconstruction quality (e.g., 41.55dB vs. 41.49dB PSNR at 1080p) while decoding 3.6x-4.7x faster with up to 11x less memory on a single H100 GPU. With architecture-aware inference optimizations, the speedup widens to 12x.
Figures
Reference graph
Works this paper leans on
-
[1]
Cosmos world foun- dation model platform for physical ai.arXiv preprint arXiv:2501.03575, 2025
Niket Agarwal, Arslan Ali, Maciej Bala, Yogesh Balaji, Erik Barker, Tiffany Cai, Prithvijit Chattopadhyay, Yongxin Chen, Yin Cui, Yifan Ding, et al. Cosmos world foun- dation model platform for physical ai.arXiv preprint arXiv:2501.03575, 2025. 1, 2
Pith/arXiv arXiv 2025
-
[2]
Joshua Ainslie, James Lee-Thorp, Michiel De Jong, Yury Zemlyanskiy, Federico Lebr ´on, and Sumit Sanghai. Gqa: Training generalized multi-query transformer models from multi-head checkpoints.arXiv preprint arXiv:2305.13245,
-
[3]
Taehv: Tiny autoencoder for hun- yuan video.https://github.com/madebyollin/ taehv, 2025
Ollin Boer Bohan. Taehv: Tiny autoencoder for hun- yuan video.https://github.com/madebyollin/ taehv, 2025. 2, 3, 7, 13, 16, 17, 18
2025
-
[4]
A short note about kinetics- 600.arXiv preprint arXiv:1808.01340, 2018
Joao Carreira, Eric Noland, Andras Banki-Horvath, Chloe Hillier, and Andrew Zisserman. A short note about kinetics- 600.arXiv preprint arXiv:1808.01340, 2018. 6, 13
Pith/arXiv arXiv 2018
-
[5]
Deep compres- sion autoencoder for efficient high-resolution diffusion mod- els
Junyu Chen, Han Cai, Junsong Chen, Enze Xie, Shang Yang, Haotian Tang, Muyang Li, and Song Han. Deep compres- sion autoencoder for efficient high-resolution diffusion mod- els. InInternational Conference on Learning Representa- tions (ICLR), 2025. 1, 2
2025
-
[6]
Junsong Chen, Shuchen Xue, Yuyang Zhao, Jincheng Yu, Sayak Paul, Junyu Chen, Han Cai, Song Han, and Enze Xie. Sana-sprint: One-step diffusion with continuous-time con- sistency distillation.arXiv preprint arXiv:2503.09641, 2025. 1
arXiv 2025
-
[7]
Fu, Stefano Ermon, Atri Rudra, and Christopher R´e
Tri Dao, Daniel Y . Fu, Stefano Ermon, Atri Rudra, and Christopher R´e. FlashAttention: Fast and memory-efficient exact attention with IO-awareness. InConference on Neural Information Processing Systems (NeurIPS), 2022. 2, 5
2022
-
[8]
Veo: a text-to-video generation system.https: //storage.googleapis.com/deepmind-media/ veo/Veo-3-Tech-Report.pdf, 2024
DeepMind. Veo: a text-to-video generation system.https: //storage.googleapis.com/deepmind-media/ veo/Veo-3-Tech-Report.pdf, 2024. 1
2024
-
[9]
Tam- ing transformers for high-resolution image synthesis
Patrick Esser, Robin Rombach, and Bjorn Ommer. Tam- ing transformers for high-resolution image synthesis. In IEEE Conference on Computer Vision and Pattern Recog- nition (CVPR), 2021. 5, 14
2021
-
[10]
Taming transformers for high-resolution image synthe- sis.https : / / github
Patrick Esser, Robin Rombach, and Bjorn Ommer. Taming transformers for high-resolution image synthe- sis.https : / / github . com / CompVis / taming - transformers, 2021. 13
2021
-
[11]
Scaling recti- fied flow transformers for high-resolution image synthesis
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas M ¨uller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al. Scaling recti- fied flow transformers for high-resolution image synthesis. InInternational Conference on Machine Learning (ICML),
-
[12]
Dat- acomp: In search of the next generation of multimodal datasets
Samir Yitzhak Gadre, Gabriel Ilharco, Alex Fang, Jonathan Hayase, Georgios Smyrnis, Thao Nguyen, Ryan Marten, Mitchell Wortsman, Dhruba Ghosh, Jieyu Zhang, et al. Dat- acomp: In search of the next generation of multimodal datasets. InConference on Neural Information Processing Systems (NeurIPS), 2023. 6, 13
2023
-
[13]
Seedance 1.0: Exploring the boundaries of video generation models.arXiv preprint arXiv:2506.09113,
Yu Gao, Haoyuan Guo, Tuyen Hoang, Weilin Huang, Lu Jiang, Fangyuan Kong, Huixia Li, Jiashi Li, Liang Li, Xi- aojie Li, et al. Seedance 1.0: Exploring the boundaries of video generation models.arXiv preprint arXiv:2506.09113,
-
[14]
On the content bias in fr ´echet video distance
Songwei Ge, Aniruddha Mahapatra, Gaurav Parmar, Jun- Yan Zhu, and Jia-Bin Huang. On the content bias in fr ´echet video distance. InIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2024. 6, 7, 15
2024
-
[15]
Generative Adversarial Nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative Adversarial Nets. InConference on Neural Information Processing Systems (NeurIPS), 2014. 5, 6, 13
2014
-
[16]
Ltx-video: Realtime video latent diffusion.arXiv preprint arXiv:2501.00103,
Yoav HaCohen, Nisan Chiprut, Benny Brazowski, Daniel Shalem, Dudu Moshe, Eitan Richardson, Eran Levin, Guy Shiran, Nir Zabari, Ori Gordon, et al. Ltx-video: Realtime video latent diffusion.arXiv preprint arXiv:2501.00103,
-
[17]
Learnings from scaling visual tokenizers for reconstruction and generation
Philippe Hansen-Estruch, David Yan, Ching-Yao Chuang, Orr Zohar, Jialiang Wang, Tingbo Hou, Tao Xu, Sriram Vishwanath, Peter Vajda, and Xinlei Chen. Learnings from scaling visual tokenizers for reconstruction and generation. InInternational Conference on Machine Learning (ICML),
-
[18]
Unified latents (ul): How to train your latents
Jonathan Heek, Emiel Hoogeboom, Thomas Mensink, and Tim Salimans. Unified latents (ul): How to train your latents. arXiv preprint arXiv:2602.17270, 2026. 15
arXiv 2026
-
[19]
Reducing the dimensionality of data with neural networks.science,
Geoffrey E Hinton and Ruslan R Salakhutdinov. Reducing the dimensionality of data with neural networks.science,
-
[20]
Denoising dif- fusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising dif- fusion probabilistic models. InConference on Neural Infor- mation Processing Systems (NeurIPS), 2020. 1
2020
-
[21]
Xun Huang, Zhengqi Li, Guande He, Mingyuan Zhou, and Eli Shechtman. Self forcing: Bridging the train- test gap in autoregressive video diffusion.arXiv preprint arXiv:2506.08009, 2025. 1
Pith/arXiv arXiv 2025
-
[22]
Image-to-image translation with conditional adver- sarial networks
Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros. Image-to-image translation with conditional adver- sarial networks. InIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017. 5, 13, 14
2017
-
[23]
Distilling diffusion models into condi- tional gans
Minguk Kang, Richard Zhang, Connelly Barnes, Sylvain Paris, Suha Kwak, Jaesik Park, Eli Shechtman, Jun-Yan Zhu, and Taesung Park. Distilling diffusion models into condi- tional gans. InEuropean Conference on Computer Vision (ECCV), 2024. 1
2024
-
[24]
R. Keys. Cubic convolution interpolation for digital image processing.IEEE Transactions on Acoustics, Speech, and Signal Processing, 1981. 6, 13
1981
-
[25]
Consistency trajectory mod- els: Learning probability flow ode trajectory of diffusion
Dongjun Kim, Chieh-Hsin Lai, Wei-Hsiang Liao, Naoki Mu- rata, Yuhta Takida, Toshimitsu Uesaka, Yutong He, Yuki Mitsufuji, and Stefano Ermon. Consistency trajectory mod- els: Learning probability flow ode trajectory of diffusion. arXiv preprint arXiv:2310.02279, 2023. 1
Pith/arXiv arXiv 2023
-
[26]
Auto-encoding varia- tional bayes.arXiv preprint arXiv:1312.6114, 2013
Diederik P Kingma and Max Welling. Auto-encoding varia- tional bayes.arXiv preprint arXiv:1312.6114, 2013. 1
Pith/arXiv arXiv 2013
-
[27]
Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, et al. Hunyuanvideo: A systematic framework for large video generative models.arXiv preprint arXiv:2412.03603, 2024. 1, 2, 7, 15
Pith/arXiv arXiv 2024
-
[28]
Black Forest Labs, Stephen Batifol, Andreas Blattmann, Frederic Boesel, Saksham Consul, Cyril Diagne, Tim Dock- horn, Jack English, Zion English, Patrick Esser, et al. Flux. 1 kontext: Flow matching for in-context image generation and editing in latent space.arXiv preprint arXiv:2506.15742,
-
[29]
Flex- attention for efficient high-resolution vision-language mod- els
Junyan Li, Delin Chen, Tianle Cai, Peihao Chen, Yining Hong, Zhenfang Chen, Yikang Shen, and Chuang Gan. Flex- attention for efficient high-resolution vision-language mod- els. InEuropean Conference on Computer Vision (ECCV),
-
[30]
Wf-vae: Enhancing video vae by wavelet-driven energy flow for latent video diffusion model
Zongjian Li, Bin Lin, Yang Ye, Liuhan Chen, Xinhua Cheng, Shenghai Yuan, and Li Yuan. Wf-vae: Enhancing video vae by wavelet-driven energy flow for latent video diffusion model. InIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2025. 2
2025
-
[31]
Open-sora plan: Open-source large video generation model.arXiv preprint arXiv:2412.00131, 2024
Bin Lin, Yunyang Ge, Xinhua Cheng, Zongjian Li, Bin Zhu, Shaodong Wang, Xianyi He, Yang Ye, Shenghai Yuan, Li- uhan Chen, et al. Open-sora plan: Open-source large video generation model.arXiv preprint arXiv:2412.00131, 2024. 1, 2
Pith/arXiv arXiv 2024
-
[32]
Diffusion adversarial post-training for one-step video generation
Shanchuan Lin, Xin Xia, Yuxi Ren, Ceyuan Yang, Xuefeng Xiao, and Lu Jiang. Diffusion adversarial post-training for one-step video generation. InInternational Conference on Machine Learning (ICML), 2025. 1
2025
-
[33]
Shanchuan Lin, Ceyuan Yang, Hao He, Jianwen Jiang, Yuxi Ren, Xin Xia, Yang Zhao, Xuefeng Xiao, and Lu Jiang. Autoregressive adversarial post-training for real-time inter- active video generation.arXiv preprint arXiv:2506.09350,
-
[34]
Kunhao Liu, Wenbo Hu, Jiale Xu, Ying Shan, and Shijian Lu. Rolling forcing: Autoregressive long video diffusion in real time.arXiv preprint arXiv:2509.25161, 2025. 1
Pith/arXiv arXiv 2025
-
[35]
Im- proving reconstruction of representation autoencoder.arXiv preprint arXiv:2602.08620, 2026
Siyu Liu, Chujie Qin, Hubery Yin, Qixin Yan, Zheng-Peng Duan, Chen Li, Jing Lyu, Chun-Le Guo, and Chongyi Li. Im- proving reconstruction of representation autoencoder.arXiv preprint arXiv:2602.08620, 2026. 2
arXiv 2026
-
[36]
Decoupled Weight De- cay Regularization
Ilya Loshchilov and Frank Hutter. Decoupled Weight De- cay Regularization. InInternational Conference on Learning Representations (ICLR), 2019. 14
2019
-
[37]
Atoken: A unified tokenizer for vision.arXiv preprint arXiv:2509.14476, 2025
Jiasen Lu, Liangchen Song, Mingze Xu, Byeongjoo Ahn, Yanjun Wang, Chen Chen, Afshin Dehghan, and Yinfei Yang. Atoken: A unified tokenizer for vision.arXiv preprint arXiv:2509.14476, 2025. 2, 3, 7, 13, 16, 17, 18
arXiv 2025
-
[38]
On distillation of guided diffusion models
Chenlin Meng, Robin Rombach, Ruiqi Gao, Diederik Kingma, Stefano Ermon, Jonathan Ho, and Tim Salimans. On distillation of guided diffusion models. InIEEE Confer- ence on Computer Vision and Pattern Recognition (CVPR),
-
[39]
Which Training Methods for GANs do actually Converge? InInternational Conference on Machine Learning (ICML),
Lars Mescheder, Sebastian Nowozin, and Andreas Geiger. Which Training Methods for GANs do actually Converge? InInternational Conference on Machine Learning (ICML),
-
[40]
Improved denoising dif- fusion probabilistic models
Alex Nichol and Prafulla Dhariwal. Improved denoising dif- fusion probabilistic models. InInternational Conference on Machine Learning (ICML), 2021. 1
2021
-
[41]
Video generation models as world simula- tors.https://openai.com/research/video- generation - models - as - world - simulators,
OpenAI. Video generation models as world simula- tors.https://openai.com/research/video- generation - models - as - world - simulators,
-
[42]
Dinov2: Learning robust visual features without supervision
Maxime Oquab, Timoth ´ee Darcet, Th ´eo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023. 2, 15
Pith/arXiv arXiv 2023
-
[43]
Scalable diffusion mod- els with transformers
William Peebles and Saining Xie. Scalable diffusion mod- els with transformers. InIEEE International Conference on Computer Vision (ICCV), 2023. 1
2023
-
[44]
Hierarchical text-conditional image gen- eration with clip latents.arXiv preprint arXiv:2204.06125,
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image gen- eration with clip latents.arXiv preprint arXiv:2204.06125,
-
[45]
High-resolution image syn- thesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj¨orn Ommer. High-resolution image syn- thesis with latent diffusion models. InIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 1, 2
2022
-
[46]
Progressive Distillation for Fast Sampling of Diffusion Models
Tim Salimans and Jonathan Ho. Progressive Distillation for Fast Sampling of Diffusion Models. InInternational Con- ference on Learning Representations (ICLR), 2022. 1
2022
-
[47]
Fast high- resolution image synthesis with latent adversarial diffusion distillation
Axel Sauer, Frederic Boesel, Tim Dockhorn, Andreas Blattmann, Patrick Esser, and Robin Rombach. Fast high- resolution image synthesis with latent adversarial diffusion distillation. InSIGGRAPH Asia 2024 Conference Papers,
2024
-
[48]
Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network
Wenzhe Shi, Jose Caballero, Ferenc Husz ´ar, Johannes Totz, Andrew P Aitken, Rob Bishop, Daniel Rueckert, and Zehan Wang. Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network. In IEEE Conference on Computer Vision and Pattern Recogni- tion (CVPR), 2016. 5
2016
-
[49]
Joonghyuk Shin, Zhengqi Li, Richard Zhang, Jun-Yan Zhu, Jaesik Park, Eli Schechtman, and Xun Huang. Motion- stream: Real-time video generation with interactive motion controls.arXiv preprint arXiv:2511.01266, 2025. 1
arXiv 2025
-
[50]
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Ab- hishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equa- tions.arXiv preprint arXiv:2011.13456, 2020. 1
Pith/arXiv arXiv 2011
-
[51]
Consistency Models
Yang Song, Prafulla Dhariwal, Mark Chen, and Ilya Sutskever. Consistency Models. InInternational Conference on Machine Learning (ICML), 2023. 1
2023
-
[52]
Ltx-2: The complete ai creative engine for video production.https://ltx.studio/blog/ltx- 2- the- complete- ai- creative- engine- for- video-production, 2025
LTX Studio. Ltx-2: The complete ai creative engine for video production.https://ltx.studio/blog/ltx- 2- the- complete- ai- creative- engine- for- video-production, 2025. 1, 2
2025
-
[53]
Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 2024
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 2024. 4
2024
-
[54]
Vidtok: A versatile and open-source video tokenizer.arXiv preprint arXiv:2412.13061, 2024
Anni Tang, Tianyu He, Junliang Guo, Xinle Cheng, Li Song, and Jiang Bian. Vidtok: A versatile and open-source video tokenizer.arXiv preprint arXiv:2412.13061, 2024. 2
Pith/arXiv arXiv 2024
-
[55]
Mochi 1.https :/ /github .com/ genmoai/models, 2024
Genmo Team. Mochi 1.https :/ /github .com/ genmoai/models, 2024. 2
2024
-
[56]
Gemma: Open models based on gemini research and tech- nology.arXiv preprint arXiv:2403.08295, 2024
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivi `ere, Mihir Sanjay Kale, Juliette Love, et al. Gemma: Open models based on gemini research and tech- nology.arXiv preprint arXiv:2403.08295, 2024. 4
Pith/arXiv arXiv 2024
-
[57]
Krea realtime 14b: Real-time, long-form ai video generation
Krea Team. Krea realtime 14b: Real-time, long-form ai video generation. Blog post, Krea AI, 2025. 1
2025
-
[58]
Magi-1: Autoregressive video genera- tion at scale.arXiv preprint arXiv:2505.13211, 2025
Hansi Teng, Hongyu Jia, Lei Sun, Lingzhi Li, Maolin Li, Mingqiu Tang, Shuai Han, Tianning Zhang, WQ Zhang, Weifeng Luo, et al. Magi-1: Autoregressive video genera- tion at scale.arXiv preprint arXiv:2505.13211, 2025. 2, 7
Pith/arXiv arXiv 2025
-
[59]
Shengbang Tong, Boyang Zheng, Ziteng Wang, Bingda Tang, Nanye Ma, Ellis Brown, Jihan Yang, Rob Fergus, Yann LeCun, and Saining Xie. Scaling text-to-image dif- fusion transformers with representation autoencoders.arXiv preprint arXiv:2601.16208, 2026. 2
arXiv 2026
-
[60]
Michael Tschannen, Alexey Gritsenko, Xiao Wang, Muham- mad Ferjad Naeem, Ibrahim Alabdulmohsin, Nikhil Parthasarathy, Talfan Evans, Lucas Beyer, Ye Xia, Basil Mustafa, et al. Siglip 2: Multilingual vision-language en- coders with improved semantic understanding, localization, and dense features.arXiv preprint arXiv:2502.14786, 2025. 2, 15
Pith/arXiv arXiv 2025
-
[61]
Image processing in python.CSI Communica- tions, 23(2), 2012
P Umesh. Image processing in python.CSI Communica- tions, 23(2), 2012. 6, 13
2012
-
[62]
Fvd: A new metric for video generation
Thomas Unterthiner, Sjoerd van Steenkiste, Karol Kurach, Rapha¨el Marinier, Marcin Michalski, and Sylvain Gelly. Fvd: A new metric for video generation. InDGS@ICLR,
-
[63]
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. InConference on Neural Information Processing Systems (NeurIPS), 2017. 2, 4
2017
-
[64]
Phenaki: Variable length video generation from open domain textual descriptions
Ruben Villegas, Mohammad Babaeizadeh, Pieter-Jan Kin- dermans, Hernan Moraldo, Han Zhang, Mohammad Taghi Saffar, Santiago Castro, Julius Kunze, and Dumitru Erhan. Phenaki: Variable length video generation from open domain textual descriptions. InInternational Conference on Learn- ing Representations (ICLR), 2023. 2
2023
-
[65]
Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314, 2025
Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianx- iao Yang, Jianyuan Zeng, Jiayu Wang, Jingfeng Zhang, Jin- gren Zhou, Jinkai Wang, Jixuan Chen, Kai Zhu, Kang Zhao, Keyu Yan, Lianghua Huang, Mengyang Feng, Ningyi Zhang, Pandeng Li, Pingyu Wu, Ruihang Chu, Ruili Feng, Shiwei Zhang, Siyang Sun, Tao Fan...
Pith/arXiv arXiv 2025
-
[66]
Omnitokenizer: A joint image- video tokenizer for visual generation
Junke Wang, Yi Jiang, Zehuan Yuan, Bingyue Peng, Zux- uan Wu, and Yu-Gang Jiang. Omnitokenizer: A joint image- video tokenizer for visual generation. InConference on Neu- ral Information Processing Systems (NeurIPS), 2024. 2, 7
2024
-
[67]
Enze Xie, Junsong Chen, Junyu Chen, Han Cai, Haotian Tang, Yujun Lin, Zhekai Zhang, Muyang Li, Ligeng Zhu, Yao Lu, et al. Sana: Efficient high-resolution image syn- thesis with linear diffusion transformers.arXiv preprint arXiv:2410.10629, 2024. 1
Pith/arXiv arXiv 2024
-
[68]
Videovae+: Large motion video autoencoding with cross-modal video vae
Yazhou Xing, Yang Fei, Yingqing He, Jingye Chen, Jiaxin Xie, Xiaowei Chi, and Qifeng Chen. Videovae+: Large motion video autoencoding with cross-modal video vae. In IEEE International Conference on Computer Vision (ICCV),
-
[69]
Zhucun Xue, Jiangning Zhang, Teng Hu, Haoyang He, Yinan Chen, Yuxuan Cai, Yabiao Wang, Chengjie Wang, Yong Liu, Xiangtai Li, and Dacheng Tao. Ultravideo: High-quality uhd video dataset with comprehensive captions.arXiv preprint arXiv:2506.13691, 2025. 6, 7
Pith/arXiv arXiv 2025
-
[70]
Cogvideox: Text-to-video diffusion models with an expert transformer
Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiao- han Zhang, Guanyu Feng, Da Yin, Yuxuan.Zhang, Weihan Wang, Yean Cheng, Bin Xu, Xiaotao Gu, Yuxiao Dong, and Jie Tang. Cogvideox: Text-to-video diffusion models with an expert transformer. InInternational Conference on Learning Representations (ICLR...
2025
-
[71]
Fasterdit: Towards faster diffusion transformers train- ing without architecture modification
Jingfeng Yao, Cheng Wang, Wenyu Liu, and Xinggang Wang. Fasterdit: Towards faster diffusion transformers train- ing without architecture modification. InConference on Neu- ral Information Processing Systems (NeurIPS), 2024. 1
2024
-
[72]
Reconstruc- tion vs
Jingfeng Yao, Bin Yang, and Xinggang Wang. Reconstruc- tion vs. generation: Taming optimization dilemma in latent diffusion models. InIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2025. 1
2025
-
[73]
Improved distribution matching distillation for fast image synthesis
Tianwei Yin, Micha ¨el Gharbi, Taesung Park, Richard Zhang, Eli Shechtman, Fredo Durand, and Bill Freeman. Improved distribution matching distillation for fast image synthesis. InConference on Neural Information Processing Systems (NeurIPS), 2024. 1
2024
-
[74]
From slow bidirectional to fast autoregressive video diffusion mod- els
Tianwei Yin, Qiang Zhang, Richard Zhang, William T Free- man, Fredo Durand, Eli Shechtman, and Xun Huang. From slow bidirectional to fast autoregressive video diffusion mod- els. InIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2025. 1
2025
-
[75]
Vector-quantized image modeling with im- proved VQGAN
Jiahui Yu, Xin Li, Jing Yu Koh, Han Zhang, Ruoming Pang, James Qin, Alexander Ku, Yuanzhong Xu, Jason Baldridge, and Yonghui Wu. Vector-quantized image modeling with im- proved VQGAN. InInternational Conference on Learning Representations (ICLR), 2022. 2
2022
-
[76]
Language model beats diffusion - tokenizer is key to visual generation
Lijun Yu, Jose Lezama, Nitesh Bharadwaj Gundavarapu, Luca Versari, Kihyuk Sohn, David Minnen, Yong Cheng, Agrim Gupta, Xiuye Gu, Alexander G Hauptmann, Boqing Gong, Ming-Hsuan Yang, Irfan Essa, David A Ross, and Lu Jiang. Language model beats diffusion - tokenizer is key to visual generation. InInternational Conference on Learning Representations (ICLR), 2024. 2
2024
-
[77]
Root mean square layer nor- malization
Biao Zhang and Rico Sennrich. Root mean square layer nor- malization. InConference on Neural Information Processing Systems (NeurIPS), 2019. 4
2019
-
[78]
The unreasonable effectiveness of deep features as a perceptual metric
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shecht- man, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. InIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018. 5, 6, 7, 13, 14
2018
-
[79]
Waver: Wave your way to lifelike video genera- tion.arXiv preprint arXiv:2508.15761, 2025
Yifu Zhang, Hao Yang, Yuqi Zhang, Yifei Hu, Fengda Zhu, Chuang Lin, Xiaofeng Mei, Yi Jiang, Bingyue Peng, and Ze- huan Yuan. Waver: Wave your way to lifelike video genera- tion.arXiv preprint arXiv:2508.15761, 2025. 1, 2
Pith/arXiv arXiv 2025
-
[80]
Cv- vae: A compatible video vae for latent generative video mod- els
Sijie Zhao, Yong Zhang, Xiaodong Cun, Shaoshu Yang, Muyao Niu, Xiaoyu Li, Wenbo Hu, and Ying Shan. Cv- vae: A compatible video vae for latent generative video mod- els. InConference on Neural Information Processing Sys- tems (NeurIPS), 2024. 2
2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.