REVIEW 4 major objections 4 minor 122 references
KVAE: Family of Tokenizers for Multimodal Generative Models
T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read KVAE claims its continuous-latent tokenizers for image, video, and audio match or surpass six frontier open-source tokenizers on reconstruction and generation metrics.
desk verdict Useful engineering report with released models, but the 'matches or surpasses' headline overreaches: the tokenizer-swap comparisons are pipeline-specific and the audio alignment details are withheld. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing design choices are a continuous Gaussian bottleneck in every tokenizer (no vector quantization), an attention-free causal Conv3D stack for video with spatial RMSNorm and an asymmetric decoder, and, for audio, a DAC-derived convolutional backbone re-strided to [2,3,4,5,8] to reach 960x temporal compression, a single self-attention block at the 50 Hz bottleneck, and an alignment regularizer that pulls the latent toward a frozen audio foundation model. The paper's selection mechanism is the correlation decay slope (CDS), the negative slope of the cosine-similarity-versus-distance fit over the latent grid, used to screen candidates cheaply; final choices are always confirmed by training a generation model on top of the frozen tokenizer.
What would settle it
Train the same downstream generator at a very different scale (for example, an 8B DiT) on KVAE-4x16x16 versus HunyuanVideo-1.5, or on 128-channel versus 64-channel audio latents, under the same data and steps: the paper's own analysis predicts these rankings can shift with generator capacity, so a reversal at another scale would show the claimed superiority is conditional on the generator rather than intrinsic to the tokenizers.
Extended reading notes
Core claim
At its core the paper claims that a tokenizer's value for latent diffusion lies less in reconstruction fidelity than in 'diffusability'—how readily a diffusion model can learn to denoise its latent space—and that this property can be engineered and screened. Under a controlled protocol, a fixed downstream generator (Kandinsky-5's 2B DiT for image and video, a 0.6B DiT for audio) is trained on each tokenizer with data, captions, and steps held constant, so that quality differences are attributed to the latent space alone. With that protocol, KVAE-4x16x16 surpasses HunyuanVideo-1.5's tokenizer in text-to-video generation and is preferred in side-by-side evaluation; KVAE-Audio with 64 channels at 50 Hz is preferred over all three audio baselines on all three judged criteria; and KVAE-2D-2.0 leads its FLUX baselines on semantic quality at equal steps. The paper introduces the correlation decay slope (CDS), a spatial cosine-similarity decay measure, as a cheap screening statistic that correlated strongly (r=0.906) with subjective visual quality across 14 image-tokenizer configurations.
Load-bearing premise
The ranking rests on the assumption that the fixed in-house generator used for the tokenizer swap—Kandinsky-5's 2B DiT for image and video, a 0.6B DiT for audio—is neutral across different latent geometries (patch size, channel count, frame rate), so that every measured difference is attributable to the tokenizer rather than to interactions with the pipeline.
Editorial extensions
If this is right
- KVAE's released checkpoints can be dropped into existing latent-diffusion pipelines as open-source components, with the reported results suggesting they would not degrade and often improve generation quality.
- Higher compression ratios, such as 4x16x16 with 64 channels, bring faster convergence in the tested 2B image and video generator, which translates to shorter training runs and lower compute cost.
- A full-band 48 kHz audio tokenizer with a 50 Hz latent removes the need for a vocoder and makes joint text-to-video-and-audio generation possible from a single continuous latent space.
- The optimal number of latent channels is not intrinsic: the paper finds 64 channels best for audio at a 0.6B generator scale while 64 channels improves video, showing the choice must be made jointly with compression factor and downstream model size.
- CDS or similar latent diagnostics can screen tokenizer candidates before the costly step of training a full diffusion generator, as long as the correlation is re-checked on new configurations.
Reading between the lines
- At larger generator scales, the reported rankings may shift: the paper itself ties the 64-channel audio optimum to the 0.6B generator, so a wider audio latent could win with a bigger DiT.
- The headline 'matches or surpasses' is established with one in-house generator family; on different diffusion backbones the ranking may compress or reverse, so the claim is most credible for pipelines close to Kandinsky-5's.
- The CDS screening result is an in-sample correlation (n=14) and, as the paper notes, does not establish out-of-sample prediction; pre-registering CDS thresholds and testing them on new configurations would settle its value.
- Two stated gaps limit the audio result's reproducibility: the frozen alignment model and loss are deferred to another publication, and the FAD backbones run at 16–32 kHz, so the objective metrics cannot confirm the advertised full-band generation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript introduces KVAE, a family of continuous-latent tokenizers for image, video, and audio, designed for text-conditioned latent diffusion. For video, it presents two causal 3D tokenizers (4x8x8/16ch and 4x16x16/64ch); for image, a 2D tokenizer (8x8/32ch); for audio, a 48 kHz full-band waveform tokenizer with a 50 Hz/64-channel latent. The paper reports reconstruction metrics (PSNR, LPIPS, SSIM, PESQ, spectral distances) and generation evaluations (FID, CLIP, CLAP, FAD, side-by-side preference) against open tokenizers from Wan, HunyuanVideo, FLUX, MovieGen, StableAudio/SAME, and MMAudio, claiming these results 'match or surpass' the baselines. It also proposes a correlation-decay-slope (CDS) diffusability diagnostic and describes ablations on channel count, normalization, decoder width, attention, and perceptual losses. The final section concludes that the released tokenizers are competitive drop-in components for latent diffusion systems.
Significance. If the findings hold, the KVAE tokenizers are a useful open contribution: they are released with code and weights, the comparisons are carried out under a fixed downstream generator with external baselines, and the paper is unusually candid about caveats (e.g., the 0.6B-scale dependence of the audio channel optimum in Sec. 6.7, and the structural insensitivity of FAD backbones to the band above 16 kHz in Sec. 6.6). The controlled tokenizer-swap protocol is a valuable methodological template, and the explicit distinction between reconstruction fidelity, diffusability, and generation quality is a strength. However, the broad competitive claim is only partially supported by the presented evidence: the tokenizer swaps are not neutral to latent frame rate, channel count, and generator scale, and the central audio generation comparison lacks uncertainty quantification and specification of the alignment regularizer. The manuscript is therefore a solid technical report whose central claim needs either additional experiments or careful reframing.
major comments (4)
- [Sec. 6.6 and Sec. 6.7] The audio tokenizer-swap comparison of Table 4 does not isolate the latent space as claimed. The fixed 0.6B DiT is trained on MMAudio (40 channels, 43.07 Hz), KVAE-Audio (64 channels, 50 Hz), DAC-VAE MovieGen (128 channels, 25 Hz), and SAME-L (256 channels, about 10.8 Hz), so per second of audio the transformer processes approximately 43, 50, 25, or 11 tokens, respectively. The paper does not state how batch size or sequence length were equalized (e.g., seconds-per-batch versus tokens-per-batch), and attention cost and effective model capacity differ by up to a factor of about 4.6x across condition. Section 6.7 itself acknowledges that the optimal channel count is a property of the generator scale, so the reported ranking may be a property of the 0.6B pipeline rather than of the tokenizer itself. Please either re-run the comparison with equalized token budgets or at multiple generator scales, or restrict the claim to 'competitive under the Kandinsky-5 0.6B pipeline'.
- [Sec. 6.3 and Sec. 6.4] The audio model's training objective includes an alignment regularizer Lalign against a frozen audio foundation model F, but the identity of F and the form of Lalign are explicitly deferred to 'the dedicated publication' (Sec. 6.3 and Sec. 6.4). Since this term is part of the objective used to train the released checkpoints, the method is not reproducible from the manuscript and the contribution of the alignment term to the reported generation quality cannot be independently assessed. Please provide at least a precise specification of F and Lalign (or a complete pseudo-code description) in an appendix, or remove the alignment term from the reported model and retrain/re-evaluate without it.
- [Tables 3 and 4, Sec. 4.2-4.4] The headline comparative claims are based on point estimates without error bars, confidence intervals, or significance tests. For example, in Table 4 (Song Describer), the MMAudio baseline has a higher CLAP score (0.356 vs 0.339) and a lower FAD-PANNs (5.412 vs 7.971), and in Table 4 (LibriSpeech) DAC-VAE MovieGen has a higher CLAP (0.413 vs 0.389); in Table 3(c) EARS, the text calls a 0.31 dB SI-SDR deficit 'within noise' without defining the noise level. Similarly, the side-by-side evaluations in Sec. 4.2-4.4 and Sec. 6.6 report win rates and 'preferred over all three baselines on all three criteria' without stating the number of annotators, the number of prompts, the inter-annotator agreement, or the statistical test used. These omissions make it impossible to distinguish genuine differences from noise, especially for the subjective claims that the text says 'carry the main weight of the comparison.' Please add uncertainty quantification and the experimental protocol details, or temper the claims accordingly.
- [Sec. 4.2 and Sec. 4.3] The same neutrality concern applies to the visual tokenizer swaps. Section 4.2 states that patch size is adjusted to balance the number of spatial tokens, but the compared tokenizers still differ in channel count (e.g., 16 vs 64 for the 4x8x8 vs 4x16x16 comparison) and in temporal compression, and all image/video generation comparisons use a single 2B generator. The paper's own Sec. 5.2 ablation shows that channel count interacts with convergence and final quality, so the observed ranking may be specific to the Kandinsky-5 2B pipeline at the evaluated resolutions. To support the abstract's broad 'matches or surpasses frontier opensource tokenizers' claim, the visual comparisons should either include at least one additional generator scale or be explicitly framed as pipeline-specific evidence.
minor comments (4)
- [Sec. 4.4, Table 2] The OmniDoc-TokenBench table lists FID values (KVAE 1.74 vs FLUX.1-dev 0.554 and FLUX.2-dev 0.73) but the text only claims superiority on PSNR, SSIM, and NED; please define what FID is measuring in this reconstruction context and explain why the KVAE value is worse if the comparison is meant to support the 'surpassing' claim.
- [Sec. 5.1] The CDS analysis, while honestly caveated, is presented as 'crucial' for model selection despite being based on a single in-sample cross-sectional correlation (r=0.906 over 14 configurations) and one joint-training trajectory; please add an explicit out-of-sample test or clearly label CDS as a heuristic rather than a validated selection criterion.
- [Throughout] There are several typos and formatting inconsistencies: 'KV AE' is sometimes written 'KV AE' and sometimes 'KV AE-...' in text, 'charachteristic' appears in Sec. 3.1, 'HunyaunVideo' in Sec. 4.1, and the reference list contains a placeholder '[89] VERIFY author list' that must be resolved before publication.
- [Sec. 6.4] The crop-length schedule is presented as a design choice but the specific schedule (e.g., how many steps at 0.38 s, how the length increases to 5 s, and how the batch size is adjusted) is not reported; since the paper emphasizes sharing training details, please provide the exact schedule or a reference to the repository where it is defined.
Circularity Check
No significant circularity: the headline ranking rests on controlled tokenizer-swap experiments against external baselines rather than on fitted constants or self-citations.
full rationale
The derivation chain is self-contained: KVAE tokenizers are trained with reconstruction, adversarial, KL and alignment objectives (Secs. 3.3 and 6.4), and the central claims are evaluated by (a) reconstruction metrics on public benchmarks (MCL-JCV, OmniDoc-TokenBench, AudioSet, MUSDB18-HQ, EARS) and (b) generation via a controlled tokenizer swap in which one downstream generator is retrained per tokenizer with the same data, captions, architecture, and step counts (Secs. 4.2, 4.3, and 6.6). No parameter of KVAE is fitted to these benchmark scores and then reported as a prediction; the reported numbers are direct measurements. The CDS statistic used for candidate screening is explicitly flagged as an in-sample association that "does not establish out-of-sample prediction" (Sec. 5.1), so the paper does not present the CDS correlation as independent validation. The downstream generator (Kandinsky-5 [4]) and evaluation utilities ([49]) are self-citations, but they are evaluation infrastructure rather than the target result; the tokenizer comparisons are against third-party baselines with released weights (Wan, HunyuanVideo, FLUX, MMAudio, DAC-VAE MovieGen, SAME-L) under the same protocol. Potential concerns that the fixed-generator comparison is sensitive to latent token counts, channel counts, and generator scale (for example, the audio swap sees roughly 43, 50, 25, or 11 tokens per second across baselines, and Sec. 6.7 shows the optimal audio channel count depends on generator scale) are validity and fairness risks for the drop-in claim, not equation-level circularity. Similarly, selecting 64 audio channels via the Sec. 6.6 protocol before reporting Table 4 is a selection-on-evaluation risk, not a fitted-input-called-prediction reduction. No self-definitional reduction, imported uniqueness, ansatz-by-citation, or renaming pattern is present.
Assumptions & free parameters
free parameters (5)
- Latent channel counts =
16 and 64 (video), 32 (image), 64 (audio)
- Audio objective weights w_KL, w_adv, w_perc, w_align, w_spec =
Not reported numerically
- Audio stride ladder [2,3,4,5,8] =
Product 960, giving 50 Hz at 48 kHz
- CDS linear-fit window =
Manhattan distances delta = 1,...,8
- Audio crop-length schedule =
0.38 s to 5 s across training stages
assumptions (5)
- domain assumption Standard VAE, KL, GAN and flow-matching training objectives produce latents that are good substrates for diffusion.
- domain assumption The fixed downstream generator is a neutral instrument for comparing tokenizers.
- domain assumption CDS is a valid proxy for subjective generation quality.
- domain assumption The bandwidth filter correctly identifies files that are truly full-band 48 kHz.
- standard math Adam converges reliably for all training stages.
Cite this review
Pith. "Pith review of KVAE: Family of Tokenizers for Multimodal Generative Models." pith.science (2026). https://pith.science/paper/7RKTK473
@misc{pith2026260805798,
author = {Pith},
title = {Pith review of: KVAE: Family of Tokenizers for Multimodal Generative Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/7RKTK473}},
note = {Machine review of arXiv:2608.05798}
}
read the original abstract
Latent diffusion modeling (LDM), a prominent paradigm, utilizes tokenizers to map input signal to compressed representation. This dependency positions tokenizer as an integral part of generation process itself, since it affects learning speed, quality of synthesized samples and lay foundation for later applications. This report presents series of KVAE tokenizers for audio, image and video, all designed for subsequent text-conditioned generation: KVAE-Audio, a continuous full-band 48 kHz tokenizer with a 50 Hz latent of 64 channels; KVAE-3D -- two causal video tokenizers for 4x16x16 and 4x8x8 compression; KVAE-2D, an image model, compressing input by factor of 8 with 32 channels. We demonstrate that reconstruction (PSNR, LPIPS, PESQ, etc.) and generation results on objective (Frechet Distance, CLIP score, CLAP score, etc.) and subjective (side-by-side evaluation) metrics matches or surpasses frontier opensource tokenizers, such as VAEs from Wan-2.2, HunyuanVideo-1.5, FLUX.2, MovieGen, StableAudio and MMAudio. Considering difficulty of development, we share with community training details, model selection method and ablation on design choices. The code is publicly available at https://github.com/kandinskylab/kvae and https://github.com/kandinskylab/kvae-audio.
Figures
Figures from the paper (10 more)
Reference graph
Works this paper leans on
-
[1]
Bitdance: Scaling autoregressive generative models with binary tokens, 2026
Yuang Ai, Jiaming Han, Shaobin Zhuang, Weijia Mao, Xuefeng Hu, Ziyan Yang, Zhen- heng Yang, Yali Wang, Huaibo Huang, Xiangyu Yue, and Hao Chen. Bitdance: Scaling autoregressive generative models with binary tokens, 2026. 1
2026
-
[2]
OmniDoc-TokenBench
Alibaba Group. OmniDoc-TokenBench. https://github.com/alibaba/ OmniDoc-TokenBench, 2026. Official benchmark repository. 8
2026
-
[3]
AOM Common Test Conditions v5.0
Alliance for Open Media. AOM Common Test Conditions v5.0. Input Document CWG-D103o, Alliance for Open Media (AOMedia), 8 2023. Codec Working Group. 6
2023
-
[4]
Kandinsky 5.0: A family of foundation models for image and video generation,
Vladimir Arkhipkin, Vladimir Korviakov, Nikolai Gerasimenko, Denis Parkhomenko, Vi- acheslav Vasilev, Alexey Letunovskiy, Nikolai Vaulin, Maria Kovaleva, Ivan Kirillov, Lev Novitskiy, Denis Koposov, Nikita Kiselev, Alexander Varlamov, Dmitrii Mikhailov, Vladimir Polovnikov, Andrey Shutkin, Julia Agafonova, Ilya Vasiliev, Anastasiia Kargapoltseva, Anna Dmi...
-
[5]
Qwen2.5-vl technical report, 2025
Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. Qwen2.5-vl technical report, 2025. 6
2025
-
[7]
Stable video diffusion: Scaling latent video diffusion models to large datasets,
Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram V oleti, Adam Letts, Varun Jampani, and Robin Rombach. Stable video diffusion: Scaling latent video diffusion models to large datasets,
-
[8]
Align your latents: High-resolution video synthesis with latent diffusion models, 2023
Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your latents: High-resolution video synthesis with latent diffusion models, 2023. 3
2023
-
[9]
Bruinsma, Ana Lucic, Megan Stanley, Anna Vaughan, Johannes Brandstetter, Patrick Garvan, Maik Riechert, Jonathan A
Cristian Bodnar, Wessel P. Bruinsma, Ana Lucic, Megan Stanley, Anna Vaughan, Johannes Brandstetter, Patrick Garvan, Maik Riechert, Jonathan A. Weyn, Haiyu Dong, Jayesh K. Gupta, Kit Thambiratnam, Alexander T. Archibald, Chun-Chieh Wu, Elizabeth Heider, Max Welling, Richard E. Turner, and Paris Perdikaris. A foundation model for the earth system, 2024. 1
2024
Show all 122 references
-
[10]
Bradley and Milton E
Ralph A. Bradley and Milton E. Terry. Rank analysis of incomplete block designs: I. the method of paired comparisons.Biometrika, 39(3/4):324–345, 1952. https://doi.org/10. 2307/2334029. 9
1952
-
[11]
Efros, and Tero Karras
Tim Brooks, Janne Hellsten, Miika Aittala, Ting-Chun Wang, Timo Aila, Jaakko Lehtinen, Ming-Yu Liu, Alexei A. Efros, and Tero Karras. Generating long videos of dynamic scenes,
-
[12]
Video generation models as world simulators
Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, Clarence Ng, Ricky Wang, and Aditya Ramesh. Video generation models as world simulators. 2024. 3
2024
-
[13]
Deep compression autoencoder for efficient high-resolution diffusion models, 2025
Junyu Chen, Han Cai, Junsong Chen, Enze Xie, Shang Yang, Haotian Tang, Muyang Li, Yao Lu, and Song Han. Deep compression autoencoder for efficient high-resolution diffusion models, 2025. 3, 4
2025
-
[14]
Dc-videogen: Efficient video generation with deep compression video autoencoder, 2025
Junyu Chen, Wenkun He, Yuchao Gu, Yuyang Zhao, Jincheng Yu, Junsong Chen, Dongyun Zou, Yujun Lin, Zhekai Zhang, Muyang Li, Haocheng Xi, Ligeng Zhu, Enze Xie, Song Han, and Han Cai. Dc-videogen: Efficient video generation with deep compression video autoencoder, 2025. 2
2025
-
[15]
Dc-ae 1.5: Accelerating diffusion model convergence with structured latent space, 2025
Junyu Chen, Dongyun Zou, Wenkun He, Junsong Chen, Enze Xie, Song Han, and Han Cai. Dc-ae 1.5: Accelerating diffusion model convergence with structured latent space, 2025. 3
2025
-
[16]
Taming multimodal joint training for high-quality video-to-audio synthesis.arXiv preprint arXiv:2412.15322, 2024
Ho Kei Cheng, Masato Ishii, Akio Hayakawa, Takashi Shibuya, Alexander Schwing, and Yuki Mitsufuji. Taming multimodal joint training for high-quality video-to-audio synthesis.arXiv preprint arXiv:2412.15322, 2024. 12, 15
2024 arXiv
-
[17]
Chien, Liuzixuan Lin, Hai Nguyen, Varsha Rao, Tristan Sharma, and Rajini Wijayawardana
Andrew A. Chien, Liuzixuan Lin, Hai Nguyen, Varsha Rao, Tristan Sharma, and Rajini Wijayawardana. Reducing the carbon impact of generative ai inference (today and in 2035). In Proceedings of the 2nd Workshop on Sustainable Computer Systems (HotCarbon ’23), Boston, MA, USA, Jul...
2023
-
[18]
Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinso...
2022
-
[19]
Adversarial video generation on complex datasets, 2019
Aidan Clark, Jeff Donahue, and Karen Simonyan. Adversarial video generation on complex datasets, 2019. 2
2019
-
[20]
High fidelity neural audio compression.arXiv preprint arXiv:2210.13438, 2022
Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi. High fidelity neural audio compression.arXiv preprint arXiv:2210.13438, 2022. 12
2022 arXiv
-
[21]
Irc-gan: Introspective recurrent convolutional gan for text-to-video generation
Kangle Deng, Tianyi Fei, Xin Huang, and Yuxin Peng. Irc-gan: Introspective recurrent convolutional gan for text-to-video generation. InProceedings of the 28th International Joint Conference on Artificial Intelligence (IJCAI), pages 2216–2222, 2019. 2
2019
-
[22]
Taming transformers for high-resolution image synthesis, 2021
Patrick Esser, Robin Rombach, and Björn Ommer. Taming transformers for high-resolution image synthesis, 2021. 5
2021
-
[23]
Parker, C
Zach Evans, Julian D. Parker, C. J. Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons. Stable audio open.arXiv preprint arXiv:2407.14358, 2024. 12
2024 arXiv
-
[24]
Parker, Matthew Rice, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons
Zach Evans, Julian D. Parker, Matthew Rice, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons. Stable audio 3, 2026. 12, 15
2026
-
[25]
The prism hypothesis: Harmonizing semantic and pixel representations via unified autoencoding, 2026
Weichen Fan, Haiwen Diao, Quan Wang, Dahua Lin, and Ziwei Liu. The prism hypothesis: Harmonizing semantic and pixel representations via unified autoencoding, 2026. 3
2026
-
[26]
Gemmeke, Daniel P
Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Chan- ning Moore, Manoj Plakal, and Marvin Ritter. Audio Set: An ontology and human-labeled dataset for audio events. InIEEE International Conference on Acoustics, Speech and Signal Processing ...
2017
-
[27]
BigVGAN: A universal neural vocoder with large-scale training
Sang gil Lee, Wei Ping, Boris Ginsburg, Bryan Catanzaro, and Sungroh Yoon. BigVGAN: A universal neural vocoder with large-scale training. InThe Eleventh International Conference on Learning Representations (ICLR), 2023. 12 20
2023
-
[28]
Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial networks, 2014. 2
2014
-
[29]
Veo 3.1: Our leading video generation model
Google DeepMind. Veo 3.1: Our leading video generation model. Google DeepMind Official Product Page, October 2025. Accessed: 2026-07-13. 3
2025
-
[30]
Ltx-2: Efficient joint audio-visual foundation model, 2026
Yoav HaCohen, Benny Brazowski, Nisan Chiprut, Yaki Bitterman, Andrew Kvochko, Avishai Berkowitz, Daniel Shalem, Daphna Lifschitz, Dudu Moshe, Eitan Porat, Eitan Richardson, Guy Shiran, Itay Chachy, Jonathan Chetboun, Michael Finkelson, Michael Kupchick, Nir Zabari, Nitzan Guet...
2026
-
[31]
Shawn Hershey, Sourish Chaudhuri, Daniel P. W. Ellis, Jort F. Gemmeke, Aren Jansen, R. Channing Moore, Manoj Plakal, Devin Platt, Rif A. Saurous, Bryan Seybold, Malcolm Slaney, Ron J. Weiss, and Kevin Wilson. CNN architectures for large-scale audio classification. InIEEE Inter...
-
[32]
Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 30, 2017
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 30, 2017. 6
2017
-
[33]
Kingma, Ben Poole, Mohammad Norouzi, David J
Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P. Kingma, Ben Poole, Mohammad Norouzi, David J. Fleet, and Tim Salimans. Imagen video: High definition video generation with diffusion models, 2022. 3
2022
-
[34]
Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J. Fleet. Video diffusion models, 2022. 3
2022
-
[35]
Cogvideo: Large-scale pretraining for text-to-video generation via transformers, 2022
Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang. Cogvideo: Large-scale pretraining for text-to-video generation via transformers, 2022. 2
2022
-
[36]
Tangoflux: Super fast and faithful text to audio generation with flow matching and clap-ranked preference optimization,
Chia-Yu Hung, Navonil Majumder, Zhifeng Kong, Ambuj Mehrish, Amir Ali Bagherzadeh, Chuan Li, Rafael Valle, Bryan Catanzaro, and Soujanya Poria. Tangoflux: Super fast and faithful text to audio generation with flow matching and clap-ranked preference optimization,
-
[37]
Perceptual evaluation of speech quality (PESQ).International Telecommunication Union, 2001
ITU-T Recommendation P.862. Perceptual evaluation of speech quality (PESQ).International Telecommunication Union, 2001. 15
2001
-
[38]
Video pixel networks, 2016
Nal Kalchbrenner, Aaron van den Oord, Karen Simonyan, Ivo Danihelka, Oriol Vinyals, Alex Graves, and Koray Kavukcuoglu. Video pixel networks, 2016. 2
2016
-
[39]
Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms.Interspeech,
Kevin Kilgour, Mauricio Zuluaga, Dominik Roblek, and Matthew Sharifi. Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms.Interspeech,
-
[40]
AudioCaps: Generating captions for audios in the wild
Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim. AudioCaps: Generating captions for audios in the wild. InProceedings of NAACL-HLT, 2019. 16
2019
-
[41]
Kingma and Jimmy Ba
Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In International Conference on Learning Representations (ICLR), 2015. 5, 14
2015
-
[42]
Auto-encoding variational bayes, 2022
Diederik P Kingma and Max Welling. Auto-encoding variational bayes, 2022. 3
2022
-
[43]
Klingai enters the 3.0 era: All in one, one for all! kling 3.0 model now fully rolled out
Kling AI. Klingai enters the 3.0 era: All in one, one for all! kling 3.0 model now fully rolled out. Kling AI Official Release Notes, February 2026. Accessed: 2026-07-13. 3
2026
-
[44]
Carbon Emissions in the Tailpipe of Gen- erative AI.Harvard Data Science Review, 15(Special Issue 5), aug 20 2024
Tamara Kneese and Meg Young. Carbon Emissions in the Tailpipe of Gen- erative AI.Harvard Data Science Review, 15(Special Issue 5), aug 20 2024. https://hdsr.mitpress.mit.edu/pub/fscsqwx4. 2 21
2024
-
[45]
Ross, Bryan Seybold, and Lu Jiang
Dan Kondratyuk, Lijun Yu, Xiuye Gu, José Lezama, Jonathan Huang, Grant Schindler, Rachel Hornung, Vighnesh Birodkar, Jimmy Yan, Ming-Chang Chiu, Krishna Somandepalli, Hassan Akbari, Yair Alon, Yong Cheng, Josh Dillon, Agrim Gupta, Meera Hahn, Anja Hauth, David Hendon, Alonso M...
-
[46]
HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae. HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis. InAdvances in Neural Information Processing Systems, 2020. 14
2020
-
[47]
Plumb- ley
Qiuqiang Kong, Yin Cao, Turab Iqbal, Yuxuan Wang, Wenwu Wang, and Mark D. Plumb- ley. PANNs: Large-scale pretrained audio neural networks for audio pattern recognition. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 28:2880–2894, 2020. 16
2020
-
[48]
Hunyuanvideo: A systematic framework for large video generative models,
Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, Kathrina Wu, Qin Lin, Junkun Yuan, Yanxin Long, Aladdin Wang, Andong Wang, Changlin Li, Duojun Huang, Fang Yang, Hao Tan, Hongmei Wang, Jacob Song, Jiawang Bai, J...
-
[49]
Kandinsky video tools
Denis Koposov, Anna Dmitrienko, Ivan Kirillov, Kirill Chernyshev, Denis Parkhomenko, and Vladimir Korviakov. Kandinsky video tools. https://github.com/gen-ai-team/ kandinsky-video-tools, 2025. 7
2025
-
[50]
Efficient training of audio transformers with patchout
Khaled Koutini, Jan Schlüter, Hamid Eghbal-zadeh, and Gerhard Widmer. Efficient training of audio transformers with patchout. InInterspeech, 2022. 16
2022
-
[51]
Eq-vae: Equivariance regularized latent space for improved generative image modeling, 2025
Theodoros Kouzelis, Ioannis Kakogeorgiou, Spyros Gidaris, and Nikos Komodakis. Eq-vae: Equivariance regularized latent space for improved generative image modeling, 2025. 5
2025
-
[52]
High-fidelity audio compression with improved RVQGAN
Rithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar, and Kundan Kumar. High-fidelity audio compression with improved RVQGAN. InAdvances in Neural Information Processing Systems, 2023. 12, 13, 14
2023
-
[53]
REPA-E: Unlocking vae for end-to-end tuning of latent diffusion transformers
Xingjian Leng, Jaskirat Singh, Yunzhong Hou, Zhenchang Xing, Saining Xie, and Liang Zheng. REPA-E: Unlocking vae for end-to-end tuning of latent diffusion transformers. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 18262–18272, 2025. https://...
2025
-
[54]
Diffusionbench: On holistic evaluation of diffusion transformers with a unified training framework bridging imagenet and text-to-image, 2026
Xingjian Leng, Jaskirat Singh, Zhanhao Liang, Ethan Smith, Martin Bell, Aninda Saha, Yuhui Yuan, and Liang Zheng. Diffusionbench: On holistic evaluation of diffusion transformers with a unified training framework bridging imagenet and text-to-image, 2026. https://arxiv. org/ab...
2026 arXiv
-
[55]
Wf-vae: Enhancing video vae by wavelet-driven energy flow for latent video diffusion model,
Zongjian Li, Bin Lin, Yang Ye, Liuhan Chen, Xinhua Cheng, Shenghai Yuan, and Li Yuan. Wf-vae: Enhancing video vae by wavelet-driven energy flow for latent video diffusion model,
-
[56]
Generating novel, designable, and diverse protein structures by equivariantly diffusing oriented residue clouds, 2023
Yeqing Lin and Mohammed AlQuraishi. Generating novel, designable, and diverse protein structures by equivariantly diffusing oriented residue clouds, 2023. 1
2023
-
[57]
Plumbley
Haohe Liu, Yi Yuan, Xubo Liu, Xinhao Mei, Qiuqiang Kong, Qiao Tian, Yuping Wang, Wenwu Wang, Yuxuan Wang, and Mark D. Plumbley. AudioLDM 2: Learning holistic audio generation with self-supervised pretraining.arXiv preprint arXiv:2308.05734, 2023. 12 22
2023 arXiv
-
[58]
Delving into latent spectral biasing of video vaes for superior diffusability.arXiv preprint arXiv:2512.05394, 2025.https://arxiv.org/abs/2512.05394
Shizhan Liu, Xinran Deng, Zhuoyi Yang, Jiayan Teng, Xiaotao Gu, and Jie Tang. Delving into latent spectral biasing of video vaes for superior diffusability.arXiv preprint arXiv:2512.05394, 2025.https://arxiv.org/abs/2512.05394. 3, 9
2025 arXiv
-
[59]
Di Ma, Fan Zhang, and David R. Bull. Bvi-dvc: A training database for deep video compres- sion.IEEE Transactions on Multimedia, 24:3847–3858, 2022. 6
2022
-
[60]
The song describer dataset: A corpus of audio captions for music-and- language evaluation
Ilaria Manco, Benno Weck, Seungheon Doh, Minz Won, Yixiao Zhang, Dmitry Bogdanov, Yusong Wu, Ke Chen, Philip Tovstogan, Emmanouil Benetos, Elio Quinton, György Fazekas, and Juhan Nam. The song describer dataset: A corpus of audio captions for music-and- language evaluation. In...
2023
-
[61]
Balasubramanian
Gaurav Mittal, Tanya Marwah, and Vineeth N. Balasubramanian. Sync-draw: Automatic video generation using deep recurrent attentive architectures. InProceedings of the 25th ACM international conference on Multimedia, MM ’17, page 1096–1104. ACM, October 2017. 2
2017
-
[62]
Transition matching distillation for fast video generation, 2026
Weili Nie, Julius Berner, Nanye Ma, Chao Liu, Saining Xie, and Arash Vahdat. Transition matching distillation for fast video generation, 2026. 2
2026
-
[63]
Blaschko, Albert Ali Salah, and Itir Onal Ertugrul
Mang Ning, Mingxiao Li, Le Zhang, Lanmiao Liu, Matthew B. Blaschko, Albert Ali Salah, and Itir Onal Ertugrul. Spectrum matching: a unified perspective for superior diffusability in latent diffusion.arXiv preprint arXiv:2603.14645, 2026. https://arxiv.org/abs/2603.14645. 3, 9
2026
-
[64]
Alpamayo-r1: Bridging reasoning and action prediction for generalizable autonomous driving in the long tail, 2026
NVIDIA, :, Yan Wang, Wenjie Luo, Junjie Bai, Yulong Cao, Tong Che, Ke Chen, Yuxiao Chen, Jenna Diamond, Yifan Ding, Wenhao Ding, Liang Feng, Greg Heinrich, Jack Huang, Peter Karkus, Boyi Li, Pinyi Li, Tsung-Yi Lin, Dongran Liu, Ming-Yu Liu, Langechuan Liu, Zhijian Liu, Jason L...
2026
-
[65]
Cosmos tokenizer: A suite of image and video neural tokenizers.arXiv preprint arXiv:2501.03575, 2025
NVIDIA, Fitsum Reda, Jinwei Gu, Xian Liu, Songwei Ge, Ting-Chun Wang, Haoxiang Wang, and Ming-Yu Liu. Cosmos tokenizer: A suite of image and video neural tokenizers.arXiv preprint arXiv:2501.03575, 2025. 3, 4, 6
2025 arXiv
-
[66]
To create what you tell: Generating videos from captions, 2018
Yingwei Pan, Zhaofan Qiu, Ting Yao, Houqiang Li, and Tao Mei. To create what you tell: Generating videos from captions, 2018. 2
2018
-
[67]
Librispeech: An ASR corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. Librispeech: An ASR corpus based on public domain audio books. InIEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2015. 16
2015
-
[68]
Parker, Zach Evans, CJ Carr, Zack Zukowski, Josiah Taylor, Matthew Rice, and Jordi Pons
Julian D. Parker, Zach Evans, CJ Carr, Zack Zukowski, Josiah Taylor, Matthew Rice, and Jordi Pons. SAME: A semantically-aligned music autoencoder, 2026. 12, 15
2026
-
[69]
Sampson, Shikai Li, Simone Parmeggiani, Steve Fine, Tara Fowler, Vladan Petrovic, and Yuming Du
Adam Polyak, Amit Zohar, Andrew Brown, Andros Tjandra, Animesh Sinha, Ann Lee, Apoorv Vyas, Bowen Shi, Chih-Yao Ma, Ching-Yao Chuang, David Yan, Dhruv Choudhary, Dingkang Wang, Geet Sethi, Guan Pang, Haoyu Ma, Ishan Misra, Ji Hou, Jialiang Wang, Kiran Jagadeesh, Kunpeng Li, Lu...
2025
-
[70]
Andersson, Andrew El-Kadi, Dominic Masters, Timo Ewalds, Jacklynn Stott, Shakir Mohamed, Peter Battaglia, Remi Lam, and Matthew Willson
Ilan Price, Alvaro Sanchez-Gonzalez, Ferran Alet, Tom R. Andersson, Andrew El-Kadi, Dominic Masters, Timo Ewalds, Jacklynn Stott, Shakir Mohamed, Peter Battaglia, Remi Lam, and Matthew Willson. Gencast: Diffusion-based ensemble forecasting for medium-range weather, 2024. 1
2024
-
[71]
Qwen-Audio-V AE technical report.arXiv preprint arXiv:2607.11738, 2026
Qwen Team. Qwen-Audio-V AE technical report.arXiv preprint arXiv:2607.11738, 2026. 12
2026 arXiv
-
[72]
Qwen-Image-V AE-2.0 technical report.arXiv preprint arXiv:2605.13565, 2026
Qwen Team. Qwen-Image-V AE-2.0 technical report.arXiv preprint arXiv:2605.13565, 2026. https://arxiv.org/abs/2605.13565. 8
2026 arXiv
-
[73]
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. InInternational conference on machine learning, pag...
2021
-
[74]
MUSDB18-HQ — an uncompressed version of MUSDB18, 2019
Zafar Rafii, Antoine Liutkus, Fabian-Robert Stöter, Stylianos Ioannis Mimilakis, and Rachel Bittner. MUSDB18-HQ — an uncompressed version of MUSDB18, 2019. Zenodo. 14
2019
-
[75]
EARS: An anechoic fullband speech dataset benchmarked for speech enhancement and dereverberation
Julius Richter, Yi-Chiao Wu, Steven Krenn, Simon Welker, Bunlong Lay, Shinji Watanabe, Alexander Richard, and Timo Gerkmann. EARS: An anechoic fullband speech dataset benchmarked for speech enhancement and dereverberation. InInterspeech, 2024. 15
2024
-
[76]
High-resolution image synthesis with latent diffusion models, 2022
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models, 2022. 2
2022
-
[77]
High-resolution image synthesis with latent diffusion models, 2022
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models, 2022. 3
2022
-
[78]
Runway gen-4: Ai video generation with world consistency
Runway Research. Runway gen-4: Ai video generation with world consistency. Runway Research Publications, March 2025. Accessed: 2026-07-13. 3
2025
-
[79]
Flow to the mode: Mode-seeking diffusion autoencoders for state-of-the-art image tokenization, 2025
Kyle Sargent, Kyle Hsu, Justin Johnson, Li Fei-Fei, and Jiajun Wu. Flow to the mode: Mode-seeking diffusion autoencoders for state-of-the-art image tokenization, 2025. 3
2025
-
[80]
Make- a-video: Text-to-video generation without text-video data, 2022
Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, Devi Parikh, Sonal Gupta, and Yaniv Taigman. Make- a-video: Text-to-video generation without text-video data, 2022. 3
2022
-
[81]
What matters for representation alignment: Global information or spatial structure?arXiv preprint arXiv:2512.10794, 2025
Jaskirat Singh, Xingjian Leng, Zongze Wu, Liang Zheng, Richard Zhang, Eli Shechtman, and Saining Xie. What matters for representation alignment: Global information or spatial structure?arXiv preprint arXiv:2512.10794, 2025. https://arxiv.org/abs/2512.10794. 9
2025
-
[82]
Improving the diffusability of autoencoders, 2025
Ivan Skorokhodov, Sharath Girish, Benran Hu, Willi Menapace, Yanyu Li, Rameen Abdal, Sergey Tulyakov, and Aliaksandr Siarohin. Improving the diffusability of autoencoders, 2025. 2, 3, 18
2025
-
[83]
Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2, 2022
Ivan Skorokhodov, Sergey Tulyakov, and Mohamed Elhoseiny. Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2, 2022. 2
2022
-
[84]
Ucf101: A dataset of 101 human actions classes from videos in the wild, 2012
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. Ucf101: A dataset of 101 human actions classes from videos in the wild, 2012. 2
2012
-
[85]
Scenediffuser++: City-scale traffic simulation via a generative world model, 2025
Shuhan Tan, John Lambert, Hong Jeon, Sakshum Kulshrestha, Yijing Bai, Jing Luo, Dragomir Anguelov, Mingxing Tan, and Chiyu Max Jiang. Scenediffuser++: City-scale traffic simulation via a generative world model, 2025. 1
2025
-
[86]
Z-image: An efficient image generation foundation model with single-stream diffusion transformer,
Image Team, Huanqia Cai, Sihan Cao, Ruoyi Du, Peng Gao, Steven Hoi, Zhaohui Hou, Shijie Huang, Dengyang Jiang, Xin Jin, Liangchen Li, Zhen Li, Zhong-Yu Li, David Liu, Dongyang Liu, Junhan Shi, Qilong Wu, Feng Yu, Chi Zhang, Shifeng Zhang, and Shilin Zhou. Z-image: An efficient...
-
[87]
Lyria Team, Antoine Caillon, Brian McWilliams, Cassie Tarakajian, Ian Simon, Ilaria Manco, Jesse Engel, Noah Constant, Yunpeng Li, Timo I. Denk, Alberto Lalama, Andrea Agostinelli, Cheng-Zhi Anna Huang, Ethan Manilow, George Brower, Hakan Erdogan, Heidi Lei, Itai Rolnick, Ivan...
2025
-
[88]
Nextstep-1: Toward autoregressive image generation with continuous tokens at scale, 2025
NextStep Team, Chunrui Han, Guopeng Li, Jingwei Wu, Quan Sun, Yan Cai, Yuang Peng, Zheng Ge, Deyu Zhou, Haomiao Tang, Hongyu Zhou, Kenkun Liu, Ailin Huang, Bin Wang, Changxin Miao, Deshan Sun, En Yu, Fukun Yin, Gang Yu, Hao Nie, Haoran Lv, Hanpeng Hu, Jia Wang, Jian Zhou, Jian...
2025
-
[89]
HunyuanVideo-Foley: Multimodal diffusion with representation alignment for high-fidelity foley audio generation.arXiv preprint arXiv:2508.16930, 2025
Tencent Hunyuan. HunyuanVideo-Foley: Multimodal diffusion with representation alignment for high-fidelity foley audio generation.arXiv preprint arXiv:2508.16930, 2025. VERIFY author list. 12
2025 arXiv
-
[90]
Reducio! generating 1k video within 16 seconds using extremely compressed motion latents, 2025
Rui Tian, Qi Dai, Jianmin Bao, Kai Qiu, Yifan Yang, Chong Luo, Zuxuan Wu, and Yu-Gang Jiang. Reducio! generating 1k video within 16 seconds using extremely compressed motion latents, 2025. 3
2025
-
[91]
Metaxas, and Sergey Tulyakov
Yu Tian, Jian Ren, Menglei Chai, Kyle Olszewski, Xi Peng, Dimitris N. Metaxas, and Sergey Tulyakov. A good image generator is what you need for high-resolution video synthesis, 2021. 2
2021
-
[92]
Meta audiobox aesthetics: Unified automatic quality assessment for speech, music, and sound
Andros Tjandra, Yi-Chiao Wu, Baishan Guo, John Hoffman, Brian Ellis, Apoorv Vyas, Bowen Shi, Sanyuan Chen, Matt Le, Nick Zacharov, Carleigh Wood, Ann Lee, and Wei-Ning Hsu. Meta audiobox aesthetics: Unified automatic quality assessment for speech, music, and sound. arXiv prepr...
2025 arXiv
-
[93]
Ssdd: Single-step diffusion decoder for efficient image tokenization, 2026
Théophane Vallaeys, Jakob Verbeek, and Matthieu Cord. Ssdd: Single-step diffusion decoder for efficient image tokenization, 2026. 3
2026
-
[94]
Conditional image generation with pixelcnn decoders, 2016
Aaron van den Oord, Nal Kalchbrenner, Oriol Vinyals, Lasse Espeholt, Alex Graves, and Koray Kavukcuoglu. Conditional image generation with pixelcnn decoders, 2016. 2
2016
-
[95]
Neural discrete representation learning, 2018
Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. Neural discrete representation learning, 2018. 2
2018
-
[96]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need, 2023. 2
2023
-
[97]
Generating videos with scene dynamics, 2016
Carl V ondrick, Hamed Pirsiavash, and Antonio Torralba. Generating videos with scene dynamics, 2016. 2
2016
-
[98]
Wan: Open and advanced large-scale video generative models, 2025
Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, Jianyuan Zeng, Jiayu Wang, Jingfeng Zhang, Jingren Zhou, Jinkai Wang, Jixuan Chen, Kai Zhu, Kang Zhao, Keyu Yan, Lianghua Huang, Mengyang Feng, Ningyi Zhang, Pande...
2025
-
[99]
Wan-2.2 anouncement
Wan Team. Wan-2.2 anouncement. https://wan.video/blog/wan2.2, July 2025. Ac- cessed: 2026-06-01. 3, 4, 5, 11 25
2025
-
[100]
Haiqiang Wang, Weihao Gan, Sudeng Hu, Joe Yuchieh Lin, Lina Jin, Longguang Song, Ping Wang, Ioannis Katsavounidis, Anne Aaron, and C.-C. Jay Kuo. Mcl-jcv: A jnd-based h.264/avc video quality assessment dataset. In2016 IEEE International Conference on Image Processing (ICIP), p...
2016
-
[101]
Videomae v2: Scaling video masked autoencoders with dual masking
Limin Wang, Bingkun Huang, Zhiyu Zhao, Zhan Tong, Yinan He, Yi Wang, Yali Wang, and Yu Qiao. Videomae v2: Scaling video masked autoencoders with dual masking. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 14549–14560,
-
[102]
Emu3: Next-token prediction is all you need, 2024
Xinlong Wang, Xiaosong Zhang, Zhengxiong Luo, Quan Sun, Yufeng Cui, Jinsheng Wang, Fan Zhang, Yueze Wang, Zhen Li, Qiying Yu, Yingli Zhao, Yulong Ao, Xuebin Min, Tao Li, Boya Wu, Bo Zhao, Bowen Zhang, Liangdong Wang, Guang Liu, Zheqi He, Xi Yang, Jingjing Liu, Yonghua Lin, Tie...
2024
-
[103]
Internvideo2: Scaling foundation models for multimodal video understanding
Yi Wang, Kunchang Li, Xinhao Li, Jiashuo Yu, Yinan He, Guo Chen, Baoqi Pei, Rongkun Zheng, Zun Wang, Yansong Shi, et al. Internvideo2: Scaling foundation models for multimodal video understanding. InEuropean conference on computer vision, pages 396–416. Springer,
-
[104]
Hunyuanvideo 1.5 technical report, 2025
Bing Wu, Chang Zou, Changlin Li, Duojun Huang, Fang Yang, Hao Tan, Jack Peng, Jianbing Wu, Jiangfeng Xiong, Jie Jiang, Linus, Patrol, Peizhen Zhang, Peng Chen, Penghao Zhao, Qi Tian, Songtao Liu, Weijie Kong, Weiyan Wang, Xiao He, Xin Li, Xinchi Deng, Xuefei Zhe, Yang Li, Yanx...
2025
-
[105]
Q-align: Teaching lmms for visual scoring via discrete text-defined levels
Haoning Wu, Zicheng Zhang, Weixia Zhang, Chaofeng Chen, Liang Liao, Chunyi Li, Yixuan Gao, Annan Wang, Erli Zhang, Wenxiu Sun, et al. Q-align: Teaching lmms for visual scoring via discrete text-defined levels. InInternational Conference on Machine Learning, pages 54015–54029. ...
2024
-
[106]
H3ae: High compression, high speed, and high quality autoencoder for video diffusion models, 2025
Yushu Wu, Yanyu Li, Ivan Skorokhodov, Anil Kag, Willi Menapace, Sharath Girish, Aliaksandr Siarohin, Yanzhi Wang, and Sergey Tulyakov. H3ae: High compression, high speed, and high quality autoencoder for video diffusion models, 2025. 4
2025
-
[107]
Large-scale contrastive language-audio pretraining with feature fusion and keyword- to-caption augmentation
Yusong Wu, Ke Chen, Tianyu Zhang, Yuchen Hui, Taylor Berg-Kirkpatrick, and Shlomo Dubnov. Large-scale contrastive language-audio pretraining with feature fusion and keyword- to-caption augmentation. InIEEE International Conference on Acoustics, Speech and Signal Processing (IC...
2023
-
[108]
Grok imagine API
xAI. Grok imagine API. https://x.ai/news/grok-imagine-api , January 2026. Ac- cessed: 2026-06-01. 3
2026
-
[109]
Learning to generate time-lapse videos using multi-stage dynamic generative adversarial networks, 2018
Wei Xiong, Wenhan Luo, Lin Ma, Wei Liu, and Jiebo Luo. Learning to generate time-lapse videos using multi-stage dynamic generative adversarial networks, 2018. 2
2018
-
[110]
Videogpt: Video generation using vq-vae and transformers, 2021
Wilson Yan, Yunzhi Zhang, Pieter Abbeel, and Aravind Srinivas. Videogpt: Video generation using vq-vae and transformers, 2021. 2
2021
-
[111]
Qwen3 technical report, 2025
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jia...
2025
-
[112]
Cogvideox: Text-to-video diffusion models with an expert transformer, 2025
Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, Da Yin, Yuxuan Zhang, Weihan Wang, Yean Cheng, Bin Xu, Xiaotao Gu, Yuxiao Dong, and Jie Tang. Cogvideox: Text-to-video diffusion models with an ex...
2025
-
[113]
Reconstruction vs
Jingfeng Yao, Bin Yang, and Xinggang Wang. Reconstruction vs. generation: Taming opti- mization dilemma in latent diffusion models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025. https://arxiv.org/abs/2501.01423. 3, 10, 12, 18
2025 arXiv
-
[114]
Reconstruction vs
Jingfeng Yao, Bin Yang, and Xinggang Wang. Reconstruction vs. generation: Taming opti- mization dilemma in latent diffusion models, 2025. 5
2025
-
[115]
Gundavarapu, Luca Versari, Kihyuk Sohn, David Minnen, Yong Cheng, Vighnesh Birodkar, Agrim Gupta, Xiuye Gu, Alexander G
Lijun Yu, José Lezama, Nitesh B. Gundavarapu, Luca Versari, Kihyuk Sohn, David Minnen, Yong Cheng, Vighnesh Birodkar, Agrim Gupta, Xiuye Gu, Alexander G. Hauptmann, Boqing Gong, Ming-Hsuan Yang, Irfan Essa, David A. Ross, and Lu Jiang. Language model beats diffusion – tokenize...
2024
-
[116]
Representation alignment for generation: Training diffusion transformers is easier than you think.arXiv preprint arXiv:2410.06940, 2024
Sihyun Yu, Sangkyung Kwak, Huiwon Jang, Jongheon Jeong, Jonathan Huang, Jinwoo Shin, and Saining Xie. Representation alignment for generation: Training diffusion transformers is easier than you think.arXiv preprint arXiv:2410.06940, 2024. 12
-
[117]
SoundStream: An end-to-end neural audio codec
Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi. SoundStream: An end-to-end neural audio codec. InIEEE/ACM Transactions on Audio, Speech, and Language Processing, 2021. 12
2021
-
[118]
Gonzalez, Jianfei Chen, and Jun Zhu
Jintao Zhang, Kaiwen Zheng, Kai Jiang, Haoxu Wang, Ion Stoica, Joseph E. Gonzalez, Jianfei Chen, and Jun Zhu. Turbodiffusion: Accelerating video diffusion models by 100-200 times,
-
[119]
Efros, Eli Shechtman, and Oliver Wang
Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shechtman, and Oliver Wang. The unrea- sonable effectiveness of deep features as a perceptual metric, 2018. 5
2018
-
[120]
Maisi-v2: Accelerated 3d high-resolution medical image synthesis with rectified flow and region-specific contrastive loss, 2025
Can Zhao, Pengfei Guo, Dong Yang, Yucheng Tang, Yufan He, Benjamin Simon, Mason Belue, Stephanie Harmon, Baris Turkbey, and Daguang Xu. Maisi-v2: Accelerated 3d high-resolution medical image synthesis with rectified flow and region-specific contrastive loss, 2025. 1
2025
-
[121]
Cv-vae: A compatible video vae for latent generative video models, 2024
Sijie Zhao, Yong Zhang, Xiaodong Cun, Shaoshu Yang, Muyao Niu, Xiaoyu Li, Wenbo Hu, and Ying Shan. Cv-vae: A compatible video vae for latent generative video models, 2024. 3
2024
-
[122]
Diffusion transformers with representation autoencoders, 2025
Boyang Zheng, Nanye Ma, Shengbang Tong, and Saining Xie. Diffusion transformers with representation autoencoders, 2025. 3, 5, 12
2025
-
[123]
Diffusing in the right space: A systematic study of latent diffusability.arXiv preprint arXiv:2606.03578, 2026
Tianxiong Zhong, Xingye Tian, Xuebo Wang, Xin Tao, and Pengfei Wan. Diffusing in the right space: A systematic study of latent diffusability.arXiv preprint arXiv:2606.03578, 2026. https://arxiv.org/abs/2606.03578. 9 27
2026 arXiv
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.