Pith. sign in

REVIEW 4 major objections 4 minor 122 references

KVAE: Family of Tokenizers for Multimodal Generative Models

T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read KVAE claims its continuous-latent tokenizers for image, video, and audio match or surpass six frontier open-source tokenizers on reconstruction and generation metrics.

desk verdict Useful engineering report with released models, but the 'matches or surpasses' headline overreaches: the tokenizer-swap comparisons are pipeline-specific and the audio alignment details are withheld. read the letter →

arxiv 2608.05798 v1 pith:7RKTK473 submitted 2026-08-06 cs.CV cs.LGcs.SD

classification cs.CVcs.LGcs.SD
keywords latentdiffusiontokenizervariationalautoencodertext-to-imagegenerationtext-to-videotext-to-audiodiffusabilitycorrelationdecayslope
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper presents KVAE, a family of continuous-latent tokenizers for audio, image, and video, all built for text-conditioned latent diffusion. It claims that these tokenizers match or surpass the released tokenizers of six frontier open-source systems—Wan-2.2, HunyuanVideo-1.5, FLUX.2, MovieGen, StableAudio, and MMAudio—on both reconstruction and generation quality, measured objectively (FID, CLIP, CLAP, Fréchet distance, PESQ) and by human side-by-side preference. The models span a 48 kHz full-band audio tokenizer with a 50 Hz latent, two causal video tokenizers at 4x8x8 and 4x16x16 compression, and an 8x8 image tokenizer. The authors also share training details, a model-selection method built on a correlation-decay statistic, and ablations of design choices.

What carries the argument

The load-bearing design choices are a continuous Gaussian bottleneck in every tokenizer (no vector quantization), an attention-free causal Conv3D stack for video with spatial RMSNorm and an asymmetric decoder, and, for audio, a DAC-derived convolutional backbone re-strided to [2,3,4,5,8] to reach 960x temporal compression, a single self-attention block at the 50 Hz bottleneck, and an alignment regularizer that pulls the latent toward a frozen audio foundation model. The paper's selection mechanism is the correlation decay slope (CDS), the negative slope of the cosine-similarity-versus-distance fit over the latent grid, used to screen candidates cheaply; final choices are always confirmed by training a generation model on top of the frozen tokenizer.

What would settle it

Train the same downstream generator at a very different scale (for example, an 8B DiT) on KVAE-4x16x16 versus HunyuanVideo-1.5, or on 128-channel versus 64-channel audio latents, under the same data and steps: the paper's own analysis predicts these rankings can shift with generator capacity, so a reversal at another scale would show the claimed superiority is conditional on the generator rather than intrinsic to the tokenizers.

Watch

Extended reading notes

Core claim

At its core the paper claims that a tokenizer's value for latent diffusion lies less in reconstruction fidelity than in 'diffusability'—how readily a diffusion model can learn to denoise its latent space—and that this property can be engineered and screened. Under a controlled protocol, a fixed downstream generator (Kandinsky-5's 2B DiT for image and video, a 0.6B DiT for audio) is trained on each tokenizer with data, captions, and steps held constant, so that quality differences are attributed to the latent space alone. With that protocol, KVAE-4x16x16 surpasses HunyuanVideo-1.5's tokenizer in text-to-video generation and is preferred in side-by-side evaluation; KVAE-Audio with 64 channels at 50 Hz is preferred over all three audio baselines on all three judged criteria; and KVAE-2D-2.0 leads its FLUX baselines on semantic quality at equal steps. The paper introduces the correlation decay slope (CDS), a spatial cosine-similarity decay measure, as a cheap screening statistic that correlated strongly (r=0.906) with subjective visual quality across 14 image-tokenizer configurations.

Load-bearing premise

The ranking rests on the assumption that the fixed in-house generator used for the tokenizer swap—Kandinsky-5's 2B DiT for image and video, a 0.6B DiT for audio—is neutral across different latent geometries (patch size, channel count, frame rate), so that every measured difference is attributable to the tokenizer rather than to interactions with the pipeline.

Editorial extensions

If this is right

  • KVAE's released checkpoints can be dropped into existing latent-diffusion pipelines as open-source components, with the reported results suggesting they would not degrade and often improve generation quality.
  • Higher compression ratios, such as 4x16x16 with 64 channels, bring faster convergence in the tested 2B image and video generator, which translates to shorter training runs and lower compute cost.
  • A full-band 48 kHz audio tokenizer with a 50 Hz latent removes the need for a vocoder and makes joint text-to-video-and-audio generation possible from a single continuous latent space.
  • The optimal number of latent channels is not intrinsic: the paper finds 64 channels best for audio at a 0.6B generator scale while 64 channels improves video, showing the choice must be made jointly with compression factor and downstream model size.
  • CDS or similar latent diagnostics can screen tokenizer candidates before the costly step of training a full diffusion generator, as long as the correlation is re-checked on new configurations.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • At larger generator scales, the reported rankings may shift: the paper itself ties the 64-channel audio optimum to the 0.6B generator, so a wider audio latent could win with a bigger DiT.
  • The headline 'matches or surpasses' is established with one in-house generator family; on different diffusion backbones the ranking may compress or reverse, so the claim is most credible for pipelines close to Kandinsky-5's.
  • The CDS screening result is an in-sample correlation (n=14) and, as the paper notes, does not establish out-of-sample prediction; pre-registering CDS thresholds and testing them on new configurations would settle its value.
  • Two stated gaps limit the audio result's reproducibility: the frozen alignment model and loss are deferred to another publication, and the FAD backbones run at 16–32 kHz, so the objective metrics cannot confirm the advertised full-band generation.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The manuscript introduces KVAE, a family of continuous-latent tokenizers for image, video, and audio, designed for text-conditioned latent diffusion. For video, it presents two causal 3D tokenizers (4x8x8/16ch and 4x16x16/64ch); for image, a 2D tokenizer (8x8/32ch); for audio, a 48 kHz full-band waveform tokenizer with a 50 Hz/64-channel latent. The paper reports reconstruction metrics (PSNR, LPIPS, SSIM, PESQ, spectral distances) and generation evaluations (FID, CLIP, CLAP, FAD, side-by-side preference) against open tokenizers from Wan, HunyuanVideo, FLUX, MovieGen, StableAudio/SAME, and MMAudio, claiming these results 'match or surpass' the baselines. It also proposes a correlation-decay-slope (CDS) diffusability diagnostic and describes ablations on channel count, normalization, decoder width, attention, and perceptual losses. The final section concludes that the released tokenizers are competitive drop-in components for latent diffusion systems.

Significance. If the findings hold, the KVAE tokenizers are a useful open contribution: they are released with code and weights, the comparisons are carried out under a fixed downstream generator with external baselines, and the paper is unusually candid about caveats (e.g., the 0.6B-scale dependence of the audio channel optimum in Sec. 6.7, and the structural insensitivity of FAD backbones to the band above 16 kHz in Sec. 6.6). The controlled tokenizer-swap protocol is a valuable methodological template, and the explicit distinction between reconstruction fidelity, diffusability, and generation quality is a strength. However, the broad competitive claim is only partially supported by the presented evidence: the tokenizer swaps are not neutral to latent frame rate, channel count, and generator scale, and the central audio generation comparison lacks uncertainty quantification and specification of the alignment regularizer. The manuscript is therefore a solid technical report whose central claim needs either additional experiments or careful reframing.

major comments (4)
  1. [Sec. 6.6 and Sec. 6.7] The audio tokenizer-swap comparison of Table 4 does not isolate the latent space as claimed. The fixed 0.6B DiT is trained on MMAudio (40 channels, 43.07 Hz), KVAE-Audio (64 channels, 50 Hz), DAC-VAE MovieGen (128 channels, 25 Hz), and SAME-L (256 channels, about 10.8 Hz), so per second of audio the transformer processes approximately 43, 50, 25, or 11 tokens, respectively. The paper does not state how batch size or sequence length were equalized (e.g., seconds-per-batch versus tokens-per-batch), and attention cost and effective model capacity differ by up to a factor of about 4.6x across condition. Section 6.7 itself acknowledges that the optimal channel count is a property of the generator scale, so the reported ranking may be a property of the 0.6B pipeline rather than of the tokenizer itself. Please either re-run the comparison with equalized token budgets or at multiple generator scales, or restrict the claim to 'competitive under the Kandinsky-5 0.6B pipeline'.
  2. [Sec. 6.3 and Sec. 6.4] The audio model's training objective includes an alignment regularizer Lalign against a frozen audio foundation model F, but the identity of F and the form of Lalign are explicitly deferred to 'the dedicated publication' (Sec. 6.3 and Sec. 6.4). Since this term is part of the objective used to train the released checkpoints, the method is not reproducible from the manuscript and the contribution of the alignment term to the reported generation quality cannot be independently assessed. Please provide at least a precise specification of F and Lalign (or a complete pseudo-code description) in an appendix, or remove the alignment term from the reported model and retrain/re-evaluate without it.
  3. [Tables 3 and 4, Sec. 4.2-4.4] The headline comparative claims are based on point estimates without error bars, confidence intervals, or significance tests. For example, in Table 4 (Song Describer), the MMAudio baseline has a higher CLAP score (0.356 vs 0.339) and a lower FAD-PANNs (5.412 vs 7.971), and in Table 4 (LibriSpeech) DAC-VAE MovieGen has a higher CLAP (0.413 vs 0.389); in Table 3(c) EARS, the text calls a 0.31 dB SI-SDR deficit 'within noise' without defining the noise level. Similarly, the side-by-side evaluations in Sec. 4.2-4.4 and Sec. 6.6 report win rates and 'preferred over all three baselines on all three criteria' without stating the number of annotators, the number of prompts, the inter-annotator agreement, or the statistical test used. These omissions make it impossible to distinguish genuine differences from noise, especially for the subjective claims that the text says 'carry the main weight of the comparison.' Please add uncertainty quantification and the experimental protocol details, or temper the claims accordingly.
  4. [Sec. 4.2 and Sec. 4.3] The same neutrality concern applies to the visual tokenizer swaps. Section 4.2 states that patch size is adjusted to balance the number of spatial tokens, but the compared tokenizers still differ in channel count (e.g., 16 vs 64 for the 4x8x8 vs 4x16x16 comparison) and in temporal compression, and all image/video generation comparisons use a single 2B generator. The paper's own Sec. 5.2 ablation shows that channel count interacts with convergence and final quality, so the observed ranking may be specific to the Kandinsky-5 2B pipeline at the evaluated resolutions. To support the abstract's broad 'matches or surpasses frontier opensource tokenizers' claim, the visual comparisons should either include at least one additional generator scale or be explicitly framed as pipeline-specific evidence.
minor comments (4)
  1. [Sec. 4.4, Table 2] The OmniDoc-TokenBench table lists FID values (KVAE 1.74 vs FLUX.1-dev 0.554 and FLUX.2-dev 0.73) but the text only claims superiority on PSNR, SSIM, and NED; please define what FID is measuring in this reconstruction context and explain why the KVAE value is worse if the comparison is meant to support the 'surpassing' claim.
  2. [Sec. 5.1] The CDS analysis, while honestly caveated, is presented as 'crucial' for model selection despite being based on a single in-sample cross-sectional correlation (r=0.906 over 14 configurations) and one joint-training trajectory; please add an explicit out-of-sample test or clearly label CDS as a heuristic rather than a validated selection criterion.
  3. [Throughout] There are several typos and formatting inconsistencies: 'KV AE' is sometimes written 'KV AE' and sometimes 'KV AE-...' in text, 'charachteristic' appears in Sec. 3.1, 'HunyaunVideo' in Sec. 4.1, and the reference list contains a placeholder '[89] VERIFY author list' that must be resolved before publication.
  4. [Sec. 6.4] The crop-length schedule is presented as a design choice but the specific schedule (e.g., how many steps at 0.38 s, how the length increases to 5 s, and how the batch size is adjusted) is not reported; since the paper emphasizes sharing training details, please provide the exact schedule or a reference to the repository where it is defined.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the headline ranking rests on controlled tokenizer-swap experiments against external baselines rather than on fitted constants or self-citations.

full rationale

The derivation chain is self-contained: KVAE tokenizers are trained with reconstruction, adversarial, KL and alignment objectives (Secs. 3.3 and 6.4), and the central claims are evaluated by (a) reconstruction metrics on public benchmarks (MCL-JCV, OmniDoc-TokenBench, AudioSet, MUSDB18-HQ, EARS) and (b) generation via a controlled tokenizer swap in which one downstream generator is retrained per tokenizer with the same data, captions, architecture, and step counts (Secs. 4.2, 4.3, and 6.6). No parameter of KVAE is fitted to these benchmark scores and then reported as a prediction; the reported numbers are direct measurements. The CDS statistic used for candidate screening is explicitly flagged as an in-sample association that "does not establish out-of-sample prediction" (Sec. 5.1), so the paper does not present the CDS correlation as independent validation. The downstream generator (Kandinsky-5 [4]) and evaluation utilities ([49]) are self-citations, but they are evaluation infrastructure rather than the target result; the tokenizer comparisons are against third-party baselines with released weights (Wan, HunyuanVideo, FLUX, MMAudio, DAC-VAE MovieGen, SAME-L) under the same protocol. Potential concerns that the fixed-generator comparison is sensitive to latent token counts, channel counts, and generator scale (for example, the audio swap sees roughly 43, 50, 25, or 11 tokens per second across baselines, and Sec. 6.7 shows the optimal audio channel count depends on generator scale) are validity and fairness risks for the drop-in claim, not equation-level circularity. Similarly, selecting 64 audio channels via the Sec. 6.6 protocol before reporting Table 4 is a selection-on-evaluation risk, not a fitted-input-called-prediction reduction. No self-definitional reduction, imported uniqueness, ansatz-by-citation, or renaming pattern is present.

Assumptions & free parameters 5 free parameters · 5 assumptions · 0 invented entities

The central claims are empirical, so the most honest accounting is the hand-chosen hyperparameters plus the domain assumptions behind the evaluation protocol. The audio model carries the largest undisclosed burden: the alignment regularizer and perceptual details are withheld, and the optimal channel count is stated to depend on generator scale.

free parameters (5)
  • Latent channel counts = 16 and 64 (video), 32 (image), 64 (audio)
    Selected by ablation rather than derived. Sec. 5.2 and 6.7 show these values trade off reconstruction against generation; the audio 64-channel optimum is stated to depend on the 0.6B generator scale.
  • Audio objective weights w_KL, w_adv, w_perc, w_align, w_spec = Not reported numerically
    Chosen empirically; Sec. 6.4 says the useful range for w_perc and w_align is narrow and must be found jointly, but exact values are withheld.
  • Audio stride ladder [2,3,4,5,8] = Product 960, giving 50 Hz at 48 kHz
    Hand-designed factorization to place the latent at 50 Hz; alternative ladders are not compared in the main text.
  • CDS linear-fit window = Manhattan distances delta = 1,...,8
    The statistic g(delta) = alpha + beta*delta uses the iREPA default range; CDS values and model selection depend on this window choice.
  • Audio crop-length schedule = 0.38 s to 5 s across training stages
    Hand-tuned curriculum justified by slow envelopes and segment lengths; it affects reconstruction and generation quality.
assumptions (5)
  • domain assumption Standard VAE, KL, GAN and flow-matching training objectives produce latents that are good substrates for diffusion.
    Invoked throughout Sec. 3-6; no proof that optimization under these losses transfers to downstream generation quality, which is why the paper measures generation separately.
  • domain assumption The fixed downstream generator is a neutral instrument for comparing tokenizers.
    Sec. 4.2 and 6.6 assume that with identical data, captions, and hyperparameters, differences in generation quality are attributable to the latent space. If the in-house Kandinsky-5 pipeline is not neutral, rankings could shift.
  • domain assumption CDS is a valid proxy for subjective generation quality.
    Sec. 5.1 reports r=0.906 in sample across 14 configurations including the final KVAE-2D-2.0, and explicitly says it does not establish out-of-sample prediction.
  • domain assumption The bandwidth filter correctly identifies files that are truly full-band 48 kHz.
    Sec. 6.4 removes a substantial fraction of upsampled audio based on estimated true bandwidth; if the estimate is wrong, the full-band reconstruction claims are affected.
  • standard math Adam converges reliably for all training stages.
    All stages use Adam with fixed learning rates; this is standard practice, not proven in the paper.

how reviews work

0 comments
Cite this review

Pith. "Pith review of KVAE: Family of Tokenizers for Multimodal Generative Models." pith.science (2026). https://pith.science/paper/7RKTK473

@misc{pith2026260805798,
  author       = {Pith},
  title        = {Pith review of: KVAE: Family of Tokenizers for Multimodal Generative Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7RKTK473}},
  note         = {Machine review of arXiv:2608.05798}
}
read the original abstract

Latent diffusion modeling (LDM), a prominent paradigm, utilizes tokenizers to map input signal to compressed representation. This dependency positions tokenizer as an integral part of generation process itself, since it affects learning speed, quality of synthesized samples and lay foundation for later applications. This report presents series of KVAE tokenizers for audio, image and video, all designed for subsequent text-conditioned generation: KVAE-Audio, a continuous full-band 48 kHz tokenizer with a 50 Hz latent of 64 channels; KVAE-3D -- two causal video tokenizers for 4x16x16 and 4x8x8 compression; KVAE-2D, an image model, compressing input by factor of 8 with 32 channels. We demonstrate that reconstruction (PSNR, LPIPS, PESQ, etc.) and generation results on objective (Frechet Distance, CLIP score, CLAP score, etc.) and subjective (side-by-side evaluation) metrics matches or surpasses frontier opensource tokenizers, such as VAEs from Wan-2.2, HunyuanVideo-1.5, FLUX.2, MovieGen, StableAudio and MMAudio. Considering difficulty of development, we share with community training details, model selection method and ablation on design choices. The code is publicly available at https://github.com/kandinskylab/kvae and https://github.com/kandinskylab/kvae-audio.

Figures

Figures reproduced from arXiv: 2608.05798 by the authors.

Figure 1
Figure 1. Frames from internal video dataset used for intermediate evaluations [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Learning curves of text-to-image generation models on 384x256 resolution [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Results of side-by-side evaluation of image generations for KVAE-4x16x16 [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (10 more)
Figure 4
Figure 4. Figure 4: Results of side-by-side evaluation of image generations for KVAE-4x8x8 [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Side-by-side evaluation of 121-frame 768x512 video generations for KVAE-4x16x16 and [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: Learning curves of text-to-video generation models on 121-frame 384x256 resolution. With [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 7
Figure 7. Figure 7: Side-by-side soft win rates for the two downstream 2B T2I DiT stacks at 200K optimization [PITH_FULL_IMAGE:figures/full_fig_p009_7.png]
Figure 8
Figure 8. Figure 8: CDS analysis for image tokenizers: (a) Bradley–Terry visual quality vs CDS ( [PITH_FULL_IMAGE:figures/full_fig_p010_8.png]
Figure 9
Figure 9. Figure 9: Normalization layer comparison on KVAE-4x8x8. [PITH_FULL_IMAGE:figures/full_fig_p011_9.png]
Figure 10
Figure 10. Figure 10: Ablation study for the 4x16x16 design. Left/Middle images: reconstruction metrics Right image: learning curves of text-to-video generation models 6 Audio tokenization Key principles of visual tokenization outlined in previous section holds for audio as well: latent di…
Figure 11
Figure 11. Figure 11: Side-by-side evaluation of audio generations for KVAE-Audio vs MMAudio 44.1kHz [PITH_FULL_IMAGE:figures/full_fig_p017_11.png]
Figure 12
Figure 12. Figure 12: Side-by-side evaluation of audio generations for KVAE-Audio vs SAME-L [PITH_FULL_IMAGE:figures/full_fig_p017_12.png]
Figure 13
Figure 13. Figure 13: Side-by-side evaluation of audio generations for KVAE-Audio vs DACVAE MovieGen [PITH_FULL_IMAGE:figures/full_fig_p017_13.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

122 extracted references · 66 canonical work pages

  1. [1]

    Bitdance: Scaling autoregressive generative models with binary tokens, 2026

    Yuang Ai, Jiaming Han, Shaobin Zhuang, Weijia Mao, Xuefeng Hu, Ziyan Yang, Zhen- heng Yang, Yali Wang, Huaibo Huang, Xiangyu Yue, and Hao Chen. Bitdance: Scaling autoregressive generative models with binary tokens, 2026. 1

  2. [2]

    OmniDoc-TokenBench

    Alibaba Group. OmniDoc-TokenBench. https://github.com/alibaba/ OmniDoc-TokenBench, 2026. Official benchmark repository. 8

  3. [3]

    AOM Common Test Conditions v5.0

    Alliance for Open Media. AOM Common Test Conditions v5.0. Input Document CWG-D103o, Alliance for Open Media (AOMedia), 8 2023. Codec Working Group. 6

  4. [4]

    Kandinsky 5.0: A family of foundation models for image and video generation,

    Vladimir Arkhipkin, Vladimir Korviakov, Nikolai Gerasimenko, Denis Parkhomenko, Vi- acheslav Vasilev, Alexey Letunovskiy, Nikolai Vaulin, Maria Kovaleva, Ivan Kirillov, Lev Novitskiy, Denis Koposov, Nikita Kiselev, Alexander Varlamov, Dmitrii Mikhailov, Vladimir Polovnikov, Andrey Shutkin, Julia Agafonova, Ilya Vasiliev, Anastasiia Kargapoltseva, Anna Dmi...

  5. [5]

    Qwen2.5-vl technical report, 2025

    Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. Qwen2.5-vl technical report, 2025. 6

  6. [7]

    Stable video diffusion: Scaling latent video diffusion models to large datasets,

    Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram V oleti, Adam Letts, Varun Jampani, and Robin Rombach. Stable video diffusion: Scaling latent video diffusion models to large datasets,

  7. [8]

    Align your latents: High-resolution video synthesis with latent diffusion models, 2023

    Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your latents: High-resolution video synthesis with latent diffusion models, 2023. 3

  8. [9]

    Bruinsma, Ana Lucic, Megan Stanley, Anna Vaughan, Johannes Brandstetter, Patrick Garvan, Maik Riechert, Jonathan A

    Cristian Bodnar, Wessel P. Bruinsma, Ana Lucic, Megan Stanley, Anna Vaughan, Johannes Brandstetter, Patrick Garvan, Maik Riechert, Jonathan A. Weyn, Haiyu Dong, Jayesh K. Gupta, Kit Thambiratnam, Alexander T. Archibald, Chun-Chieh Wu, Elizabeth Heider, Max Welling, Richard E. Turner, and Paris Perdikaris. A foundation model for the earth system, 2024. 1

Show all 122 references
  1. [10]

    Bradley and Milton E

    Ralph A. Bradley and Milton E. Terry. Rank analysis of incomplete block designs: I. the method of paired comparisons.Biometrika, 39(3/4):324–345, 1952. https://doi.org/10. 2307/2334029. 9

  2. [11]

    Efros, and Tero Karras

    Tim Brooks, Janne Hellsten, Miika Aittala, Ting-Chun Wang, Timo Aila, Jaakko Lehtinen, Ming-Yu Liu, Alexei A. Efros, and Tero Karras. Generating long videos of dynamic scenes,

  3. [12]

    Video generation models as world simulators

    Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, Clarence Ng, Ricky Wang, and Aditya Ramesh. Video generation models as world simulators. 2024. 3

  4. [13]

    Deep compression autoencoder for efficient high-resolution diffusion models, 2025

    Junyu Chen, Han Cai, Junsong Chen, Enze Xie, Shang Yang, Haotian Tang, Muyang Li, Yao Lu, and Song Han. Deep compression autoencoder for efficient high-resolution diffusion models, 2025. 3, 4

  5. [14]

    Dc-videogen: Efficient video generation with deep compression video autoencoder, 2025

    Junyu Chen, Wenkun He, Yuchao Gu, Yuyang Zhao, Jincheng Yu, Junsong Chen, Dongyun Zou, Yujun Lin, Zhekai Zhang, Muyang Li, Haocheng Xi, Ligeng Zhu, Enze Xie, Song Han, and Han Cai. Dc-videogen: Efficient video generation with deep compression video autoencoder, 2025. 2

  6. [15]

    Dc-ae 1.5: Accelerating diffusion model convergence with structured latent space, 2025

    Junyu Chen, Dongyun Zou, Wenkun He, Junsong Chen, Enze Xie, Song Han, and Han Cai. Dc-ae 1.5: Accelerating diffusion model convergence with structured latent space, 2025. 3

  7. [16]

    Taming multimodal joint training for high-quality video-to-audio synthesis.arXiv preprint arXiv:2412.15322, 2024

    Ho Kei Cheng, Masato Ishii, Akio Hayakawa, Takashi Shibuya, Alexander Schwing, and Yuki Mitsufuji. Taming multimodal joint training for high-quality video-to-audio synthesis.arXiv preprint arXiv:2412.15322, 2024. 12, 15

  8. [17]

    Chien, Liuzixuan Lin, Hai Nguyen, Varsha Rao, Tristan Sharma, and Rajini Wijayawardana

    Andrew A. Chien, Liuzixuan Lin, Hai Nguyen, Varsha Rao, Tristan Sharma, and Rajini Wijayawardana. Reducing the carbon impact of generative ai inference (today and in 2035). In Proceedings of the 2nd Workshop on Sustainable Computer Systems (HotCarbon ’23), Boston, MA, USA, Jul...

  9. [18]

    Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V

    Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinso...

  10. [19]

    Adversarial video generation on complex datasets, 2019

    Aidan Clark, Jeff Donahue, and Karen Simonyan. Adversarial video generation on complex datasets, 2019. 2

  11. [20]

    High fidelity neural audio compression.arXiv preprint arXiv:2210.13438, 2022

    Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi. High fidelity neural audio compression.arXiv preprint arXiv:2210.13438, 2022. 12

  12. [21]

    Irc-gan: Introspective recurrent convolutional gan for text-to-video generation

    Kangle Deng, Tianyi Fei, Xin Huang, and Yuxin Peng. Irc-gan: Introspective recurrent convolutional gan for text-to-video generation. InProceedings of the 28th International Joint Conference on Artificial Intelligence (IJCAI), pages 2216–2222, 2019. 2

  13. [22]

    Taming transformers for high-resolution image synthesis, 2021

    Patrick Esser, Robin Rombach, and Björn Ommer. Taming transformers for high-resolution image synthesis, 2021. 5

  14. [23]

    Parker, C

    Zach Evans, Julian D. Parker, C. J. Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons. Stable audio open.arXiv preprint arXiv:2407.14358, 2024. 12

  15. [24]

    Parker, Matthew Rice, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons

    Zach Evans, Julian D. Parker, Matthew Rice, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons. Stable audio 3, 2026. 12, 15

  16. [25]

    The prism hypothesis: Harmonizing semantic and pixel representations via unified autoencoding, 2026

    Weichen Fan, Haiwen Diao, Quan Wang, Dahua Lin, and Ziwei Liu. The prism hypothesis: Harmonizing semantic and pixel representations via unified autoencoding, 2026. 3

  17. [26]

    Gemmeke, Daniel P

    Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Chan- ning Moore, Manoj Plakal, and Marvin Ritter. Audio Set: An ontology and human-labeled dataset for audio events. InIEEE International Conference on Acoustics, Speech and Signal Processing ...

  18. [27]

    BigVGAN: A universal neural vocoder with large-scale training

    Sang gil Lee, Wei Ping, Boris Ginsburg, Bryan Catanzaro, and Sungroh Yoon. BigVGAN: A universal neural vocoder with large-scale training. InThe Eleventh International Conference on Learning Representations (ICLR), 2023. 12 20

  19. [28]

    Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio

    Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial networks, 2014. 2

  20. [29]

    Veo 3.1: Our leading video generation model

    Google DeepMind. Veo 3.1: Our leading video generation model. Google DeepMind Official Product Page, October 2025. Accessed: 2026-07-13. 3

  21. [30]

    Ltx-2: Efficient joint audio-visual foundation model, 2026

    Yoav HaCohen, Benny Brazowski, Nisan Chiprut, Yaki Bitterman, Andrew Kvochko, Avishai Berkowitz, Daniel Shalem, Daphna Lifschitz, Dudu Moshe, Eitan Porat, Eitan Richardson, Guy Shiran, Itay Chachy, Jonathan Chetboun, Michael Finkelson, Michael Kupchick, Nir Zabari, Nitzan Guet...

  22. [31]

    Shawn Hershey, Sourish Chaudhuri, Daniel P. W. Ellis, Jort F. Gemmeke, Aren Jansen, R. Channing Moore, Manoj Plakal, Devin Platt, Rif A. Saurous, Bryan Seybold, Malcolm Slaney, Ron J. Weiss, and Kevin Wilson. CNN architectures for large-scale audio classification. InIEEE Inter...

  23. [32]

    Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 30, 2017

    Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 30, 2017. 6

  24. [33]

    Kingma, Ben Poole, Mohammad Norouzi, David J

    Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P. Kingma, Ben Poole, Mohammad Norouzi, David J. Fleet, and Tim Salimans. Imagen video: High definition video generation with diffusion models, 2022. 3

  25. [34]

    Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J. Fleet. Video diffusion models, 2022. 3

  26. [35]

    Cogvideo: Large-scale pretraining for text-to-video generation via transformers, 2022

    Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang. Cogvideo: Large-scale pretraining for text-to-video generation via transformers, 2022. 2

  27. [36]

    Tangoflux: Super fast and faithful text to audio generation with flow matching and clap-ranked preference optimization,

    Chia-Yu Hung, Navonil Majumder, Zhifeng Kong, Ambuj Mehrish, Amir Ali Bagherzadeh, Chuan Li, Rafael Valle, Bryan Catanzaro, and Soujanya Poria. Tangoflux: Super fast and faithful text to audio generation with flow matching and clap-ranked preference optimization,

  28. [37]

    Perceptual evaluation of speech quality (PESQ).International Telecommunication Union, 2001

    ITU-T Recommendation P.862. Perceptual evaluation of speech quality (PESQ).International Telecommunication Union, 2001. 15

  29. [38]

    Video pixel networks, 2016

    Nal Kalchbrenner, Aaron van den Oord, Karen Simonyan, Ivo Danihelka, Oriol Vinyals, Alex Graves, and Koray Kavukcuoglu. Video pixel networks, 2016. 2

  30. [39]

    Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms.Interspeech,

    Kevin Kilgour, Mauricio Zuluaga, Dominik Roblek, and Matthew Sharifi. Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms.Interspeech,

  31. [40]

    AudioCaps: Generating captions for audios in the wild

    Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim. AudioCaps: Generating captions for audios in the wild. InProceedings of NAACL-HLT, 2019. 16

  32. [41]

    Kingma and Jimmy Ba

    Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In International Conference on Learning Representations (ICLR), 2015. 5, 14

  33. [42]

    Auto-encoding variational bayes, 2022

    Diederik P Kingma and Max Welling. Auto-encoding variational bayes, 2022. 3

  34. [43]

    Klingai enters the 3.0 era: All in one, one for all! kling 3.0 model now fully rolled out

    Kling AI. Klingai enters the 3.0 era: All in one, one for all! kling 3.0 model now fully rolled out. Kling AI Official Release Notes, February 2026. Accessed: 2026-07-13. 3

  35. [44]

    Carbon Emissions in the Tailpipe of Gen- erative AI.Harvard Data Science Review, 15(Special Issue 5), aug 20 2024

    Tamara Kneese and Meg Young. Carbon Emissions in the Tailpipe of Gen- erative AI.Harvard Data Science Review, 15(Special Issue 5), aug 20 2024. https://hdsr.mitpress.mit.edu/pub/fscsqwx4. 2 21

  36. [45]

    Ross, Bryan Seybold, and Lu Jiang

    Dan Kondratyuk, Lijun Yu, Xiuye Gu, José Lezama, Jonathan Huang, Grant Schindler, Rachel Hornung, Vighnesh Birodkar, Jimmy Yan, Ming-Chang Chiu, Krishna Somandepalli, Hassan Akbari, Yair Alon, Yong Cheng, Josh Dillon, Agrim Gupta, Meera Hahn, Anja Hauth, David Hendon, Alonso M...

  37. [46]

    HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis

    Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae. HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis. InAdvances in Neural Information Processing Systems, 2020. 14

  38. [47]

    Plumb- ley

    Qiuqiang Kong, Yin Cao, Turab Iqbal, Yuxuan Wang, Wenwu Wang, and Mark D. Plumb- ley. PANNs: Large-scale pretrained audio neural networks for audio pattern recognition. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 28:2880–2894, 2020. 16

  39. [48]

    Hunyuanvideo: A systematic framework for large video generative models,

    Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, Kathrina Wu, Qin Lin, Junkun Yuan, Yanxin Long, Aladdin Wang, Andong Wang, Changlin Li, Duojun Huang, Fang Yang, Hao Tan, Hongmei Wang, Jacob Song, Jiawang Bai, J...

  40. [49]

    Kandinsky video tools

    Denis Koposov, Anna Dmitrienko, Ivan Kirillov, Kirill Chernyshev, Denis Parkhomenko, and Vladimir Korviakov. Kandinsky video tools. https://github.com/gen-ai-team/ kandinsky-video-tools, 2025. 7

  41. [50]

    Efficient training of audio transformers with patchout

    Khaled Koutini, Jan Schlüter, Hamid Eghbal-zadeh, and Gerhard Widmer. Efficient training of audio transformers with patchout. InInterspeech, 2022. 16

  42. [51]

    Eq-vae: Equivariance regularized latent space for improved generative image modeling, 2025

    Theodoros Kouzelis, Ioannis Kakogeorgiou, Spyros Gidaris, and Nikos Komodakis. Eq-vae: Equivariance regularized latent space for improved generative image modeling, 2025. 5

  43. [52]

    High-fidelity audio compression with improved RVQGAN

    Rithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar, and Kundan Kumar. High-fidelity audio compression with improved RVQGAN. InAdvances in Neural Information Processing Systems, 2023. 12, 13, 14

  44. [53]

    REPA-E: Unlocking vae for end-to-end tuning of latent diffusion transformers

    Xingjian Leng, Jaskirat Singh, Yunzhong Hou, Zhenchang Xing, Saining Xie, and Liang Zheng. REPA-E: Unlocking vae for end-to-end tuning of latent diffusion transformers. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 18262–18272, 2025. https://...

  45. [54]

    Diffusionbench: On holistic evaluation of diffusion transformers with a unified training framework bridging imagenet and text-to-image, 2026

    Xingjian Leng, Jaskirat Singh, Zhanhao Liang, Ethan Smith, Martin Bell, Aninda Saha, Yuhui Yuan, and Liang Zheng. Diffusionbench: On holistic evaluation of diffusion transformers with a unified training framework bridging imagenet and text-to-image, 2026. https://arxiv. org/ab...

  46. [55]

    Wf-vae: Enhancing video vae by wavelet-driven energy flow for latent video diffusion model,

    Zongjian Li, Bin Lin, Yang Ye, Liuhan Chen, Xinhua Cheng, Shenghai Yuan, and Li Yuan. Wf-vae: Enhancing video vae by wavelet-driven energy flow for latent video diffusion model,

  47. [56]

    Generating novel, designable, and diverse protein structures by equivariantly diffusing oriented residue clouds, 2023

    Yeqing Lin and Mohammed AlQuraishi. Generating novel, designable, and diverse protein structures by equivariantly diffusing oriented residue clouds, 2023. 1

  48. [57]

    Plumbley

    Haohe Liu, Yi Yuan, Xubo Liu, Xinhao Mei, Qiuqiang Kong, Qiao Tian, Yuping Wang, Wenwu Wang, Yuxuan Wang, and Mark D. Plumbley. AudioLDM 2: Learning holistic audio generation with self-supervised pretraining.arXiv preprint arXiv:2308.05734, 2023. 12 22

  49. [58]

    Delving into latent spectral biasing of video vaes for superior diffusability.arXiv preprint arXiv:2512.05394, 2025.https://arxiv.org/abs/2512.05394

    Shizhan Liu, Xinran Deng, Zhuoyi Yang, Jiayan Teng, Xiaotao Gu, and Jie Tang. Delving into latent spectral biasing of video vaes for superior diffusability.arXiv preprint arXiv:2512.05394, 2025.https://arxiv.org/abs/2512.05394. 3, 9

  50. [59]

    Di Ma, Fan Zhang, and David R. Bull. Bvi-dvc: A training database for deep video compres- sion.IEEE Transactions on Multimedia, 24:3847–3858, 2022. 6

  51. [60]

    The song describer dataset: A corpus of audio captions for music-and- language evaluation

    Ilaria Manco, Benno Weck, Seungheon Doh, Minz Won, Yixiao Zhang, Dmitry Bogdanov, Yusong Wu, Ke Chen, Philip Tovstogan, Emmanouil Benetos, Elio Quinton, György Fazekas, and Juhan Nam. The song describer dataset: A corpus of audio captions for music-and- language evaluation. In...

  52. [61]

    Balasubramanian

    Gaurav Mittal, Tanya Marwah, and Vineeth N. Balasubramanian. Sync-draw: Automatic video generation using deep recurrent attentive architectures. InProceedings of the 25th ACM international conference on Multimedia, MM ’17, page 1096–1104. ACM, October 2017. 2

  53. [62]

    Transition matching distillation for fast video generation, 2026

    Weili Nie, Julius Berner, Nanye Ma, Chao Liu, Saining Xie, and Arash Vahdat. Transition matching distillation for fast video generation, 2026. 2

  54. [63]

    Blaschko, Albert Ali Salah, and Itir Onal Ertugrul

    Mang Ning, Mingxiao Li, Le Zhang, Lanmiao Liu, Matthew B. Blaschko, Albert Ali Salah, and Itir Onal Ertugrul. Spectrum matching: a unified perspective for superior diffusability in latent diffusion.arXiv preprint arXiv:2603.14645, 2026. https://arxiv.org/abs/2603.14645. 3, 9

  55. [64]

    Alpamayo-r1: Bridging reasoning and action prediction for generalizable autonomous driving in the long tail, 2026

    NVIDIA, :, Yan Wang, Wenjie Luo, Junjie Bai, Yulong Cao, Tong Che, Ke Chen, Yuxiao Chen, Jenna Diamond, Yifan Ding, Wenhao Ding, Liang Feng, Greg Heinrich, Jack Huang, Peter Karkus, Boyi Li, Pinyi Li, Tsung-Yi Lin, Dongran Liu, Ming-Yu Liu, Langechuan Liu, Zhijian Liu, Jason L...

  56. [65]

    Cosmos tokenizer: A suite of image and video neural tokenizers.arXiv preprint arXiv:2501.03575, 2025

    NVIDIA, Fitsum Reda, Jinwei Gu, Xian Liu, Songwei Ge, Ting-Chun Wang, Haoxiang Wang, and Ming-Yu Liu. Cosmos tokenizer: A suite of image and video neural tokenizers.arXiv preprint arXiv:2501.03575, 2025. 3, 4, 6

  57. [66]

    To create what you tell: Generating videos from captions, 2018

    Yingwei Pan, Zhaofan Qiu, Ting Yao, Houqiang Li, and Tao Mei. To create what you tell: Generating videos from captions, 2018. 2

  58. [67]

    Librispeech: An ASR corpus based on public domain audio books

    Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. Librispeech: An ASR corpus based on public domain audio books. InIEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2015. 16

  59. [68]

    Parker, Zach Evans, CJ Carr, Zack Zukowski, Josiah Taylor, Matthew Rice, and Jordi Pons

    Julian D. Parker, Zach Evans, CJ Carr, Zack Zukowski, Josiah Taylor, Matthew Rice, and Jordi Pons. SAME: A semantically-aligned music autoencoder, 2026. 12, 15

  60. [69]

    Sampson, Shikai Li, Simone Parmeggiani, Steve Fine, Tara Fowler, Vladan Petrovic, and Yuming Du

    Adam Polyak, Amit Zohar, Andrew Brown, Andros Tjandra, Animesh Sinha, Ann Lee, Apoorv Vyas, Bowen Shi, Chih-Yao Ma, Ching-Yao Chuang, David Yan, Dhruv Choudhary, Dingkang Wang, Geet Sethi, Guan Pang, Haoyu Ma, Ishan Misra, Ji Hou, Jialiang Wang, Kiran Jagadeesh, Kunpeng Li, Lu...

  61. [70]

    Andersson, Andrew El-Kadi, Dominic Masters, Timo Ewalds, Jacklynn Stott, Shakir Mohamed, Peter Battaglia, Remi Lam, and Matthew Willson

    Ilan Price, Alvaro Sanchez-Gonzalez, Ferran Alet, Tom R. Andersson, Andrew El-Kadi, Dominic Masters, Timo Ewalds, Jacklynn Stott, Shakir Mohamed, Peter Battaglia, Remi Lam, and Matthew Willson. Gencast: Diffusion-based ensemble forecasting for medium-range weather, 2024. 1

  62. [71]

    Qwen-Audio-V AE technical report.arXiv preprint arXiv:2607.11738, 2026

    Qwen Team. Qwen-Audio-V AE technical report.arXiv preprint arXiv:2607.11738, 2026. 12

  63. [72]

    Qwen-Image-V AE-2.0 technical report.arXiv preprint arXiv:2605.13565, 2026

    Qwen Team. Qwen-Image-V AE-2.0 technical report.arXiv preprint arXiv:2605.13565, 2026. https://arxiv.org/abs/2605.13565. 8

  64. [73]

    Learning transferable visual models from natural language supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. InInternational conference on machine learning, pag...

  65. [74]

    MUSDB18-HQ — an uncompressed version of MUSDB18, 2019

    Zafar Rafii, Antoine Liutkus, Fabian-Robert Stöter, Stylianos Ioannis Mimilakis, and Rachel Bittner. MUSDB18-HQ — an uncompressed version of MUSDB18, 2019. Zenodo. 14

  66. [75]

    EARS: An anechoic fullband speech dataset benchmarked for speech enhancement and dereverberation

    Julius Richter, Yi-Chiao Wu, Steven Krenn, Simon Welker, Bunlong Lay, Shinji Watanabe, Alexander Richard, and Timo Gerkmann. EARS: An anechoic fullband speech dataset benchmarked for speech enhancement and dereverberation. InInterspeech, 2024. 15

  67. [76]

    High-resolution image synthesis with latent diffusion models, 2022

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models, 2022. 2

  68. [77]

    High-resolution image synthesis with latent diffusion models, 2022

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models, 2022. 3

  69. [78]

    Runway gen-4: Ai video generation with world consistency

    Runway Research. Runway gen-4: Ai video generation with world consistency. Runway Research Publications, March 2025. Accessed: 2026-07-13. 3

  70. [79]

    Flow to the mode: Mode-seeking diffusion autoencoders for state-of-the-art image tokenization, 2025

    Kyle Sargent, Kyle Hsu, Justin Johnson, Li Fei-Fei, and Jiajun Wu. Flow to the mode: Mode-seeking diffusion autoencoders for state-of-the-art image tokenization, 2025. 3

  71. [80]

    Make- a-video: Text-to-video generation without text-video data, 2022

    Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, Devi Parikh, Sonal Gupta, and Yaniv Taigman. Make- a-video: Text-to-video generation without text-video data, 2022. 3

  72. [81]

    What matters for representation alignment: Global information or spatial structure?arXiv preprint arXiv:2512.10794, 2025

    Jaskirat Singh, Xingjian Leng, Zongze Wu, Liang Zheng, Richard Zhang, Eli Shechtman, and Saining Xie. What matters for representation alignment: Global information or spatial structure?arXiv preprint arXiv:2512.10794, 2025. https://arxiv.org/abs/2512.10794. 9

  73. [82]

    Improving the diffusability of autoencoders, 2025

    Ivan Skorokhodov, Sharath Girish, Benran Hu, Willi Menapace, Yanyu Li, Rameen Abdal, Sergey Tulyakov, and Aliaksandr Siarohin. Improving the diffusability of autoencoders, 2025. 2, 3, 18

  74. [83]

    Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2, 2022

    Ivan Skorokhodov, Sergey Tulyakov, and Mohamed Elhoseiny. Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2, 2022. 2

  75. [84]

    Ucf101: A dataset of 101 human actions classes from videos in the wild, 2012

    Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. Ucf101: A dataset of 101 human actions classes from videos in the wild, 2012. 2

  76. [85]

    Scenediffuser++: City-scale traffic simulation via a generative world model, 2025

    Shuhan Tan, John Lambert, Hong Jeon, Sakshum Kulshrestha, Yijing Bai, Jing Luo, Dragomir Anguelov, Mingxing Tan, and Chiyu Max Jiang. Scenediffuser++: City-scale traffic simulation via a generative world model, 2025. 1

  77. [86]

    Z-image: An efficient image generation foundation model with single-stream diffusion transformer,

    Image Team, Huanqia Cai, Sihan Cao, Ruoyi Du, Peng Gao, Steven Hoi, Zhaohui Hou, Shijie Huang, Dengyang Jiang, Xin Jin, Liangchen Li, Zhen Li, Zhong-Yu Li, David Liu, Dongyang Liu, Junhan Shi, Qilong Wu, Feng Yu, Chi Zhang, Shifeng Zhang, and Shilin Zhou. Z-image: An efficient...

  78. [87]

    Lyria Team, Antoine Caillon, Brian McWilliams, Cassie Tarakajian, Ian Simon, Ilaria Manco, Jesse Engel, Noah Constant, Yunpeng Li, Timo I. Denk, Alberto Lalama, Andrea Agostinelli, Cheng-Zhi Anna Huang, Ethan Manilow, George Brower, Hakan Erdogan, Heidi Lei, Itai Rolnick, Ivan...

  79. [88]

    Nextstep-1: Toward autoregressive image generation with continuous tokens at scale, 2025

    NextStep Team, Chunrui Han, Guopeng Li, Jingwei Wu, Quan Sun, Yan Cai, Yuang Peng, Zheng Ge, Deyu Zhou, Haomiao Tang, Hongyu Zhou, Kenkun Liu, Ailin Huang, Bin Wang, Changxin Miao, Deshan Sun, En Yu, Fukun Yin, Gang Yu, Hao Nie, Haoran Lv, Hanpeng Hu, Jia Wang, Jian Zhou, Jian...

  80. [89]

    HunyuanVideo-Foley: Multimodal diffusion with representation alignment for high-fidelity foley audio generation.arXiv preprint arXiv:2508.16930, 2025

    Tencent Hunyuan. HunyuanVideo-Foley: Multimodal diffusion with representation alignment for high-fidelity foley audio generation.arXiv preprint arXiv:2508.16930, 2025. VERIFY author list. 12

  81. [90]

    Reducio! generating 1k video within 16 seconds using extremely compressed motion latents, 2025

    Rui Tian, Qi Dai, Jianmin Bao, Kai Qiu, Yifan Yang, Chong Luo, Zuxuan Wu, and Yu-Gang Jiang. Reducio! generating 1k video within 16 seconds using extremely compressed motion latents, 2025. 3

  82. [91]

    Metaxas, and Sergey Tulyakov

    Yu Tian, Jian Ren, Menglei Chai, Kyle Olszewski, Xi Peng, Dimitris N. Metaxas, and Sergey Tulyakov. A good image generator is what you need for high-resolution video synthesis, 2021. 2

  83. [92]

    Meta audiobox aesthetics: Unified automatic quality assessment for speech, music, and sound

    Andros Tjandra, Yi-Chiao Wu, Baishan Guo, John Hoffman, Brian Ellis, Apoorv Vyas, Bowen Shi, Sanyuan Chen, Matt Le, Nick Zacharov, Carleigh Wood, Ann Lee, and Wei-Ning Hsu. Meta audiobox aesthetics: Unified automatic quality assessment for speech, music, and sound. arXiv prepr...

  84. [93]

    Ssdd: Single-step diffusion decoder for efficient image tokenization, 2026

    Théophane Vallaeys, Jakob Verbeek, and Matthieu Cord. Ssdd: Single-step diffusion decoder for efficient image tokenization, 2026. 3

  85. [94]

    Conditional image generation with pixelcnn decoders, 2016

    Aaron van den Oord, Nal Kalchbrenner, Oriol Vinyals, Lasse Espeholt, Alex Graves, and Koray Kavukcuoglu. Conditional image generation with pixelcnn decoders, 2016. 2

  86. [95]

    Neural discrete representation learning, 2018

    Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. Neural discrete representation learning, 2018. 2

  87. [96]

    Gomez, Lukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need, 2023. 2

  88. [97]

    Generating videos with scene dynamics, 2016

    Carl V ondrick, Hamed Pirsiavash, and Antonio Torralba. Generating videos with scene dynamics, 2016. 2

  89. [98]

    Wan: Open and advanced large-scale video generative models, 2025

    Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, Jianyuan Zeng, Jiayu Wang, Jingfeng Zhang, Jingren Zhou, Jinkai Wang, Jixuan Chen, Kai Zhu, Kang Zhao, Keyu Yan, Lianghua Huang, Mengyang Feng, Ningyi Zhang, Pande...

  90. [99]

    Wan-2.2 anouncement

    Wan Team. Wan-2.2 anouncement. https://wan.video/blog/wan2.2, July 2025. Ac- cessed: 2026-06-01. 3, 4, 5, 11 25

  91. [100]

    Haiqiang Wang, Weihao Gan, Sudeng Hu, Joe Yuchieh Lin, Lina Jin, Longguang Song, Ping Wang, Ioannis Katsavounidis, Anne Aaron, and C.-C. Jay Kuo. Mcl-jcv: A jnd-based h.264/avc video quality assessment dataset. In2016 IEEE International Conference on Image Processing (ICIP), p...

  92. [101]

    Videomae v2: Scaling video masked autoencoders with dual masking

    Limin Wang, Bingkun Huang, Zhiyu Zhao, Zhan Tong, Yinan He, Yi Wang, Yali Wang, and Yu Qiao. Videomae v2: Scaling video masked autoencoders with dual masking. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 14549–14560,

  93. [102]

    Emu3: Next-token prediction is all you need, 2024

    Xinlong Wang, Xiaosong Zhang, Zhengxiong Luo, Quan Sun, Yufeng Cui, Jinsheng Wang, Fan Zhang, Yueze Wang, Zhen Li, Qiying Yu, Yingli Zhao, Yulong Ao, Xuebin Min, Tao Li, Boya Wu, Bo Zhao, Bowen Zhang, Liangdong Wang, Guang Liu, Zheqi He, Xi Yang, Jingjing Liu, Yonghua Lin, Tie...

  94. [103]

    Internvideo2: Scaling foundation models for multimodal video understanding

    Yi Wang, Kunchang Li, Xinhao Li, Jiashuo Yu, Yinan He, Guo Chen, Baoqi Pei, Rongkun Zheng, Zun Wang, Yansong Shi, et al. Internvideo2: Scaling foundation models for multimodal video understanding. InEuropean conference on computer vision, pages 396–416. Springer,

  95. [104]

    Hunyuanvideo 1.5 technical report, 2025

    Bing Wu, Chang Zou, Changlin Li, Duojun Huang, Fang Yang, Hao Tan, Jack Peng, Jianbing Wu, Jiangfeng Xiong, Jie Jiang, Linus, Patrol, Peizhen Zhang, Peng Chen, Penghao Zhao, Qi Tian, Songtao Liu, Weijie Kong, Weiyan Wang, Xiao He, Xin Li, Xinchi Deng, Xuefei Zhe, Yang Li, Yanx...

  96. [105]

    Q-align: Teaching lmms for visual scoring via discrete text-defined levels

    Haoning Wu, Zicheng Zhang, Weixia Zhang, Chaofeng Chen, Liang Liao, Chunyi Li, Yixuan Gao, Annan Wang, Erli Zhang, Wenxiu Sun, et al. Q-align: Teaching lmms for visual scoring via discrete text-defined levels. InInternational Conference on Machine Learning, pages 54015–54029. ...

  97. [106]

    H3ae: High compression, high speed, and high quality autoencoder for video diffusion models, 2025

    Yushu Wu, Yanyu Li, Ivan Skorokhodov, Anil Kag, Willi Menapace, Sharath Girish, Aliaksandr Siarohin, Yanzhi Wang, and Sergey Tulyakov. H3ae: High compression, high speed, and high quality autoencoder for video diffusion models, 2025. 4

  98. [107]

    Large-scale contrastive language-audio pretraining with feature fusion and keyword- to-caption augmentation

    Yusong Wu, Ke Chen, Tianyu Zhang, Yuchen Hui, Taylor Berg-Kirkpatrick, and Shlomo Dubnov. Large-scale contrastive language-audio pretraining with feature fusion and keyword- to-caption augmentation. InIEEE International Conference on Acoustics, Speech and Signal Processing (IC...

  99. [108]

    Grok imagine API

    xAI. Grok imagine API. https://x.ai/news/grok-imagine-api , January 2026. Ac- cessed: 2026-06-01. 3

  100. [109]

    Learning to generate time-lapse videos using multi-stage dynamic generative adversarial networks, 2018

    Wei Xiong, Wenhan Luo, Lin Ma, Wei Liu, and Jiebo Luo. Learning to generate time-lapse videos using multi-stage dynamic generative adversarial networks, 2018. 2

  101. [110]

    Videogpt: Video generation using vq-vae and transformers, 2021

    Wilson Yan, Yunzhi Zhang, Pieter Abbeel, and Aravind Srinivas. Videogpt: Video generation using vq-vae and transformers, 2021. 2

  102. [111]

    Qwen3 technical report, 2025

    An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jia...

  103. [112]

    Cogvideox: Text-to-video diffusion models with an expert transformer, 2025

    Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, Da Yin, Yuxuan Zhang, Weihan Wang, Yean Cheng, Bin Xu, Xiaotao Gu, Yuxiao Dong, and Jie Tang. Cogvideox: Text-to-video diffusion models with an ex...

  104. [113]

    Reconstruction vs

    Jingfeng Yao, Bin Yang, and Xinggang Wang. Reconstruction vs. generation: Taming opti- mization dilemma in latent diffusion models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025. https://arxiv.org/abs/2501.01423. 3, 10, 12, 18

  105. [114]

    Reconstruction vs

    Jingfeng Yao, Bin Yang, and Xinggang Wang. Reconstruction vs. generation: Taming opti- mization dilemma in latent diffusion models, 2025. 5

  106. [115]

    Gundavarapu, Luca Versari, Kihyuk Sohn, David Minnen, Yong Cheng, Vighnesh Birodkar, Agrim Gupta, Xiuye Gu, Alexander G

    Lijun Yu, José Lezama, Nitesh B. Gundavarapu, Luca Versari, Kihyuk Sohn, David Minnen, Yong Cheng, Vighnesh Birodkar, Agrim Gupta, Xiuye Gu, Alexander G. Hauptmann, Boqing Gong, Ming-Hsuan Yang, Irfan Essa, David A. Ross, and Lu Jiang. Language model beats diffusion – tokenize...

  107. [116]

    Representation alignment for generation: Training diffusion transformers is easier than you think.arXiv preprint arXiv:2410.06940, 2024

    Sihyun Yu, Sangkyung Kwak, Huiwon Jang, Jongheon Jeong, Jonathan Huang, Jinwoo Shin, and Saining Xie. Representation alignment for generation: Training diffusion transformers is easier than you think.arXiv preprint arXiv:2410.06940, 2024. 12

  108. [117]

    SoundStream: An end-to-end neural audio codec

    Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi. SoundStream: An end-to-end neural audio codec. InIEEE/ACM Transactions on Audio, Speech, and Language Processing, 2021. 12

  109. [118]

    Gonzalez, Jianfei Chen, and Jun Zhu

    Jintao Zhang, Kaiwen Zheng, Kai Jiang, Haoxu Wang, Ion Stoica, Joseph E. Gonzalez, Jianfei Chen, and Jun Zhu. Turbodiffusion: Accelerating video diffusion models by 100-200 times,

  110. [119]

    Efros, Eli Shechtman, and Oliver Wang

    Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shechtman, and Oliver Wang. The unrea- sonable effectiveness of deep features as a perceptual metric, 2018. 5

  111. [120]

    Maisi-v2: Accelerated 3d high-resolution medical image synthesis with rectified flow and region-specific contrastive loss, 2025

    Can Zhao, Pengfei Guo, Dong Yang, Yucheng Tang, Yufan He, Benjamin Simon, Mason Belue, Stephanie Harmon, Baris Turkbey, and Daguang Xu. Maisi-v2: Accelerated 3d high-resolution medical image synthesis with rectified flow and region-specific contrastive loss, 2025. 1

  112. [121]

    Cv-vae: A compatible video vae for latent generative video models, 2024

    Sijie Zhao, Yong Zhang, Xiaodong Cun, Shaoshu Yang, Muyao Niu, Xiaoyu Li, Wenbo Hu, and Ying Shan. Cv-vae: A compatible video vae for latent generative video models, 2024. 3

  113. [122]

    Diffusion transformers with representation autoencoders, 2025

    Boyang Zheng, Nanye Ma, Shengbang Tong, and Saining Xie. Diffusion transformers with representation autoencoders, 2025. 3, 5, 12

  114. [123]

    Diffusing in the right space: A systematic study of latent diffusability.arXiv preprint arXiv:2606.03578, 2026

    Tianxiong Zhong, Xingye Tian, Xuebo Wang, Xin Tao, and Pengfei Wan. Diffusing in the right space: A systematic study of latent diffusability.arXiv preprint arXiv:2606.03578, 2026. https://arxiv.org/abs/2606.03578. 9 27

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.