Pith. sign in

REVIEW 90 references

DiverseDiT++: Quantifying, Analyzing, and Promoting Representation Diversity in Diffusion Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2608.03082 v1 pith:6XI2IFKB submitted 2026-08-04 cs.CV

classification cs.CV
keywords representationditsdiversityacrossblocksrepresentationsdiverseditlearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Recent advances in Diffusion Transformers (DiTs) have enabled remarkable progress in visual synthesis, benefiting from their superior scalability. To facilitate DiTs' capability of capturing meaningful internal representations, recent works such as REPA incorporate external pretrained encoders for representation alignment. However, the underlying mechanisms governing representation learning within DiTs remain poorly understood in the community. To this end, this paper first presents a systematic analysis of the representation dynamics of DiTs via quantifying the diversity of block-wise representations. Specifically, we introduce a novel metric, termed the Weighted Diversity Score (WDS), to measure the representational discrepancies across different blocks. Through extensive investigations on the evolution and influence of internal representations under various settings, we reveal that representation diversity across blocks is a critical factor for effective representation learning in DiTs. More importantly, WDS exhibits a strong correlation with synthesis quality across diverse settings, model scales, and training stages (Pearson's $r=-0.869$ with $\log(\text{FID})$), suggesting its potential as an indicator to reflect model performance and a principled guide for model optimization. Based on this key finding, we propose DiverseDiT++, a novel framework that explicitly promotes diverse representation learning. Concretely, our method incorporates long residual connections to diversify input representations across blocks and a representation diversity loss to encourage blocks to learn distinct features. Extensive experiments on ImageNet $256\times256$ and $512\times512$ demonstrate that our DiverseDiT++ yields consistent performance gains and convergence acceleration when applied to different backbones with various sizes,...

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

90 extracted references · 27 canonical work pages

  1. [1]

    Rep- resentation alignment for generation: Training diffusion transformers is easier than you think,

    S. Yu, S. Kwak, H. Jang, J. Jeong, J. Huang, J. Shin, and S. Xie, “Rep- resentation alignment for generation: Training diffusion transformers is easier than you think,” inInt. Conf. Learn. Represent., 2024. 1, 2, 3, 4, 5, 8, 9, 16, 18

  2. [2]

    Diffuse and disperse: Image generation with representation regularization,

    R. Wang and K. He, “Diffuse and disperse: Image generation with representation regularization,”arXiv preprint arXiv:2506.09027, 2025. 1, 2, 3, 8, 9, 10, 11, 16, 19, 22

  3. [3]

    Deep unsupervised learning using nonequilibrium thermodynamics,

    J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli, “Deep unsupervised learning using nonequilibrium thermodynamics,” inInt. Conf. Mach. Learn.PMLR, 2015, pp. 2256–2265. 1, 3

  4. [4]

    Denoising diffusion probabilistic models,

    J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” Adv. Neural Inform. Process. Syst., vol. 33, pp. 6840–6851, 2020. 1, 3, 4

  5. [5]

    Scalable diffusion models with transformers,

    W. Peebles and S. Xie, “Scalable diffusion models with transformers,” inInt. Conf. Comput. Vis., 2023, pp. 4195–4205. 1, 3, 8, 9, 18

  6. [6]

    All are worth words: A vit backbone for diffusion models,

    F. Bao, S. Nie, K. Xue, Y . Cao, C. Li, H. Su, and J. Zhu, “All are worth words: A vit backbone for diffusion models,” inIEEE Conf. Comput. Vis. Pattern Recog., 2023, pp. 22 669–22 679. 1, 3

  7. [7]

    Sana: Efficient high-resolution text-to-image syn- thesis with linear diffusion transformers,

    E. Xie, J. Chen, J. Chen, H. Cai, H. Tang, Y . Lin, Z. Zhang, M. Li, L. Zhu, Y . Luet al., “Sana: Efficient high-resolution text-to-image syn- thesis with linear diffusion transformers,” inInt. Conf. Learn. Represent.,

  8. [8]

    High- resolution image synthesis with latent diffusion models,

    R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High- resolution image synthesis with latent diffusion models,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 10 684–10 695. 1, 3, 8, 9, 16, 18

Show all 90 references
  1. [9]

    Cogvideox: Text-to-video diffusion models with an expert transformer,

    Z. Yang, J. Teng, W. Zheng, M. Ding, S. Huang, J. Xu, Y . Yang, W. Hong, X. Zhang, G. Fenget al., “Cogvideox: Text-to-video diffusion models with an expert transformer,” inInt. Conf. Learn. Represent.,

  2. [10]

    Photorealistic video generation with diffusion models,

    A. Gupta, L. Yu, K. Sohn, X. Gu, M. Hahn, F.-F. Li, I. Essa, L. Jiang, and J. Lezama, “Photorealistic video generation with diffusion models,” inEur. Conf. Comput. Vis.Springer, 2024, pp. 393–411. 1

  3. [11]

    Diffusion based representation learning,

    S. Mittal, K. Abstreiter, S. Bauer, B. Sch ¨olkopf, and A. Mehrjou, “Diffusion based representation learning,” inInt. Conf. Mach. Learn. PMLR, 2023, pp. 24 963–24 982. 1, 3

  4. [12]

    Deconstructing denoising diffusion models for self-supervised learning,

    X. Chen, Z. Liu, S. Xie, and K. He, “Deconstructing denoising diffusion models for self-supervised learning,” inInt. Conf. Learn. Represent.,

  5. [13]

    Denoising diffusion autoencoders are unified self-supervised learners,

    W. Xiang, H. Yang, D. Huang, and Y . Wang, “Denoising diffusion autoencoders are unified self-supervised learners,” inInt. Conf. Comput. Vis., 2023, pp. 15 802–15 812. 1, 3

  6. [14]

    REPA-E: Unlocking vae for end-to-end tuning with latent diffusion transformers,

    X. Leng, J. Singh, Y . Hou, Z. Xing, S. Xie, and L. Zheng, “REPA-E: Unlocking vae for end-to-end tuning with latent diffusion transformers,” inInt. Conf. Comput. Vis., 2025, pp. 18 262–18 272. 1, 3, 8, 9, 18

  7. [15]

    Representation entanglement for generation: Training diffusion transformers is much easier than you think,

    G. Wu, S. Zhang, R. Shi, S. Gao, Z. Chen, L. Wang, Z. Chen, H. Gao, Y . Tang, M.-M. Chenget al., “Representation entanglement for generation: Training diffusion transformers is much easier than you think,”Adv. Neural Inform. Process. Syst., vol. 38, pp. 7714–7743, 2026. 1, 3, 8, 9, 18

  8. [16]

    No other representation component is needed: Diffusion transformers can provide representation guidance by themselves,

    D. Jiang, M. Wang, L. Li, L. Zhang, H. Wang, W. Wei, G. Dai, Y . Zhang, and J. Wang, “No other representation component is needed: Diffusion transformers can provide representation guidance by themselves,”arXiv preprint arXiv:2505.02831, 2025. 1, 2, 3, 8, 9, 11, 16, 19, 22

  9. [17]

    Similarity of neural network representations revisited,

    S. Kornblith, M. Norouzi, H. Lee, and G. Hinton, “Similarity of neural network representations revisited,” inInt. Conf. Mach. Learn.PMLR, 2019, pp. 3519–3529. 1, 3, 5, 15, 17, 18

  10. [18]

    Mean flows for one- step generative modeling,

    Z. Geng, M. Deng, X. Bai, J. Z. Kolter, and K. He, “Mean flows for one- step generative modeling,”Adv. Neural Inform. Process. Syst., vol. 38, pp. 75 460–75 482, 2026. 2, 8, 9, 10, 16, 19

  11. [19]

    Diversedit: Towards diverse representation learning in diffusion transformers,

    M. Yang, Z. Tan, B. Li, X. Yang, H. Chen, and H. Li, “Diversedit: Towards diverse representation learning in diffusion transformers,” in IEEE Conf. Comput. Vis. Pattern Recog., 2026. 3, 9

  12. [20]

    Representation learning: A review and new perspectives,

    Y . Bengio, A. Courville, and P. Vincent, “Representation learning: A review and new perspectives,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 35, no. 8, pp. 1798–1828, 2013. 3

  13. [21]

    Disentangled rep- resentation learning,

    X. Wang, H. Chen, S. Tang, Z. Wu, and W. Zhu, “Disentangled rep- resentation learning,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 46, no. 12, pp. 9677–9696, 2024. 3

  14. [22]

    Deep diversity- enhanced feature representation of hyperspectral images,

    J. Hou, Z. Zhu, J. Hou, H. Liu, H. Zeng, and D. Meng, “Deep diversity- enhanced feature representation of hyperspectral images,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 46, no. 12, pp. 8123–8138, 2024. 3

  15. [23]

    Self-supervised visual feature learning with deep neural networks: A survey,

    L. Jing and Y . Tian, “Self-supervised visual feature learning with deep neural networks: A survey,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 43, no. 11, pp. 4037–4058, 2021. 3

  16. [24]

    Bootstrap your own latent-a new approach to self-supervised learning,

    J.-B. Grill, F. Strub, F. Altch ´e, C. Tallec, P. Richemond, E. Buchatskaya, C. Doersch, B. Avila Pires, Z. Guo, M. Gheshlaghi Azaret al., “Bootstrap your own latent-a new approach to self-supervised learning,” 13 inAdv. Neural Inform. Process. Syst., vol. 33, 2020, pp. 21 271...

  17. [25]

    Dinov2: Learning robust visual features without supervision,

    M. Oquab, T. Darcet, T. Moutakanni, H. V o, M. Szafraniec, V . Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Noubyet al., “Dinov2: Learning robust visual features without supervision,”Transactions on Machine Learning Research Journal, 2024. 3, 5, 8, 16, 17

  18. [26]

    An empirical study of training self- supervised vision transformers,

    X. Chen, S. Xie, and K. He, “An empirical study of training self- supervised vision transformers,” inInt. Conf. Comput. Vis., 2021, pp. 9640–9649. 3, 5, 8, 17

  19. [27]

    Auto-encoding variational bayes,

    D. P. Kingma and M. Welling, “Auto-encoding variational bayes,” in Int. Conf. Learn. Represent., 2014. 3

  20. [28]

    Masked autoencoders are scalable vision learners,

    K. He, X. Chen, S. Xie, Y . Li, P. Doll ´ar, and R. Girshick, “Masked autoencoders are scalable vision learners,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 16 000–16 009. 3, 8, 17, 18

  21. [29]

    Simmim: A simple framework for masked image modeling,

    Z. Xie, Z. Zhang, Y . Cao, Y . Lin, J. Bao, Z. Yao, Q. Dai, and H. Hu, “Simmim: A simple framework for masked image modeling,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 9653–9663. 3

  22. [30]

    Diffusion model as representation learner,

    X. Yang and X. Wang, “Diffusion model as representation learner,” in Int. Conf. Comput. Vis., 2023, pp. 18 938–18 949. 3

  23. [31]

    Unsupervised representation learning from pre-trained diffusion probabilistic models,

    Z. Zhang, Z. Zhao, and Z. Lin, “Unsupervised representation learning from pre-trained diffusion probabilistic models,” inAdv. Neural Inform. Process. Syst., vol. 35, 2022, pp. 22 117–22 130. 3

  24. [32]

    Learning transferable visual models from natural language supervision,

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clarket al., “Learning transferable visual models from natural language supervision,” inInt. Conf. Mach. Learn., 2021, pp. 8748–8763. 3

  25. [33]

    Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

    J. Li, D. Li, C. Xiong, and S. Hoi, “Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,” inInt. Conf. Mach. Learn.PMLR, 2022, pp. 12 888–12 900. 3

  26. [34]

    Sigmoid loss for language image pre-training,

    X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer, “Sigmoid loss for language image pre-training,” inInt. Conf. Comput. Vis., 2023, pp. 11 975–11 986. 3

  27. [35]

    Siglip 2: Multilingual vision-language encoders with improved semantic understanding, localization, and dense features,

    M. Tschannen, A. Gritsenko, X. Wang, M. F. Naeem, I. Alabdul- mohsin, N. Parthasarathy, T. Evans, L. Beyer, Y . Xia, B. Mustafa et al., “Siglip 2: Multilingual vision-language encoders with improved semantic understanding, localization, and dense features,”arXiv preprint arXiv...

  28. [36]

    ImageNet: A large-scale hierarchical image database,

    J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A large-scale hierarchical image database,” inIEEE Conf. Comput. Vis. Pattern Recog., 2009, pp. 248–255. 3

  29. [37]

    Assessing generative models via precision and recall,

    M. S. Sajjadi, O. Bachem, M. Lucic, O. Bousquet, and S. Gelly, “Assessing generative models via precision and recall,”Adv. Neural Inform. Process. Syst., vol. 31, 2018. 3

  30. [38]

    Im- proved precision and recall metric for assessing generative models,

    T. Kynk ¨a¨anniemi, T. Karras, S. Laine, J. Lehtinen, and T. Aila, “Im- proved precision and recall metric for assessing generative models,” in Adv. Neural Inform. Process. Syst., vol. 32, 2019. 3, 8, 17

  31. [39]

    Reliability of cka as a similarity measure in deep learning,

    M. Davari, S. Horoi, A. Natik, G. Lajoie, G. Wolf, and E. Belilovsky, “Reliability of cka as a similarity measure in deep learning,” inInt. Conf. Learn. Represent., 2023. 3, 5, 17

  32. [40]

    Gans trained by a two time-scale update rule converge to a local nash equilibrium,

    M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, “Gans trained by a two time-scale update rule converge to a local nash equilibrium,”Adv. Neural Inform. Process. Syst., vol. 30, 2017. 3, 6, 8, 15, 17

  33. [41]

    Exposing flaws of generative model evaluation metrics and their unfair treatment of diffusion models,

    G. Stein, J. Cresswell, R. Hosseinzadeh, Y . Sui, B. Ross, V . Villecroze, Z. Liu, A. L. Caterini, E. Taylor, and G. Loaiza-Ganem, “Exposing flaws of generative model evaluation metrics and their unfair treatment of diffusion models,”Adv. Neural Inform. Process. Syst., vol. 36...

  34. [42]

    Generating images with sparse representations,

    C. Nash, J. Menick, S. Dieleman, and P. Battaglia, “Generating images with sparse representations,” inInt. Conf. Mach. Learn.PMLR, 2021, pp. 7958–7968. 3, 8, 17

  35. [43]

    Denoising diffusion implicit models,

    J. Song, C. Meng, and S. Ermon, “Denoising diffusion implicit models,” inInt. Conf. Learn. Represent., 2021. 3

  36. [44]

    Diffusion models in vision: A survey,

    F.-A. Croitoru, V . Hondru, R. T. Ionescu, and M. Shah, “Diffusion models in vision: A survey,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 9, pp. 10 850–10 869, 2023. 3

  37. [45]

    Pixart-α: Fast training of diffusion transformer for photorealistic text-to-image synthesis,

    J. Chen, J. Yu, C. Ge, L. Yao, E. Xie, Y . Wu, Z. Wang, J. Kwok, P. Luo, H. Luet al., “Pixart-α: Fast training of diffusion transformer for photorealistic text-to-image synthesis,” inInt. Conf. Learn. Represent.,

  38. [46]

    Hunyuanvideo: A systematic framework for large video generative models,

    W. Kong, Q. Tian, Z. Zhang, R. Min, Z. Dai, J. Zhou, J. Xiong, X. Li, B. Wu, J. Zhanget al., “Hunyuanvideo: A systematic framework for large video generative models,”arXiv preprint arXiv:2412.03603, 2024. 3

  39. [47]

    Flow matching for generative modeling,

    Y . Lipman, R. T. Chen, H. Ben-Hamu, M. Nickel, and M. Le, “Flow matching for generative modeling,” inInt. Conf. Learn. Represent., 2023. 3, 4

  40. [48]

    Flow straight and fast: Learning to generate and transfer data with rectified flow,

    X. Liu, C. Gonget al., “Flow straight and fast: Learning to generate and transfer data with rectified flow,” inInt. Conf. Learn. Represent.,

  41. [49]

    Diffusion models beat gans on image synthesis,

    P. Dhariwal and A. Nichol, “Diffusion models beat gans on image synthesis,” inAdv. Neural Inform. Process. Syst., vol. 34, 2021, pp. 8780–8794. 3, 8, 9, 17, 18

  42. [50]

    Stable video diffusion: Scaling latent video diffusion models to large datasets,

    A. Blattmann, T. Dockhorn, S. Kulal, D. Mendelevitch, M. Kilian, D. Lorenz, Y . Levi, Z. English, V . V oleti, A. Lettset al., “Stable video diffusion: Scaling latent video diffusion models to large datasets,”arXiv preprint arXiv:2311.15127, 2023. 3

  43. [51]

    SiT: Exploring flow and diffusion-based generative models with scalable interpolant transformers,

    N. Ma, M. Goldstein, M. S. Albergo, N. M. Boffi, E. Vanden-Eijnden, and S. Xie, “SiT: Exploring flow and diffusion-based generative models with scalable interpolant transformers,” inEur. Conf. Comput. Vis. Springer, 2024, pp. 23–40. 3, 8, 9, 16, 18

  44. [52]

    Dreamteacher: Pretraining image backbones with deep generative models,

    D. Li, H. Ling, A. Kar, D. Acuna, S. W. Kim, K. Kreis, A. Torralba, and S. Fidler, “Dreamteacher: Pretraining image backbones with deep generative models,” inInt. Conf. Comput. Vis., 2023, pp. 16 698–16 708. 3

  45. [53]

    Diffusion models and representation learning: A survey,

    M. Fuest, P. Ma, M. Gui, J. Schusterbauer, V . T. Hu, and B. Ommer, “Diffusion models and representation learning: A survey,”IEEE Trans. Pattern Anal. Mach. Intell., 2026. 3

  46. [54]

    SARA: Structural and adversarial representation alignment for training-efficient diffusion models,

    H. Chen, J. Wang, Z. Tan, and H. Li, “SARA: Structural and adversarial representation alignment for training-efficient diffusion models,”arXiv preprint arXiv:2503.08253, 2025. 3

  47. [55]

    Aligning text to image in diffusion models is easier than you think,

    J.-Y . Lee, B. Cha, J. Kim, and J. C. Ye, “Aligning text to image in diffusion models is easier than you think,”Adv. Neural Inform. Process. Syst., vol. 38, pp. 157 106–157 136, 2026. 3

  48. [56]

    Learning diffusion models with flexible representation guidance,

    C. Wang, C. Zhou, S. Gupta, J. Lin, S. Jegelka, S. Bates, and T. Jaakkola, “Learning diffusion models with flexible representation guidance,”Adv. Neural Inform. Process. Syst., vol. 38, pp. 131 176–131 222, 2026. 3, 4, 10, 19

  49. [57]

    What matters for representation alignment: Global information or spatial structure?

    J. Singh, X. Leng, Z. Wu, L. Zheng, R. Zhang, E. Shechtman, and S. Xie, “What matters for representation alignment: Global information or spatial structure?” inInt. Conf. Learn. Represent., 2026. 3, 9

  50. [58]

    Generative flows on discrete state-spaces: enabling multimodal flows with applications to protein co-design,

    A. Campbell, J. Yim, R. Barzilay, T. Rainforth, and T. Jaakkola, “Generative flows on discrete state-spaces: enabling multimodal flows with applications to protein co-design,” inInt. Conf. Mach. Learn., 2024, pp. 5453–5512. 3, 10, 16, 18, 19

  51. [59]

    Ro- bust deep learning–based protein sequence design using proteinmpnn,

    J. Dauparas, I. Anishchenko, N. Bennett, H. Bai, R. J. Ragotte, L. F. Milles, B. I. Wicky, A. Courbet, R. J. de Haas, N. Bethelet al., “Ro- bust deep learning–based protein sequence design using proteinmpnn,” Science, vol. 378, no. 6615, pp. 49–56, 2022. 3, 10, 16, 19

  52. [60]

    MiDi: Mixed graph and 3d denoising diffusion for molecule generation,

    C. Vignac, N. Osman, L. Toni, and P. Frossard, “MiDi: Mixed graph and 3d denoising diffusion for molecule generation,” inJoint European Con- ference on Machine Learning and Knowledge Discovery in Databases, 2023, pp. 560–576. 4

  53. [61]

    Navigating the design space of equivariant diffusion-based generative models for de novo 3d molecule generation,

    T. Le, J. Cremer, F. Noe, D.-A. Clevert, and K. T. Sch ¨utt, “Navigating the design space of equivariant diffusion-based generative models for de novo 3d molecule generation,” inInt. Conf. Learn. Represent., 2024. 4

  54. [62]

    Semlaflow–efficient 3d molecular generation with latent attention and equivariant flow matching,

    R. Irwin, A. Tibo, J. P. Janet, and S. Olsson, “Semlaflow–efficient 3d molecular generation with latent attention and equivariant flow matching,”arXiv preprint arXiv:2406.07266, 2024. 4, 11, 17, 19

  55. [63]

    Accurate structure prediction of biomolecular interactions with AlphaFold 3,

    J. Abramson, J. Adler, J. Dunber, R. Evanset al., “Accurate structure prediction of biomolecular interactions with AlphaFold 3,”Nature, vol. 630, pp. 493–500, 2024. 4, 19

  56. [64]

    Uni-mol: A universal 3d molecular representation learning framework,

    G. Zhou, Z. Gao, Q. Ding, H. Zheng, H. Xu, Z. Wei, L. Zhang, and G. Ke, “Uni-mol: A universal 3d molecular representation learning framework,” inInt. Conf. Learn. Represent., 2023. 4, 19

  57. [65]

    Measuring statistical dependence with hilbert-schmidt norms,

    A. Gretton, O. Bousquet, A. Smola, and B. Sch ¨olkopf, “Measuring statistical dependence with hilbert-schmidt norms,” inInternational Conference on Algorithmic Learning Theory, 2005, pp. 63–77. 5, 18

  58. [66]

    Masked autoencoders are scalable vision learners,

    K. He, X. Chen, S. Xie, Y . Li, P. Doll ´ar, and R. Girshick, “Masked autoencoders are scalable vision learners,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022. 5

  59. [67]

    Imagenet: A large-scale hierarchical image database,

    J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” inIEEE Conf. Comput. Vis. Pattern Recog., 2009, pp. 248–255. 8

  60. [68]

    Improved techniques for training gans,

    T. Salimans, I. Goodfellow, W. Zaremba, V . Cheung, A. Radford, and X. Chen, “Improved techniques for training gans,” inAdv. Neural Inform. Process. Syst., vol. 29, 2016. 8, 17

  61. [69]

    Classifier-free diffusion guidance,

    J. Ho and T. Salimans, “Classifier-free diffusion guidance,”arXiv preprint arXiv:2207.12598, 2022. 8

  62. [70]

    Understanding diffusion objectives as the elbo with simple data augmentation,

    D. Kingma and R. Gao, “Understanding diffusion objectives as the elbo with simple data augmentation,”Adv. Neural Inform. Process. Syst., vol. 36, pp. 65 484–65 516, 2023. 8, 9, 18 14

  63. [71]

    Cascaded diffusion models for high fidelity image generation,

    J. Ho, C. Saharia, W. Chan, D. J. Fleet, M. Norouzi, and T. Salimans, “Cascaded diffusion models for high fidelity image generation,”Journal of Machine Learning Research, vol. 23, no. 47, pp. 1–33, 2022. 8, 9, 18

  64. [72]

    Sd- dit: Unleashing the power of self-supervised discrimination in diffusion transformer,

    R. Zhu, Y . Pan, Y . Li, T. Yao, Z. Sun, T. Mei, and C. W. Chen, “Sd- dit: Unleashing the power of self-supervised discrimination in diffusion transformer,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 8435–8445. 8, 9, 18

  65. [73]

    Fast train- ing of diffusion models with masked transformers,

    H. Zheng, W. Nie, A. Vahdat, and A. Anandkumar, “Fast train- ing of diffusion models with masked transformers,”arXiv preprint arXiv:2306.09305, 2023. 8, 9, 18

  66. [74]

    Mdtv2: Masked dif- fusion transformer is a strong image synthesizer,

    S. Gao, P. Zhou, M.-M. Cheng, and S. Yan, “Mdtv2: Masked dif- fusion transformer is a strong image synthesizer,”arXiv preprint arXiv:2303.14389, 2023. 8, 9, 18

  67. [75]

    Improved techniques for training consistency models,

    Y . Song and P. Dhariwal, “Improved techniques for training consistency models,” inInt. Conf. Learn. Represent., 2024. 8, 10, 19

  68. [76]

    One step diffusion via shortcut models,

    K. Frans, D. Hafner, S. Levine, and P. Abbeel, “One step diffusion via shortcut models,” inInt. Conf. Learn. Represent., 2025. 8, 10, 19

  69. [77]

    Inductive moment matching,

    L. Zhou, S. Ermon, and J. Song, “Inductive moment matching,” inInt. Conf. Mach. Learn.PMLR, 2025, pp. 78 651–78 686. 8, 10, 19

  70. [78]

    Fine-tuning discrete diffusion models via reward optimization with applications to dna and protein design,

    C. Wang, M. Uehara, Y . He, A. Wang, A. Lal, T. Jaakkola, S. Levine, A. Regev, H. Wang, and T. Biancalani, “Fine-tuning discrete diffusion models via reward optimization with applications to dna and protein design,” inInt. Conf. Learn. Represent., 2025. 10, 16

  71. [79]

    Evolutionary-scale prediction of atomic-level protein structure with a language model,

    Z. Lin, H. Akin, R. Rao, B. Hie, Z. Zhuet al., “Evolutionary-scale prediction of atomic-level protein structure with a language model,” Science, vol. 379, no. 6637, pp. 1123–1130, 2023. 10, 18, 23

  72. [80]

    GEOM, energy-annotated molecular conformations for property prediction and molecular genera- tion,

    S. Axelrod and R. Gomez-Bombarelli, “GEOM, energy-annotated molecular conformations for property prediction and molecular genera- tion,”Scientific data, vol. 9, no. 1, p. 185, 2022. 11, 17, 24

  73. [81]

    Quantum chemistry structures and properties of 134 kilo molecules,

    R. Ramakrishnan, P. O. Dral, M. Rupp, and O. A. V on Lilienfeld, “Quantum chemistry structures and properties of 134 kilo molecules,” Scientific data, vol. 1, no. 1, pp. 1–7, 2014. 11, 25

  74. [82]

    The effective rank: A measure of effective dimensionality,

    O. Roy and M. Vetterli, “The effective rank: A measure of effective dimensionality,”European Signal Processing Conference, pp. 606–610,

  75. [83]

    Visualizing and understanding convolu- tional networks,

    M. D. Zeiler and R. Fergus, “Visualizing and understanding convolu- tional networks,” inEur. Conf. Comput. Vis., 2014. 15

  76. [84]

    Sampling is as easy as learning the score: theory for diffusion models with minimal data assumptions,

    S. Chen, S. Chewi, J. Li, Y . Li, A. Salim, and A. R. Zhang, “Sampling is as easy as learning the score: theory for diffusion models with minimal data assumptions,”arXiv preprint arXiv:2209.11215, 2022. 15

  77. [85]

    Decoupled weight decay regularization,

    I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” arXiv preprint arXiv:1711.05101, 2017. 16

  78. [86]

    Return of unconditional generation: A self- supervised representation generation method,

    T. Li, D. Katabi, and K. He, “Return of unconditional generation: A self- supervised representation generation method,” inAdv. Neural Inform. Process. Syst., vol. 37, 2024, pp. 125 441–125 468. 17

  79. [87]

    Momentum contrast for unsupervised visual representation learning,

    K. He, H. Fan, Y . Wu, S. Xie, and R. Girshick, “Momentum contrast for unsupervised visual representation learning,” inIEEE Conf. Comput. Vis. Pattern Recog., 2020, pp. 9729–9738. 17

  80. [88]

    Improved baselines with mo- mentum contrastive learning,

    X. Chen, H. Fan, R. Girshick, and K. He, “Improved baselines with mo- mentum contrastive learning,”arXiv preprint arXiv:2003.04297, 2020. 17

  81. [89]

    Rethinking the inception architecture for computer vision,

    C. Szegedy, V . Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the inception architecture for computer vision,” inIEEE Conf. Comput. Vis. Pattern Recog., 2016, pp. 2818–2826. 17

  82. [90]

    Improved denoising diffusion probabilis- tic models,

    A. Q. Nichol and P. Dhariwal, “Improved denoising diffusion probabilis- tic models,” inInt. Conf. Mach. Learn.PMLR, 2021, pp. 8162–8171. 19 15 APPENDIXA APPENDIXOVERVIEW This supplementary material is organized as follows: First, we provide a heuristic analysis of the WDS–FID ...

Pith tools