REVIEW 3 major objections 5 minor 62 references
UDT: Reconciling U-Nets and Diffusion Transformers with Data-Adaptive Token Reduction
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read UDT's token-merging U-Net reaches SiT's 1400-epoch FID in 40 epochs
desk verdict Real architectural contribution with an over-sold headline: single-run FIDs make the exact 40x speedup unverified, but the convergence story holds up. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is Token Merging (ToMe), here repurposed from an inference-speed technique into the down/upsampling operator of a U-shape transformer. It partitions tokens by bipartite soft matching, merges the most similar key-pairs with size-weighted averaging, applies proportional attention to correct for merged-token sizes, and unmerges by copying back along recorded indices. In UDT it progressively reduces tokens from 256 to a bottleneck of 112 ($N_{\text{Merge}} = 112$) with the hidden dimension kept at the DiT value $D$, then restores tokens symmetrically; the first encoder and last decoder block stay at full resolution and are skip-connected.
What would settle it
Measure the copy-back discrepancy at the bottleneck: unmerge tokens with the recorded indices and compare them with the original full-resolution features, on images dominated by fine texture such as fur, fabric, or foliage. If the reconstruction error grows sharply with merge rate or texture density, and FID degrades correspondingly, the information-preservation assumption fails; if unmerged features are near-identical, the assumption holds.
Extended reading notes
Core claim
The central claim is that a U-Net-shaped diffusion transformer whose downsampling and upsampling are implemented by token merging and unmerging preserves the DiT's isotropic token dimension and self-attention dynamics while giving it a true encoder-decoder hierarchy. Merging is driven by key similarity: redundant tokens such as backgrounds and flat regions are fused by weighted averaging, their merge indices are recorded, and the decoder copies merged tokens back to their original positions with skip connections carrying full-resolution information. The authors show that this beats both isotropic DiTs and earlier U-Net DiTs that use fixed 2x2 neighborhood downsamples with learnable projections, and that it aligns naturally with representation alignment because bottleneck features can be unmerged back to full token resolution for patch-wise matching.
Load-bearing premise
The load-bearing premise is that merging tokens by key similarity and later copying them back preserves the fine-grained information a diffusion model must reconstruct; if merging irreversibly discards detail-critical tokens, the encoder-decoder would lose exactly what later layers need.
Editorial extensions
If this is right
- UDT-XL/2+ with REPA reaches FID 7.6 without classifier-free guidance at 40 epochs, roughly 40x faster convergence than the SiT-XL/2 baseline's 1400 epochs, and 7.7 FID out-of-the-box at 80 epochs.
- With classifier-free guidance the model reports FID 1.38 after 320 epochs using the standard latent VAE and 1.35 after 500 epochs with an improved VAE, competitive with models trained two to three times longer.
- Because the down/upsampling is parameter-free and keeps the token dimension, the same blocks can replace DiT, pixel-space JiT, and the visual branch of MMDiT, improving FID in each case.
- Longer token sequences such as patch size 1 and 512x512 images are handled by aggressive early merging, with roughly 1.3-2.2x cost increase instead of 4.5x, reaching FID 1.71 at 512x512 from scratch.
- On 10% of ImageNet, UDT-L/2 reaches FID 10.6 at 500 epochs, beating SiT-L/2 trained on the full dataset at 80 epochs while using fewer training images.
Reading between the lines
- Beyond the paper: a testable reading is that data adaptivity, not mere resolution reduction, drives the gain; ablating token merging with random pooling at the same rates would settle whether similarity-based fusion or just hierarchy matters.
- Beyond the paper: the authors explicitly leave video generation and 2K resolution untested, but because token merging adapts to redundancy, one would expect the largest gains on high-resolution images where backgrounds occupy most of the frame; that expectation is not supported by the paper's evidence.
- Beyond the paper: the copy-back unmerge means every original token still receives a prediction, so the merge indices naturally give a pooling/unpooling pair that could be reused by dense prediction heads or segmentation objectives.
- Beyond the paper: the 10%-data result hints that merging acts as a structural prior for scarce-data regimes; a controlled experiment varying dataset size would test whether the convergence advantage grows as data shrinks.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper introduces UDT, a U-Net-shaped diffusion transformer that replaces fixed learnable spatial downsampling with data-adaptive token merging (ToMe) and exact unmerging, while preserving the token hidden dimension. The authors report that UDT accelerates convergence substantially, e.g., UDT-XL/2+REPA reaches FID 7.6 at 40 epochs without CFG versus SiT-XL/2's 7.9 FID at 1400 epochs (~40x), and achieves strong CFG results (FID 1.38 at 320 epochs with SD-VAE and 1.35 at 500 epochs with VA-VAE). The paper includes ablations of NMerge, r schedules, merging components, advanced techniques, and a range of drop-in replacements (DiT, JiT, LightningDiT, MMDiT), plus reduced-data and 512x512 experiments.
Significance. The core architectural idea is simple and plausible: merging semantically similar tokens is a transformer-native way to create a U-Net hierarchy, and the paper provides credible evidence that this outperforms learned projection-based downsampling in the same U-shape (Table 1b), as well as strong results across model sizes. If the convergence claims survive statistical scrutiny, the contribution is significant because it offers a drop-in, parameter-free (in the learned-parameter sense) U-Net DiT backbone compatible with REPA and T2I models. The paper also ships code and extensive experiments, which is a strength. The main caveat is that the headline quantitative claims rest on single-run FID point estimates without repeated-seed variation, so the magnitudes of the speedups are not yet pinned down.
major comments (3)
- [Section 4.2, Tables 2-4, Fig. 1(b), Appendix A.2] The central speedup claim (7.6 FID at 40 epochs vs 7.9 FID at 1400 epochs) is supported only by single-run FID estimates; Appendix A.2 specifies seed=0 for evaluation but no training-seed variation, standard deviations, or confidence intervals are reported anywhere. A 0.3 FID margin is of the same order as typical run-to-run variation for class-conditional ImageNet training at this scale, and the 40-epoch point appears only in a figure curve rather than in a numerical table. Please add at least three training seeds for the headline configuration (and ideally for Tables 2 and 4), report mean +/- std or confidence intervals, and provide a checkpoint-level table for the 40/80/320/500 epoch numbers. Without this, the '40x faster convergence' claim is not statistically distinguishable from 'comparable performance in far fewer epochs.'
- [Abstract and Section 4.2] The 7.9 FID baseline is ambiguous. The abstract attributes 7.9 to 'SiT ... at 1400 epochs (w/o CFG)', while Table 4 reports SiT-XL/2+REPA at 800 epochs with FID 7.9, and Section 4.2 says both '20-40x faster' and 'outperforming SiT-XL/2 + REPA'. Please state explicitly which checkpoint and configuration (SiT vs SiT+REPA, 800 vs 1400 epochs, CFG/no-CFG) is used for each speedup factor, and compute the factors consistently. If the relevant baseline is SiT-XL/2 at 1400 epochs, the source of that number should be cited rather than inferred from Table 4.
- [Appendix A.1, Table 10] The r-schedule notation contains apparent typos that make the architecture specification incomplete; for example, UDT-L/2 reads '15 (Enc 2-5), 12 (Enc 2-13)' and UDT-XL/2 reads '12 (Enc 2), 14 (Enc 6-11)', leaving the intervening encoder blocks unspecified. Please list the exact per-block r values for every model size, since the schedule is a central design choice and the paper elsewhere emphasizes that the improvement is purely architectural.
minor comments (5)
- [Fig. 1(b) and Fig. 4] Please distinguish curves by markers or line styles in addition to color, since color-only distinction is hard to read and the figures may be printed in grayscale.
- [Appendix A.2] State the training seed(s) used for each run, not only the evaluation seed, and report the number of training runs that each reported FID is based on.
- [Table 3 caption] Indicate clearly that the reported 'epochs' are training epochs and that FID is computed on 50K samples without class-balanced sampling unless stated; Appendix G should be referenced from the caption.
- [Appendix G, Fig. 7 caption] The term 'auto-guidance' appears in the figure caption but is never defined in the text; either define it or remove the reference.
- [Section 3 and Table 1(b)] The 'data-adaptive' terminology is used for a fixed key-similarity heuristic adopted from ToMe; please clarify that no per-sample learned adaptation is involved, and consider adding a convolutional projection baseline in Table 1(b) to strengthen the comparison against learned downsampling.
Circularity Check
No significant circularity: the result is an empirical architecture comparison against external baselines; the sole self-citation is non-load-bearing.
full rationale
The paper's central claim is an empirical architectural comparison: UDT reaches FID 7.6 at 40 epochs versus SiT-XL/2+REPA's 7.9 at 1400 epochs (Fig. 1, Tables 2/4). This is not a formal derivation, and no equation defines the claimed speedup in terms of the compared quantities. Token merging is explicitly imported from Bolya et al. (ToMe) and is independently established; the paper ablates its components in Table 1(c). REPA is an external method from Yu et al., not a self-citation. The only self-citation is reference [60], mentioned in the related-work survey as 'methods that promote linear separability [60]'; it is not load-bearing for UDT's architecture or results. Hyperparameters such as NMerge=112 and the r schedules are tuned on the B/2 model and transferred across scales, but they are not fitted to the headline FID values and are not renamed as predictions. The statistical caveat that the headline FIDs are single-run point estimates is a robustness/correctness concern, not circularity.
Assumptions & free parameters
free parameters (3)
- NMerge (bottleneck token count) =
112
- r schedule (token merge rate per encoder block) =
Varies by model size and patch size, e.g., 36 for B/2, 12/15 for XL/2, 512 then 256 for B/1
- CFG weight and guidance interval =
e.g., w=1.7, interval [0,0.7] for XL/2 on ImageNet; w=2.8, [0.3,1.0] for VA-VAE models
assumptions (3)
- standard math ToMe's bipartite matching with key-based similarity and proportional attention (Eq. 1) is a valid way to merge tokens without unacceptable information loss.
- domain assumption The 50K-sample FID evaluation protocol, as used by ADM, SiT and REPA, is a reliable proxy for generation quality.
- domain assumption Merged tokens can be exactly unmerged by copying the merged token back to the original positions, and this preserves the information needed for per-token denoising.
Cite this review
Pith. "Pith review of UDT: Reconciling U-Nets and Diffusion Transformers with Data-Adaptive Token Reduction." pith.science (2026). https://pith.science/paper/M25SIPR5
@misc{pith2026260801298,
author = {Pith},
title = {Pith review of: UDT: Reconciling U-Nets and Diffusion Transformers with Data-Adaptive Token Reduction},
year = {2026},
howpublished = {\url{https://pith.science/paper/M25SIPR5}},
note = {Machine review of arXiv:2608.01298}
}
read the original abstract
Diffusion Transformers (DiTs) have emerged as a core architecture in generative modeling due to their scalability and adaptability to multimodal tasks. DiTs comprise isotropic transformer blocks, and learn representations progressively across depth, where the denoising objective drives later layers to focus on fine-detail reconstruction. This results in degraded representation quality and an imbalanced encoder-decoder behavior. Prior approaches such as representation alignment (REPA) mitigate this by encouraging stronger early representations via training regularization. Alternatively, U-Net-style DiT architectures introduce explicit multi-scale encoder-decoder structures for improved convergence. But they build on standard U-Net wisdom via learnable operators for spatial downsampling, which are not well-suited to transformer architectures, introducing inefficiencies and compatibility issues with components such as cross-attention and representation regularization. In this work, we propose UDT, a U-Net diffusion transformer that combines the representation power of DiTs with the encoding-decoding benefits of U-Nets, through data-adaptive token merging for downsampling and upsampling, while preserving the DiT token dimension. Our baseline UDT architecture outperforms existing U-Net DiTs and achieves performance comparable to REPA across all model sizes. Furthermore, using architectural optimization and REPA, UDT outperforms SiT's 7.9 FID at 1400 epochs (w/o CFG) within 40 epochs (~ 40x faster convergence) for XL model size on 256x256 ImageNet. Finally, it achieves strong image generation performance with CFG, reaching FID of 1.38 (320 epochs) with SD-VAE and 1.35 (500 epochs) with VA-VAE, providing a new backbone for DiTs with strong empirical benefits.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
G. Alain and Y . Bengio. Understanding intermediate layers using linear classifier probes. InProc. Int. Conf. Learn. Represent., 2016
work page 2016
-
[2]
F. Bao, S. Nie, K. Xue, Y . Cao, C. Li, H. Su, and J. Zhu. All are worth words: A ViT backbone for diffusion models. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog., pages 22669–22679, 2023
work page 2023
-
[3]
D. Baranchuk, A. V oynov, I. Rubachev, V . Khrulkov, and A. Babenko. Label-efficient semantic segmentation with diffusion models. InProc. Int. Conf. Learn. Represent., 2022
work page 2022
-
[4]
D. Bolya, C.-Y . Fu, X. Dai, P. Zhang, C. Feichtenhofer, and J. Hoffman. Token merging: Your ViT but faster. InProc. Int. Conf. Learn. Represent., 2022
work page 2022
-
[5]
D. Bolya and J. Hoffman. Token merging for fast stable diffusion. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog., pages 4599–4603, 2023
work page 2023
-
[6]
J. Chen, C. Ge, E. Xie, Y . Wu, L. Yao, X. Ren, Z. Wang, P. Luo, H. Lu, and Z. Li. PixArt-σ: Weak-to-strong training of diffusion transformer for 4K text-to-image generation. InProc. Eur. Conf. Comput. Vis., pages 74–91. Springer, 2024
work page 2024
-
[7]
X. Chen, Z. Liu, S. Xie, and K. He. Deconstructing denoising diffusion models for self-supervised learning. InProc. Int. Conf. Learn. Represent., 2025
work page 2025
-
[8]
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. Imagenet: A large-scale hierarchical image database. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog., pages 248–255, 2009
work page 2009
Show all 62 references
-
[9]
Dhariwal and A
P. Dhariwal and A. Nichol. Diffusion models beat GANs on image synthesis. InProc. Adv. Neural Inf. Process. Syst., pages 8780–8794, 2021
2021
-
[10]
X. Ding, X. Zhang, N. Ma, J. Han, G. Ding, and J. Sun. Repvgg: Making vgg-style convnets great again. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recog., pages 13733–13742, 2021
2021
-
[11]
Dosovitskiy, L
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. InProc. Int. Conf. Learn. Represent., 2021
2021
-
[12]
Esser, S
P. Esser, S. Kulal, A. Blattmann, R. Entezari, J. Müller, H. Saini, Y . Levi, D. Lorenz, A. Sauer, F. Boesel, et al. Scaling rectified flow transformers for high-resolution image synthesis. InProc. Int. Conf. Mach. Learn., 2024
2024
-
[13]
K. He, X. Chen, S. Xie, Y . Li, P. Dollár, and R. Girshick. Masked autoencoders are scalable vision learners. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog., pages 16000–16009, 2022
2022
-
[14]
Heusel, H
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter. GANs trained by a two time-scale update rule converge to a local nash equilibrium. InProc. Adv. Neural Inf. Process. Syst., 2017
2017
-
[15]
J. Ho, A. Jain, and P. Abbeel. Denoising diffusion probabilistic models. InProc. Adv. Neural Inf. Process. Syst., pages 6840–6851, 2020
2020
-
[16]
J. Ho, C. Saharia, W. Chan, D. J. Fleet, M. Norouzi, and T. Salimans. Cascaded diffusion models for high fidelity image generation.J. Mach. Learn. Res., 23(47):1–33, 2022
2022
-
[17]
Ho and T
J. Ho and T. Salimans. Classifier-free diffusion guidance. InProc. NeurIPS Workshop DGMs Appl., 2021
2021
-
[18]
Hoogeboom, J
E. Hoogeboom, J. Heek, and T. Salimans. Simple diffusion: End-to-end diffusion for high resolution images. InProc. Int. Conf. Mach. Learn., pages 13213–13232, 2023
2023
-
[19]
Jiang, M
D. Jiang, M. Wang, L. Li, L. Zhang, H. Wang, W. Wei, G. Dai, Y . Zhang, and J. Wang. No other representation component is needed: Diffusion transformers can provide representation guidance by themselves. InProc. Int. Conf. Learn. Represent., 2026
2026
-
[20]
Karras, M
T. Karras, M. Aittala, T. Aila, and S. Laine. Elucidating the design space of diffusion-based generative models. InProc. Adv. Neural Inf. Process. Syst., pages 26565–26577, 2022
2022
-
[21]
Kinga, J
D. Kinga, J. B. Adam, et al. A method for stochastic optimization. InProc. Int. Conf. Learn. Represent., 2015
2015
-
[22]
Kingma and R
D. Kingma and R. Gao. Understanding diffusion objectives as the elbo with simple data augmentation. In Proc. Adv. Neural Inf. Process. Syst., pages 65484–65516, 2023. 11
2023
-
[23]
Z. Kong, P. Dong, X. Ma, X. Meng, M. Sun, W. Niu, X. Shen, G. Yuan, B. Ren, M. Qin, et al. SPViT: enabling faster vision transformers via soft token pruning. InProc. Eur. Conf. Comput. Vis., 2022
2022
-
[24]
Krause, T
F. Krause, T. Phan, M. Gui, S. A. Baumann, V . T. Hu, and B. Ommer. TREAD: Token routing for efficient architecture-agnostic diffusion training, 2025. arXiv:2501.04765
2025
-
[25]
Kynkäänniemi, M
T. Kynkäänniemi, M. Aittala, T. Karras, S. Laine, T. Aila, and J. Lehtinen. Applying guidance in a limited interval improves sample and distribution quality in diffusion models. InProc. Adv. Neural Inf. Process. Syst., 2024
2024
-
[26]
X. Leng, J. Singh, Y . Hou, Z. Xing, S. Xie, and L. Zheng. REPA-E: Unlocking V AE for end-to-end tuning with latent diffusion transformers, 2025
2025
-
[27]
T. Li, H. Chang, S. Mishra, H. Zhang, D. Katabi, and D. Krishnan. MAGE: Masked generative encoder to unify representation learning and image synthesis. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog., pages 2142–2152, 2023
2023
-
[28]
Li and K
T. Li and K. He. Back to basics: Let denoising generative models denoise.arXiv preprint arXiv:2511.13720, 2025
2025 arXiv
-
[29]
Liang, C
Y . Liang, C. Ge, Z. Tong, Y . Song, J. Wang, and P. Xie. Not all patches are what you need: Expediting vision transformers via token reorganizations. InProc. Int. Conf. Learn. Represent., 2022
2022
-
[30]
T.-Y . Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick. Microsoft COCO: Common objects in context. InProc. Eur. Conf. Comput. Vis., pages 740–755, 2014
2014
-
[31]
Lipman, R
Y . Lipman, R. T. Q. Chen, H. Ben-Hamu, M. Nickel, and M. Le. Flow matching for generative modeling. In Proc. Int. Conf. Learn. Represent., 2023
2023
-
[32]
Litjens, T
G. Litjens, T. Kooi, B. E. Bejnordi, A. A. A. Setio, F. Ciompi, M. Ghafoorian, J. A. Van Der Laak, B. Van Ginneken, and C. I. Sánchez. A survey on deep learning in medical image analysis.Med. Image Anal., 42:60–88, 2017
2017
-
[33]
Z. Liu, Y . Lin, Y . Cao, H. Hu, Y . Wei, Z. Zhang, S. Lin, and B. Guo. Swin transformer: Hierarchical vision transformer using shifted windows. InProc. IEEE/CVF Int. Conf. Comput. Vis., pages 10012–10022, 2021
2021
-
[34]
Loshchilov and F
I. Loshchilov and F. Hutter. Decoupled weight decay regularization. InProc. Int. Conf. Learn. Represent., 2019
2019
-
[35]
N. Ma, M. Goldstein, M. S. Albergo, N. M. Boffi, E. Vanden-Eijnden, and S. Xie. SiT: Exploring flow and diffusion-based generative models with scalable interpolant transformers. InProc. Eur. Conf. Comput. Vis., pages 23–40, 2024
2024
-
[36]
Marin, J.-H
D. Marin, J.-H. R. Chang, A. Ranjan, A. Prabhu, M. Rastegari, and O. Tuzel. Token pooling in vision transformers.arXiv preprint arXiv:2110.03860, 2021
2021 arXiv
-
[37]
L. Meng, H. Li, B.-C. Chen, S. Lan, Z. Wu, Y .-G. Jiang, and S.-N. Lim. AdaViT: Adaptive vision transformers for efficient image recognition. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog., pages 12309–12318, 2022
2022
-
[38]
Oquab, T
M. Oquab, T. Darcet, T. Moutakanni, H. V . V o, M. Szafraniec, V . Khalidov, P. Fernandez, D. HAZIZA, F. Massa, A. El-Nouby, M. Assran, N. Ballas, W. Galuba, R. Howes, P.-Y . Huang, S.-W. Li, I. Misra, M. Rabbat, V . Sharma, G. Synnaeve, H. Xu, H. Jegou, J. Mairal, P. Labatut,...
2024
-
[39]
Peebles and S
W. Peebles and S. Xie. Scalable diffusion models with transformers. InProc. IEEE/CVF Int. Conf. Comput. Vis., pages 4195–4205, 2023
2023
-
[40]
Y . Rao, W. Zhao, B. Liu, J. Lu, J. Zhou, and C.-J. Hsieh. DynamicViT: Efficient vision transformers with dynamic token sparsification. InProc. Adv. Neural Inf. Process. Syst., volume 34, pages 13937–13949, 2021
2021
-
[41]
Rombach, A
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer. High-resolution image synthesis with latent diffusion models. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog., pages 10684–10695, 2022
2022
-
[42]
Ronneberger, P
O. Ronneberger, P. Fischer, and T. Brox. U-net: Convolutional networks for biomedical image segmentation. InProc. Int. Conf. Med. Image Comput. Comput.-Assist. Intervent., pages 234–241, 2015
2015
-
[43]
M. Ryoo, A. Piergiovanni, A. Arnab, M. Dehghani, and A. Angelova. Tokenlearner: Adaptive space-time tokenization for videos. InProc. Adv. Neural Inf. Process. Syst., volume 34, pages 12786–12797, 2021. 12
2021
-
[44]
Salimans, I
T. Salimans, I. Goodfellow, W. Zaremba, V . Cheung, A. Radford, and X. Chen. Improved techniques for training GANs. InProc. Adv. Neural Inf. Process. Syst., 2016
2016
-
[45]
N. Shazeer. Glu variants improve transformer.arXiv preprint arXiv:2002.05202, 2020
2002 arXiv
-
[46]
J. Su, M. Ahmed, Y . Lu, S. Pan, W. Bo, and Y . Liu. Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024
2024
-
[47]
Y . Tian, H. Chen, M. Zheng, Y . Liang, C. Xu, and Y . Wang. U-REPA: Aligning Diffusion U-Nets to ViTs. In Proc. Adv. Neural Inf. Process. Syst., 2026
2026
-
[48]
Y . Tian, Z. Tu, H. Chen, J. Hu, C. Xu, and Y . Wang. U-DiTs: Downsample tokens in u-shaped diffusion transformers. InProc. Adv. Neural Inf. Process. Syst., volume 37, pages 51994–52013, 2024
2024
-
[49]
Vaswani, N
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin. Attention is all you need. InProc. Adv. Neural Inf. Process. Syst., pages 6000–6010, 2017
2017
-
[50]
J. Wang, N. Kang, L. Yao, M. Chen, C. Wu, S. Zhang, S. Xue, Y . Liu, T. Wu, X. Liu, et al. LiT: Delving into a simple linear diffusion transformer for image generation. InICCV, pages 16068–16078, 2025
2025
-
[51]
Wang and K
R. Wang and K. He. Diffuse and disperse: Image generation with representation regularization.arXiv preprint arXiv:2506.09027, 2025
2025 arXiv
-
[52]
S. Wang, Z. Tian, W. Huang, and L. Wang. DDT: Decoupled diffusion transformer.Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recog., 2026
2026
-
[53]
Z. Wang, X. Cun, J. Bao, W. Zhou, J. Liu, and H. Li. Uformer: A general u-shaped transformer for image restoration. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog., pages 17683–17693, 2022
2022
-
[54]
Xiang, H
W. Xiang, H. Yang, D. Huang, and Y . Wang. Denoising diffusion autoencoders are unified self-supervised learners. InProc. IEEE/CVF Int. Conf. Comput. Vis., pages 15802–15812, 2023
2023
-
[55]
E. Xie, J. Chen, J. Chen, H. Cai, H. Tang, Y . Lin, Z. Zhang, M. Li, L. Zhu, Y . Lu, et al. Sana: Efficient high-resolution image synthesis with linear diffusion transformers.arXiv preprint arXiv:2410.10629, 2024
2024 arXiv
-
[56]
J. Yao, C. Wang, W. Liu, and X. Wang. FasterDiT: Towards faster diffusion transformers training without architecture modification. 37:56166–56189, 2024
2024
-
[57]
J. Yao, B. Yang, and X. Wang. Reconstruction vs. generation: Taming optimization dilemma in latent diffusion models. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog., pages 15703–15712, 2025
2025
-
[58]
H. Yin, A. Vahdat, J. M. Alvarez, A. Mallya, J. Kautz, and P. Molchanov. A-ViT: Adaptive tokens for efficient vision transformer. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog., pages 10809–10818, 2022
2022
-
[59]
S. Yu, S. Kwak, H. Jang, J. Jeong, J. Huang, J. Shin, and S. Xie. Representation alignment for generation: Training diffusion transformers is easier than you think. InProc. Int. Conf. Learn. Represent., 2025
2025
-
[60]
J. Yun, Y . U. Alçalar, and M. Akçakaya. No alignment needed for generation: Learning linearly separable representations in diffusion models, 2025. arXiv:2509.21565
2025
-
[61]
Zheng, N
B. Zheng, N. Ma, S. Tong, and S. Xie. Diffusion transformers with representation autoencoders. InProc. Int. Conf. Learn. Represent., 2026
2026
-
[62]
Zheng, W
H. Zheng, W. Nie, A. Vahdat, and A. Anandkumar. Fast training of diffusion models with masked transformers. InTrans. Mach. Learn. Res., 2024. 13 Appendix A Implementation Details A.1 Model Configurations All trainings were conducted from scratch, using 4 NVIDIA A100 GPUs for t...
2024
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.