Pith. sign in

REVIEW 3 major objections 5 minor 62 references

UDT: Reconciling U-Nets and Diffusion Transformers with Data-Adaptive Token Reduction

T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read UDT's token-merging U-Net reaches SiT's 1400-epoch FID in 40 epochs

desk verdict Real architectural contribution with an over-sold headline: single-run FIDs make the exact 40x speedup unverified, but the convergence story holds up. read the letter →

arxiv 2608.01298 v1 pith:M25SIPR5 submitted 2026-08-02 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords diffusiontransformerstokenmergingU-NetarchitectureimagegenerationrepresentationalignmentflowmatchingtrainingconvergenceNet
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that diffusion transformers can gain the convergence benefits of U-Nets without learnable spatial downsampling: replace fixed grid pooling with data-adaptive token merging in a U-shaped encoder-decoder, while keeping the token dimension constant. It claims that this single architectural change makes a flow-based transformer on ImageNet 256x256 reach, without classifier-free guidance, FID 7.6 after 40 epochs, where the isotropic baseline needs 1400 epochs for FID 7.9. With classifier-free guidance and representation alignment, it reports FID 1.38 after 320 epochs. The broader point is that representation quality and denoising capacity can both improve from an architecture that compresses redundant tokens instead of blurring neighboring ones.

What carries the argument

The central object is Token Merging (ToMe), here repurposed from an inference-speed technique into the down/upsampling operator of a U-shape transformer. It partitions tokens by bipartite soft matching, merges the most similar key-pairs with size-weighted averaging, applies proportional attention to correct for merged-token sizes, and unmerges by copying back along recorded indices. In UDT it progressively reduces tokens from 256 to a bottleneck of 112 ($N_{\text{Merge}} = 112$) with the hidden dimension kept at the DiT value $D$, then restores tokens symmetrically; the first encoder and last decoder block stay at full resolution and are skip-connected.

What would settle it

Measure the copy-back discrepancy at the bottleneck: unmerge tokens with the recorded indices and compare them with the original full-resolution features, on images dominated by fine texture such as fur, fabric, or foliage. If the reconstruction error grows sharply with merge rate or texture density, and FID degrades correspondingly, the information-preservation assumption fails; if unmerged features are near-identical, the assumption holds.

Watch

Extended reading notes

Core claim

The central claim is that a U-Net-shaped diffusion transformer whose downsampling and upsampling are implemented by token merging and unmerging preserves the DiT's isotropic token dimension and self-attention dynamics while giving it a true encoder-decoder hierarchy. Merging is driven by key similarity: redundant tokens such as backgrounds and flat regions are fused by weighted averaging, their merge indices are recorded, and the decoder copies merged tokens back to their original positions with skip connections carrying full-resolution information. The authors show that this beats both isotropic DiTs and earlier U-Net DiTs that use fixed 2x2 neighborhood downsamples with learnable projections, and that it aligns naturally with representation alignment because bottleneck features can be unmerged back to full token resolution for patch-wise matching.

Load-bearing premise

The load-bearing premise is that merging tokens by key similarity and later copying them back preserves the fine-grained information a diffusion model must reconstruct; if merging irreversibly discards detail-critical tokens, the encoder-decoder would lose exactly what later layers need.

Editorial extensions

If this is right

  • UDT-XL/2+ with REPA reaches FID 7.6 without classifier-free guidance at 40 epochs, roughly 40x faster convergence than the SiT-XL/2 baseline's 1400 epochs, and 7.7 FID out-of-the-box at 80 epochs.
  • With classifier-free guidance the model reports FID 1.38 after 320 epochs using the standard latent VAE and 1.35 after 500 epochs with an improved VAE, competitive with models trained two to three times longer.
  • Because the down/upsampling is parameter-free and keeps the token dimension, the same blocks can replace DiT, pixel-space JiT, and the visual branch of MMDiT, improving FID in each case.
  • Longer token sequences such as patch size 1 and 512x512 images are handled by aggressive early merging, with roughly 1.3-2.2x cost increase instead of 4.5x, reaching FID 1.71 at 512x512 from scratch.
  • On 10% of ImageNet, UDT-L/2 reaches FID 10.6 at 500 epochs, beating SiT-L/2 trained on the full dataset at 80 epochs while using fewer training images.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: a testable reading is that data adaptivity, not mere resolution reduction, drives the gain; ablating token merging with random pooling at the same rates would settle whether similarity-based fusion or just hierarchy matters.
  • Beyond the paper: the authors explicitly leave video generation and 2K resolution untested, but because token merging adapts to redundancy, one would expect the largest gains on high-resolution images where backgrounds occupy most of the frame; that expectation is not supported by the paper's evidence.
  • Beyond the paper: the copy-back unmerge means every original token still receives a prediction, so the merge indices naturally give a pooling/unpooling pair that could be reused by dense prediction heads or segmentation objectives.
  • Beyond the paper: the 10%-data result hints that merging acts as a structural prior for scarce-data regimes; a controlled experiment varying dataset size would test whether the convergence advantage grows as data shrinks.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper introduces UDT, a U-Net-shaped diffusion transformer that replaces fixed learnable spatial downsampling with data-adaptive token merging (ToMe) and exact unmerging, while preserving the token hidden dimension. The authors report that UDT accelerates convergence substantially, e.g., UDT-XL/2+REPA reaches FID 7.6 at 40 epochs without CFG versus SiT-XL/2's 7.9 FID at 1400 epochs (~40x), and achieves strong CFG results (FID 1.38 at 320 epochs with SD-VAE and 1.35 at 500 epochs with VA-VAE). The paper includes ablations of NMerge, r schedules, merging components, advanced techniques, and a range of drop-in replacements (DiT, JiT, LightningDiT, MMDiT), plus reduced-data and 512x512 experiments.

Significance. The core architectural idea is simple and plausible: merging semantically similar tokens is a transformer-native way to create a U-Net hierarchy, and the paper provides credible evidence that this outperforms learned projection-based downsampling in the same U-shape (Table 1b), as well as strong results across model sizes. If the convergence claims survive statistical scrutiny, the contribution is significant because it offers a drop-in, parameter-free (in the learned-parameter sense) U-Net DiT backbone compatible with REPA and T2I models. The paper also ships code and extensive experiments, which is a strength. The main caveat is that the headline quantitative claims rest on single-run FID point estimates without repeated-seed variation, so the magnitudes of the speedups are not yet pinned down.

major comments (3)
  1. [Section 4.2, Tables 2-4, Fig. 1(b), Appendix A.2] The central speedup claim (7.6 FID at 40 epochs vs 7.9 FID at 1400 epochs) is supported only by single-run FID estimates; Appendix A.2 specifies seed=0 for evaluation but no training-seed variation, standard deviations, or confidence intervals are reported anywhere. A 0.3 FID margin is of the same order as typical run-to-run variation for class-conditional ImageNet training at this scale, and the 40-epoch point appears only in a figure curve rather than in a numerical table. Please add at least three training seeds for the headline configuration (and ideally for Tables 2 and 4), report mean +/- std or confidence intervals, and provide a checkpoint-level table for the 40/80/320/500 epoch numbers. Without this, the '40x faster convergence' claim is not statistically distinguishable from 'comparable performance in far fewer epochs.'
  2. [Abstract and Section 4.2] The 7.9 FID baseline is ambiguous. The abstract attributes 7.9 to 'SiT ... at 1400 epochs (w/o CFG)', while Table 4 reports SiT-XL/2+REPA at 800 epochs with FID 7.9, and Section 4.2 says both '20-40x faster' and 'outperforming SiT-XL/2 + REPA'. Please state explicitly which checkpoint and configuration (SiT vs SiT+REPA, 800 vs 1400 epochs, CFG/no-CFG) is used for each speedup factor, and compute the factors consistently. If the relevant baseline is SiT-XL/2 at 1400 epochs, the source of that number should be cited rather than inferred from Table 4.
  3. [Appendix A.1, Table 10] The r-schedule notation contains apparent typos that make the architecture specification incomplete; for example, UDT-L/2 reads '15 (Enc 2-5), 12 (Enc 2-13)' and UDT-XL/2 reads '12 (Enc 2), 14 (Enc 6-11)', leaving the intervening encoder blocks unspecified. Please list the exact per-block r values for every model size, since the schedule is a central design choice and the paper elsewhere emphasizes that the improvement is purely architectural.
minor comments (5)
  1. [Fig. 1(b) and Fig. 4] Please distinguish curves by markers or line styles in addition to color, since color-only distinction is hard to read and the figures may be printed in grayscale.
  2. [Appendix A.2] State the training seed(s) used for each run, not only the evaluation seed, and report the number of training runs that each reported FID is based on.
  3. [Table 3 caption] Indicate clearly that the reported 'epochs' are training epochs and that FID is computed on 50K samples without class-balanced sampling unless stated; Appendix G should be referenced from the caption.
  4. [Appendix G, Fig. 7 caption] The term 'auto-guidance' appears in the figure caption but is never defined in the text; either define it or remove the reference.
  5. [Section 3 and Table 1(b)] The 'data-adaptive' terminology is used for a fixed key-similarity heuristic adopted from ToMe; please clarify that no per-sample learned adaptation is involved, and consider adding a convolutional projection baseline in Table 1(b) to strengthen the comparison against learned downsampling.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the result is an empirical architecture comparison against external baselines; the sole self-citation is non-load-bearing.

full rationale

The paper's central claim is an empirical architectural comparison: UDT reaches FID 7.6 at 40 epochs versus SiT-XL/2+REPA's 7.9 at 1400 epochs (Fig. 1, Tables 2/4). This is not a formal derivation, and no equation defines the claimed speedup in terms of the compared quantities. Token merging is explicitly imported from Bolya et al. (ToMe) and is independently established; the paper ablates its components in Table 1(c). REPA is an external method from Yu et al., not a self-citation. The only self-citation is reference [60], mentioned in the related-work survey as 'methods that promote linear separability [60]'; it is not load-bearing for UDT's architecture or results. Hyperparameters such as NMerge=112 and the r schedules are tuned on the B/2 model and transferred across scales, but they are not fitted to the headline FID values and are not renamed as predictions. The statistical caveat that the headline FIDs are single-run point estimates is a robustness/correctness concern, not circularity.

Assumptions & free parameters 3 free parameters · 3 assumptions · 0 invented entities

No new physical entities or forces are posited. The central claim rests on tuned token-merging hyperparameters and the assumption that merging does not discard detail-critical tokens. The free parameters (NMerge, r schedules, CFG settings) are the main fitted components, while the axioms are standard algorithmic assumptions inherited from ToMe and the FID evaluation protocol.

free parameters (3)
  • NMerge (bottleneck token count) = 112
    Selected by validation FID sweep in Table 1a; used across all model sizes and patch sizes.
  • r schedule (token merge rate per encoder block) = Varies by model size and patch size, e.g., 36 for B/2, 12/15 for XL/2, 512 then 256 for B/1
    Chosen per model configuration (Appendix A) and not derived from a principle; affects the encoder-decoder structure and computational savings.
  • CFG weight and guidance interval = e.g., w=1.7, interval [0,0.7] for XL/2 on ImageNet; w=2.8, [0.3,1.0] for VA-VAE models
    Tuned per model variant for system-level FID comparisons; small changes in these values can affect reported FID.
assumptions (3)
  • standard math ToMe's bipartite matching with key-based similarity and proportional attention (Eq. 1) is a valid way to merge tokens without unacceptable information loss.
    The paper relies on the ToMe algorithm and its properties, citing [4], and does not re-derive or verify the merging quality for the diffusion training setting.
  • domain assumption The 50K-sample FID evaluation protocol, as used by ADM, SiT and REPA, is a reliable proxy for generation quality.
    All model comparisons use this protocol, but no repeated seeds or confidence intervals are reported, so the single-run FID values assume measurement stability.
  • domain assumption Merged tokens can be exactly unmerged by copying the merged token back to the original positions, and this preserves the information needed for per-token denoising.
    The decoder uses unmerging with recorded indices (Section 3). If copy-back loses information that cannot be recovered by skip connections, the architecture would be limited; the paper provides empirical evidence but no formal guarantee.

how reviews work

0 comments
Cite this review

Pith. "Pith review of UDT: Reconciling U-Nets and Diffusion Transformers with Data-Adaptive Token Reduction." pith.science (2026). https://pith.science/paper/M25SIPR5

@misc{pith2026260801298,
  author       = {Pith},
  title        = {Pith review of: UDT: Reconciling U-Nets and Diffusion Transformers with Data-Adaptive Token Reduction},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/M25SIPR5}},
  note         = {Machine review of arXiv:2608.01298}
}
read the original abstract

Diffusion Transformers (DiTs) have emerged as a core architecture in generative modeling due to their scalability and adaptability to multimodal tasks. DiTs comprise isotropic transformer blocks, and learn representations progressively across depth, where the denoising objective drives later layers to focus on fine-detail reconstruction. This results in degraded representation quality and an imbalanced encoder-decoder behavior. Prior approaches such as representation alignment (REPA) mitigate this by encouraging stronger early representations via training regularization. Alternatively, U-Net-style DiT architectures introduce explicit multi-scale encoder-decoder structures for improved convergence. But they build on standard U-Net wisdom via learnable operators for spatial downsampling, which are not well-suited to transformer architectures, introducing inefficiencies and compatibility issues with components such as cross-attention and representation regularization. In this work, we propose UDT, a U-Net diffusion transformer that combines the representation power of DiTs with the encoding-decoding benefits of U-Nets, through data-adaptive token merging for downsampling and upsampling, while preserving the DiT token dimension. Our baseline UDT architecture outperforms existing U-Net DiTs and achieves performance comparable to REPA across all model sizes. Furthermore, using architectural optimization and REPA, UDT outperforms SiT's 7.9 FID at 1400 epochs (w/o CFG) within 40 epochs (~ 40x faster convergence) for XL model size on 256x256 ImageNet. Finally, it achieves strong image generation performance with CFG, reaching FID of 1.38 (320 epochs) with SD-VAE and 1.35 (500 epochs) with VA-VAE, providing a new backbone for DiTs with strong empirical benefits.

Figures

Figures reproduced from arXiv: 2608.01298 by the authors.

Figure 1
Figure 1. (a) Architecture of the proposed method UDT. It preserves the token hidden dimension (D) while gradually reducing and restoring the token sequence length via data-adaptive token merging and unmerging. (b) FID vs. Epoch on XL models for ImageNet 256 × 256 without classifier-free guidance. UDT (our baseline, marked in light pink) and UDT+ (baseline + advanced techniques, marked in pink), achieve 7.7 FID at 80 epochs a… view at source ↗
Figure 2
Figure 2. Representation Analysis. (a) PCA visualization of intermediate layers. Features are extracted from the bottleneck layer (e.g., layer 18 of 24 for SiT-L/2, layer 16 of 32 for SiT↓, and layer 12 of 24 for UDT). Features with 112 tokens in our method are unmerged for visualization. (b) Linear probing evaluation on pretrained models across layers. All experiments are conducted at noise level t = 0.1 using Large models t… view at source ↗
Figure 3
Figure 3. Visualization of token reduction: (a) Data-adaptive token reduction (ours) gradually merges [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: FID vs. Epoch on ImageNet 256 × 256 without CFG [PITH_FULL_IMAGE:figures/full_fig_p009_4.png]
Figure 5
Figure 5. Figure 5: FID vs. Wall-Clock Time (hours) on ImageNet [PITH_FULL_IMAGE:figures/full_fig_p018_5.png]
Figure 6
Figure 6. Figure 6: PCA visualization of intermediate features at [PITH_FULL_IMAGE:figures/full_fig_p019_6.png]
Figure 7
Figure 7. Figure 7: Evaluation with class-balanced 50K sam￾pling on ImageNet 256 × 256. Gray: results with auto￾guidance instead of CFG. Model Epochs Vis. Enc. FID↓ SiT-XL/2 [35] 1400 – 1.95 SiT-XL/2 + REPA [59] 800 DINOv2 1.29 DDT-XL/2† + REPA [52] 400 DINOv2 1.26 UDT-XL/2+ + REPA (Ours)…
Figure 8
Figure 8. Figure 8: Uncurated ImageNet samples at resolutions of [PITH_FULL_IMAGE:figures/full_fig_p020_8.png]
Figure 9
Figure 9. Figure 9: Uncurated ImageNet samples at resolutions of [PITH_FULL_IMAGE:figures/full_fig_p021_9.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

62 extracted references · 55 canonical work pages

  1. [1]

    Alain and Y

    G. Alain and Y . Bengio. Understanding intermediate layers using linear classifier probes. InProc. Int. Conf. Learn. Represent., 2016

  2. [2]

    F. Bao, S. Nie, K. Xue, Y . Cao, C. Li, H. Su, and J. Zhu. All are worth words: A ViT backbone for diffusion models. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog., pages 22669–22679, 2023

  3. [3]

    Baranchuk, A

    D. Baranchuk, A. V oynov, I. Rubachev, V . Khrulkov, and A. Babenko. Label-efficient semantic segmentation with diffusion models. InProc. Int. Conf. Learn. Represent., 2022

  4. [4]

    Bolya, C.-Y

    D. Bolya, C.-Y . Fu, X. Dai, P. Zhang, C. Feichtenhofer, and J. Hoffman. Token merging: Your ViT but faster. InProc. Int. Conf. Learn. Represent., 2022

  5. [5]

    Bolya and J

    D. Bolya and J. Hoffman. Token merging for fast stable diffusion. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog., pages 4599–4603, 2023

  6. [6]

    J. Chen, C. Ge, E. Xie, Y . Wu, L. Yao, X. Ren, Z. Wang, P. Luo, H. Lu, and Z. Li. PixArt-σ: Weak-to-strong training of diffusion transformer for 4K text-to-image generation. InProc. Eur. Conf. Comput. Vis., pages 74–91. Springer, 2024

  7. [7]

    X. Chen, Z. Liu, S. Xie, and K. He. Deconstructing denoising diffusion models for self-supervised learning. InProc. Int. Conf. Learn. Represent., 2025

  8. [8]

    J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. Imagenet: A large-scale hierarchical image database. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog., pages 248–255, 2009

Show all 62 references
  1. [9]

    Dhariwal and A

    P. Dhariwal and A. Nichol. Diffusion models beat GANs on image synthesis. InProc. Adv. Neural Inf. Process. Syst., pages 8780–8794, 2021

  2. [10]

    X. Ding, X. Zhang, N. Ma, J. Han, G. Ding, and J. Sun. Repvgg: Making vgg-style convnets great again. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recog., pages 13733–13742, 2021

  3. [11]

    Dosovitskiy, L

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. InProc. Int. Conf. Learn. Represent., 2021

  4. [12]

    Esser, S

    P. Esser, S. Kulal, A. Blattmann, R. Entezari, J. Müller, H. Saini, Y . Levi, D. Lorenz, A. Sauer, F. Boesel, et al. Scaling rectified flow transformers for high-resolution image synthesis. InProc. Int. Conf. Mach. Learn., 2024

  5. [13]

    K. He, X. Chen, S. Xie, Y . Li, P. Dollár, and R. Girshick. Masked autoencoders are scalable vision learners. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog., pages 16000–16009, 2022

  6. [14]

    Heusel, H

    M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter. GANs trained by a two time-scale update rule converge to a local nash equilibrium. InProc. Adv. Neural Inf. Process. Syst., 2017

  7. [15]

    J. Ho, A. Jain, and P. Abbeel. Denoising diffusion probabilistic models. InProc. Adv. Neural Inf. Process. Syst., pages 6840–6851, 2020

  8. [16]

    J. Ho, C. Saharia, W. Chan, D. J. Fleet, M. Norouzi, and T. Salimans. Cascaded diffusion models for high fidelity image generation.J. Mach. Learn. Res., 23(47):1–33, 2022

  9. [17]

    Ho and T

    J. Ho and T. Salimans. Classifier-free diffusion guidance. InProc. NeurIPS Workshop DGMs Appl., 2021

  10. [18]

    Hoogeboom, J

    E. Hoogeboom, J. Heek, and T. Salimans. Simple diffusion: End-to-end diffusion for high resolution images. InProc. Int. Conf. Mach. Learn., pages 13213–13232, 2023

  11. [19]

    Jiang, M

    D. Jiang, M. Wang, L. Li, L. Zhang, H. Wang, W. Wei, G. Dai, Y . Zhang, and J. Wang. No other representation component is needed: Diffusion transformers can provide representation guidance by themselves. InProc. Int. Conf. Learn. Represent., 2026

  12. [20]

    Karras, M

    T. Karras, M. Aittala, T. Aila, and S. Laine. Elucidating the design space of diffusion-based generative models. InProc. Adv. Neural Inf. Process. Syst., pages 26565–26577, 2022

  13. [21]

    Kinga, J

    D. Kinga, J. B. Adam, et al. A method for stochastic optimization. InProc. Int. Conf. Learn. Represent., 2015

  14. [22]

    Kingma and R

    D. Kingma and R. Gao. Understanding diffusion objectives as the elbo with simple data augmentation. In Proc. Adv. Neural Inf. Process. Syst., pages 65484–65516, 2023. 11

  15. [23]

    Z. Kong, P. Dong, X. Ma, X. Meng, M. Sun, W. Niu, X. Shen, G. Yuan, B. Ren, M. Qin, et al. SPViT: enabling faster vision transformers via soft token pruning. InProc. Eur. Conf. Comput. Vis., 2022

  16. [24]

    Krause, T

    F. Krause, T. Phan, M. Gui, S. A. Baumann, V . T. Hu, and B. Ommer. TREAD: Token routing for efficient architecture-agnostic diffusion training, 2025. arXiv:2501.04765

  17. [25]

    Kynkäänniemi, M

    T. Kynkäänniemi, M. Aittala, T. Karras, S. Laine, T. Aila, and J. Lehtinen. Applying guidance in a limited interval improves sample and distribution quality in diffusion models. InProc. Adv. Neural Inf. Process. Syst., 2024

  18. [26]

    X. Leng, J. Singh, Y . Hou, Z. Xing, S. Xie, and L. Zheng. REPA-E: Unlocking V AE for end-to-end tuning with latent diffusion transformers, 2025

  19. [27]

    T. Li, H. Chang, S. Mishra, H. Zhang, D. Katabi, and D. Krishnan. MAGE: Masked generative encoder to unify representation learning and image synthesis. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog., pages 2142–2152, 2023

  20. [28]

    Li and K

    T. Li and K. He. Back to basics: Let denoising generative models denoise.arXiv preprint arXiv:2511.13720, 2025

  21. [29]

    Liang, C

    Y . Liang, C. Ge, Z. Tong, Y . Song, J. Wang, and P. Xie. Not all patches are what you need: Expediting vision transformers via token reorganizations. InProc. Int. Conf. Learn. Represent., 2022

  22. [30]

    T.-Y . Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick. Microsoft COCO: Common objects in context. InProc. Eur. Conf. Comput. Vis., pages 740–755, 2014

  23. [31]

    Lipman, R

    Y . Lipman, R. T. Q. Chen, H. Ben-Hamu, M. Nickel, and M. Le. Flow matching for generative modeling. In Proc. Int. Conf. Learn. Represent., 2023

  24. [32]

    Litjens, T

    G. Litjens, T. Kooi, B. E. Bejnordi, A. A. A. Setio, F. Ciompi, M. Ghafoorian, J. A. Van Der Laak, B. Van Ginneken, and C. I. Sánchez. A survey on deep learning in medical image analysis.Med. Image Anal., 42:60–88, 2017

  25. [33]

    Z. Liu, Y . Lin, Y . Cao, H. Hu, Y . Wei, Z. Zhang, S. Lin, and B. Guo. Swin transformer: Hierarchical vision transformer using shifted windows. InProc. IEEE/CVF Int. Conf. Comput. Vis., pages 10012–10022, 2021

  26. [34]

    Loshchilov and F

    I. Loshchilov and F. Hutter. Decoupled weight decay regularization. InProc. Int. Conf. Learn. Represent., 2019

  27. [35]

    N. Ma, M. Goldstein, M. S. Albergo, N. M. Boffi, E. Vanden-Eijnden, and S. Xie. SiT: Exploring flow and diffusion-based generative models with scalable interpolant transformers. InProc. Eur. Conf. Comput. Vis., pages 23–40, 2024

  28. [36]

    Marin, J.-H

    D. Marin, J.-H. R. Chang, A. Ranjan, A. Prabhu, M. Rastegari, and O. Tuzel. Token pooling in vision transformers.arXiv preprint arXiv:2110.03860, 2021

  29. [37]

    L. Meng, H. Li, B.-C. Chen, S. Lan, Z. Wu, Y .-G. Jiang, and S.-N. Lim. AdaViT: Adaptive vision transformers for efficient image recognition. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog., pages 12309–12318, 2022

  30. [38]

    Oquab, T

    M. Oquab, T. Darcet, T. Moutakanni, H. V . V o, M. Szafraniec, V . Khalidov, P. Fernandez, D. HAZIZA, F. Massa, A. El-Nouby, M. Assran, N. Ballas, W. Galuba, R. Howes, P.-Y . Huang, S.-W. Li, I. Misra, M. Rabbat, V . Sharma, G. Synnaeve, H. Xu, H. Jegou, J. Mairal, P. Labatut,...

  31. [39]

    Peebles and S

    W. Peebles and S. Xie. Scalable diffusion models with transformers. InProc. IEEE/CVF Int. Conf. Comput. Vis., pages 4195–4205, 2023

  32. [40]

    Y . Rao, W. Zhao, B. Liu, J. Lu, J. Zhou, and C.-J. Hsieh. DynamicViT: Efficient vision transformers with dynamic token sparsification. InProc. Adv. Neural Inf. Process. Syst., volume 34, pages 13937–13949, 2021

  33. [41]

    Rombach, A

    R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer. High-resolution image synthesis with latent diffusion models. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog., pages 10684–10695, 2022

  34. [42]

    Ronneberger, P

    O. Ronneberger, P. Fischer, and T. Brox. U-net: Convolutional networks for biomedical image segmentation. InProc. Int. Conf. Med. Image Comput. Comput.-Assist. Intervent., pages 234–241, 2015

  35. [43]

    M. Ryoo, A. Piergiovanni, A. Arnab, M. Dehghani, and A. Angelova. Tokenlearner: Adaptive space-time tokenization for videos. InProc. Adv. Neural Inf. Process. Syst., volume 34, pages 12786–12797, 2021. 12

  36. [44]

    Salimans, I

    T. Salimans, I. Goodfellow, W. Zaremba, V . Cheung, A. Radford, and X. Chen. Improved techniques for training GANs. InProc. Adv. Neural Inf. Process. Syst., 2016

  37. [45]

    N. Shazeer. Glu variants improve transformer.arXiv preprint arXiv:2002.05202, 2020

  38. [46]

    J. Su, M. Ahmed, Y . Lu, S. Pan, W. Bo, and Y . Liu. Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024

  39. [47]

    Y . Tian, H. Chen, M. Zheng, Y . Liang, C. Xu, and Y . Wang. U-REPA: Aligning Diffusion U-Nets to ViTs. In Proc. Adv. Neural Inf. Process. Syst., 2026

  40. [48]

    Y . Tian, Z. Tu, H. Chen, J. Hu, C. Xu, and Y . Wang. U-DiTs: Downsample tokens in u-shaped diffusion transformers. InProc. Adv. Neural Inf. Process. Syst., volume 37, pages 51994–52013, 2024

  41. [49]

    Vaswani, N

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin. Attention is all you need. InProc. Adv. Neural Inf. Process. Syst., pages 6000–6010, 2017

  42. [50]

    J. Wang, N. Kang, L. Yao, M. Chen, C. Wu, S. Zhang, S. Xue, Y . Liu, T. Wu, X. Liu, et al. LiT: Delving into a simple linear diffusion transformer for image generation. InICCV, pages 16068–16078, 2025

  43. [51]

    Wang and K

    R. Wang and K. He. Diffuse and disperse: Image generation with representation regularization.arXiv preprint arXiv:2506.09027, 2025

  44. [52]

    S. Wang, Z. Tian, W. Huang, and L. Wang. DDT: Decoupled diffusion transformer.Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recog., 2026

  45. [53]

    Z. Wang, X. Cun, J. Bao, W. Zhou, J. Liu, and H. Li. Uformer: A general u-shaped transformer for image restoration. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog., pages 17683–17693, 2022

  46. [54]

    Xiang, H

    W. Xiang, H. Yang, D. Huang, and Y . Wang. Denoising diffusion autoencoders are unified self-supervised learners. InProc. IEEE/CVF Int. Conf. Comput. Vis., pages 15802–15812, 2023

  47. [55]

    E. Xie, J. Chen, J. Chen, H. Cai, H. Tang, Y . Lin, Z. Zhang, M. Li, L. Zhu, Y . Lu, et al. Sana: Efficient high-resolution image synthesis with linear diffusion transformers.arXiv preprint arXiv:2410.10629, 2024

  48. [56]

    J. Yao, C. Wang, W. Liu, and X. Wang. FasterDiT: Towards faster diffusion transformers training without architecture modification. 37:56166–56189, 2024

  49. [57]

    J. Yao, B. Yang, and X. Wang. Reconstruction vs. generation: Taming optimization dilemma in latent diffusion models. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog., pages 15703–15712, 2025

  50. [58]

    H. Yin, A. Vahdat, J. M. Alvarez, A. Mallya, J. Kautz, and P. Molchanov. A-ViT: Adaptive tokens for efficient vision transformer. InProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog., pages 10809–10818, 2022

  51. [59]

    S. Yu, S. Kwak, H. Jang, J. Jeong, J. Huang, J. Shin, and S. Xie. Representation alignment for generation: Training diffusion transformers is easier than you think. InProc. Int. Conf. Learn. Represent., 2025

  52. [60]

    J. Yun, Y . U. Alçalar, and M. Akçakaya. No alignment needed for generation: Learning linearly separable representations in diffusion models, 2025. arXiv:2509.21565

  53. [61]

    Zheng, N

    B. Zheng, N. Ma, S. Tong, and S. Xie. Diffusion transformers with representation autoencoders. InProc. Int. Conf. Learn. Represent., 2026

  54. [62]

    Zheng, W

    H. Zheng, W. Nie, A. Vahdat, and A. Anandkumar. Fast training of diffusion models with masked transformers. InTrans. Mach. Learn. Res., 2024. 13 Appendix A Implementation Details A.1 Model Configurations All trainings were conducted from scratch, using 4 NVIDIA A100 GPUs for t...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.