REVIEW 3 major objections 4 minor 131 references
This survey establishes pixel-space diffusion transformers as a distinct, scalable class of end-to-end generative models that removes the fixed VAE bottleneck by denoising directly in raw pixel space.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 17:32 UTC pith:54UJ2QT7
load-bearing objection A useful survey with a solid taxonomy, but the 'first systematic' claim is unaudited and needs a methodology section. the 3 major comments →
Pixel-Space Diffusion Transformers
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that pDiTs are not merely a return to early pixel-space diffusion but an independent, scalable paradigm: 'pixel space' means the diffusion state, prediction target, and supervision are defined in the raw image domain, while patchification is only a computational tokenization strategy, not a frozen visual compressor. It asserts that this removes an irreversible information bottleneck, allows single-stage end-to-end optimization, and provides a shared token space in which text, pixels, and task conditions can be jointly modeled by one Transformer. The survey substantiates this through a seven-category architectural taxonomy, an analysis of DDPM and flow-matching fo
What carries the argument
The load-bearing object is the pDiT formulation: a noisy image is patchified, each patch is mapped by a learnable embedding into tokens, a Transformer backbone models global dependencies under time and text conditions, and an end-to-end decoding head predicts the clean image in pixel space. The survey's analytic grid is a seven-category taxonomy — single-stream large-patch, hierarchical/hourglass, global-local decoupled, frequency-decoupled, implicit neural field decoding, cross-scale semantic anchoring, and shared-token unified multimodal architectures — with flow matching and DDPM as the continuous trajectory mathematics and the O(N^2 d) attention cost as the central scaling constraint.
Load-bearing premise
The survey's usefulness rests on the assumption that its selection of representative methods and its seven-category taxonomy faithfully cover the pDiT landscape; if important methods are omitted or miscategorized, the systematic and first-survey claims lose force.
What would settle it
A concrete check: compile an independent list of pixel-space diffusion transformer papers with stated noise schedules and loss weightings, then test whether each fits one of the seven Table I categories; any substantial class that does not fit, or a matched-compute benchmark where latent diffusion beats pixel-space models specifically on text rendering and edge fidelity, would undercut the central taxonomy and the claim that the VAE bottleneck is removed.
If this is right
- Fixed VAE or vision-foundation tokenizers become an optional interface rather than a quality ceiling, since raw-pixel supervision can preserve textures, edges, and text that compression discards.
- End-to-end training can jointly optimize representation and generation, reducing the reconstruction-generation mismatch inherent in two-stage latent diffusion.
- Unified multimodal models can be built on a shared token space where pixels, text, and task conditions are processed by a single Transformer.
- Structured computation — large patches, hierarchy, global-local decoupling, frequency separation, and implicit decoding — makes high-resolution pixel-space generation tractable.
- Evaluation must move beyond a single FID score or latency figure to include text fidelity, structural preservation, edit locality, and compute-fidelity Pareto comparisons.
Where Pith is reading between the lines
- If the pDiT advantages hold, the most promising architectures will likely combine several taxonomy mechanisms — for example, global-local decoupling with frequency-aware scheduling and dynamic token granularity — rather than rely on a single design.
- The learnability-fidelity conflict suggests a testable scaling prediction: at sufficient compute and with clean-image prediction or structured supervision, pixel-space models should match or surpass latent models on fine-detail metrics under matched training FLOPs.
- The observed risk that token-wise representation alignment degrades diversity in pixel space implies that future guidance will shift toward relation-level or stage-adaptive constraints; a direct comparison of Gram-matrix alignment versus token-wise alignment on diversity and fidelity would test this.
- If the shared-token unified paradigm is realized, multi-task gradient interference would need active management; monitoring gradient cosine-similarity conflicts during training of such a unified model could predict when understanding and generation objectives fight each other.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a survey of Pixel-Space Diffusion Transformers (pDiTs), arguing that this emerging class of generative models—defined by diffusion and supervision in the raw image space, with Transformer backbones and no fixed VAE/VQ tokenizer—offers a distinct and scalable alternative to latent diffusion. The survey covers theoretical foundations (DDPM, flow matching, complexity analysis), a seven-category architectural taxonomy (Table I), unified multimodal modeling, applications (image, video, 3D, medical), discussion, and future challenges. The central claim is that this is 'the first survey dedicated to the systematic review and comprehensive analysis of pixel-space diffusion models' (Sec. I).
Significance. If the taxonomy and coverage are reliable, the paper would serve as a useful roadmap for a rapidly evolving area. The background mathematics (Eqs. 1–11) is standard and correctly reproduced, and the paper clearly organizes the field along three dimensions: architecture, continuous generative mechanisms, and unified multimodal modeling. It also identifies substantive open problems, such as the computation–fidelity trade-off, non-invasive semantic guidance, forgetting in unified models, and native pixel-space trajectory design. The survey is balanced in acknowledging the continued strengths of latent diffusion. However, the value of the survey as a 'systematic review' depends on the completeness and fidelity of its method selection and categorization, which are not auditable in the current manuscript.
major comments (3)
- [Sec. I and Sec. III-B (Table I)] The paper's central claim of being 'the first survey dedicated to the systematic review and comprehensive analysis of pixel-space diffusion models' is not supported by any stated methodology. The manuscript does not provide a search protocol, inclusion/exclusion criteria, database sources, or a time window, so the 'systematic' claim is unverifiable. Table I lists only one or two representative methods per category, but there is no justification for why these methods were chosen or why the seven categories are mutually exclusive and collectively exhaustive. Please add a methodology subsection describing the literature search and selection process, or soften the 'systematic/first' claim accordingly.
- [Sec. III and Sec. V-A] There is an internal inconsistency in the definition of pDiT. Section III defines pDiT as a Transformer-based architecture: 'the Transformer uses tokenized representations and attention mechanisms as its primary computational backbone.' However, Section V-A states that 'The earliest applications of pDiT' include Simple Diffusion [67] and SiD2 [57], both of which are U-Net-based pixel-space diffusion models, not Transformers. The same issue appears in Fig. 2 and the timeline. This conflates 'pixel-space diffusion' with 'pixel-space diffusion Transformers.' Please either restrict the application discussion to Transformer-based models or explicitly distinguish between pixel-space diffusion in general and the pDiT subset, and adjust the taxonomy and title accordingly.
- [Sec. V-A, Sec. V-B] The application sections mix pDiT-specific claims with general pixel-space diffusion results without sufficient discrimination. For example, BlazeEdit [123] is described as 'still following the latent diffusion paradigm rather than pDiT,' which is appropriately flagged, but other methods are listed as pDiT without checking whether they use a Transformer backbone. The survey would be strengthened by a consistent criterion for which methods enter the pDiT taxonomy and which are included only as adjacent pixel-space approaches. This is necessary to make the seven-category taxonomy in Table I a faithful representation of the pDiT landscape.
minor comments (4)
- [Fig. 2] Typo: 'PixWAorld' should be 'PixWorld'.
- [References] Reference [5] has an incomplete title: 'in computer vision' appears where a topic phrase is expected.
- [Eq. (33)] The equation formatting has stray line-break characters (\r\r) inside the alignment loss expression; please clean up the LaTeX.
- [Sec. II-B] The figure overview (Fig. 1) mentions 'Time Parameterization' as a background topic, but the background section does not discuss time parameterizations explicitly beyond DDPM/flow matching. Consider adding a short paragraph to align the overview with the content.
Circularity Check
No significant circularity: the survey's taxonomy and claims are literature-based and do not reduce to the paper's own inputs.
full rationale
This paper is a survey, not a derivation. The central claims are (i) that pixel-space diffusion transformers are a distinct scalable class of generative models and (ii) that this is the first systematic survey of that area. Neither claim is derived from an equation or fitted parameter within the paper. The theoretical background (DDPM, flow matching, attention complexity) consists of standard, externally established results presented for context, not used to predict a quantity that was already assumed. The Table I taxonomy and the three organizing dimensions are interpretative classifications of external methods; they are not presented as predictions and no method is scored or fitted against the taxonomy to produce a result. The self-citations (refs. [1], [10], [26], [37], [122]) appear only as general contextual references in the introduction and the privacy discussion; none of them is invoked to justify the taxonomy, the 'first survey' novelty claim, or any technical conclusion. The absence of stated inclusion/exclusion criteria for the surveyed methods is a legitimate audit concern about completeness and representativeness, but it is not a circularity: the survey's organizational structure is not equivalent to its own literature sample by construction. No step in the paper reduces, by definition or by self-citation, to its own inputs. Accordingly, the honest finding is no significant circularity.
Axiom & Free-Parameter Ledger
axioms (3)
- domain assumption Cited primary works (Simple Diffusion, JiT, PixelDiT, HDiT, HiDream-O1-Image, etc.) are correctly described and their results are real.
- standard math The standard DDPM and flow-matching formulations (Eqs. 1-11) are standard mathematics as used in the literature.
- domain assumption The absence of an earlier dedicated survey of pixel-space diffusion models (the 'first survey' claim) is true.
read the original abstract
Latent diffusion models (LDMs) enable efficient high-resolution image synthesis by denoising in a VAE-compressed latent space. However, fixed visual tokenizers can discard fine textures and structural details, while separate representation and diffusion training creates a mismatch between reconstruction and generation objectives. These limitations have renewed interest in pixel-space diffusion, which models raw pixels directly, removes the VAE bottleneck, and supports end-to-end optimization. This formulation better matches the demands of high-fidelity generation but introduces challenges in high-dimensional modeling, including noise scheduling, loss weighting, token efficiency, and scalable architecture design. Pixel-space modeling also offers a promising basis for unified multimodal systems: raw pixels, text, and task conditions can be represented in a shared token space and jointly processed by a single Transformer, narrowing the gap between visual understanding and generation. This paper reviews Pixel-Space Diffusion Transformers (pDiTs) from the perspectives of model architecture, continuous generative mechanisms, and unified multimodal modeling. We summarize representative methods, identify key technical challenges, and discuss future directions toward high-fidelity, end-to-end vision foundation models that integrate generation and understanding.
Figures
Reference graph
Works this paper leans on
-
[1]
A sanity check for multi-in-domain face forgery detection in the real world,
J. Cheng, R. Yan, Z. Yan, Y . Gan, X. Zhang, Z. Wang, W. Peng, and L. Liang, “A sanity check for multi-in-domain face forgery detection in the real world,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026, pp. 21 306–21 315
2026
-
[2]
Edge deep learning in computer vision and medical diagnostics: a com- prehensive survey,
Y . Xu, T. M. Khan, Y . Song, and E. Meijering, “Edge deep learning in computer vision and medical diagnostics: a com- prehensive survey,”arXiv preprint arXiv:2605.06714, 2026
Pith/arXiv arXiv 2026
-
[3]
A review of pseudo-labeling for computer vision,
P. Kage, J. Rothenberger, P. Andreadis, and D. Diochnos, “A review of pseudo-labeling for computer vision,”Journal of Artificial Intelligence Research, vol. 85, 2026
2026
-
[4]
Intelligent recognition of emergency vehicles in congested traffic using computer vision,
K. P. M. Ballesteros, C. M. L. D. Cruz, J. D. R. Magbanua, M. J. M. Mancenido, and L. V . Comia, “Intelligent recognition of emergency vehicles in congested traffic using computer vision,” in2026 6th International Conference on Image Pro- cessing and Capsule Networks (ICIPCN). IEEE, 2026, pp. 34–40
2026
-
[5]
in computer vision,
Y . Ji, W. Wu, H. Chen, and Z. Liu, “in computer vision,” Artificial Intelligence in Digital Image Processing: Theories, Methods, and Applications, p. 79, 2026
2026
-
[6]
Attention mechanisms in computer vision: A survey,
M.-H. Guo, T.-X. Xu, J.-J. Liu, Z.-N. Liu, P.-T. Jiang, T.- J. Mu, S.-H. Zhang, R. R. Martin, M.-M. Cheng, and S.-M. Hu, “Attention mechanisms in computer vision: A survey,” Computational visual media, vol. 8, no. 3, pp. 331–368, 2022
2022
-
[7]
Generative adversarial networks in computer vision: A survey and taxonomy,
Z. Wang, Q. She, and T. E. Ward, “Generative adversarial networks in computer vision: A survey and taxonomy,”ACM Computing Surveys (CSUR), vol. 54, no. 2, pp. 1–38, 2021
2021
-
[8]
Picture perfect: Engaging customers with visual generative ai,
M. Heitmann, T. P. Jansen, M. Reisenbichler, and D. A. Schweidel, “Picture perfect: Engaging customers with visual generative ai,”Journal of Marketing, vol. 90, no. 4, pp. 74–96, 2026
2026
-
[9]
Around the world in 80 timesteps: A generative approach to global visual geolocation,
N. Dufour, V . Kalogeiton, D. Picard, and L. Landrieu, “Around the world in 80 timesteps: A generative approach to global visual geolocation,” inProceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 23 016–23 026
2025
-
[10]
Entropy-adaptive diffusion policy optimization with dynamic step alignment,
R. Yan, J. Cheng, Y . Gan, S. Sun, Y . Wu, Y . Yang, L. Ling, J. Lin, Y . Zhu, J. Zhouet al., “Entropy-adaptive diffusion policy optimization with dynamic step alignment,” inProceed- ings of the IEEE/CVF International Conference on Computer Vision, 2025, pp. 1924–1934
2025
-
[11]
Deep generative modelling: A comparative review of vaes, gans, normalizing flows, energy-based and autoregressive mod- els,
S. Bond-Taylor, A. Leach, Y . Long, and C. G. Willcocks, “Deep generative modelling: A comparative review of vaes, gans, normalizing flows, energy-based and autoregressive mod- els,”IEEE transactions on pattern analysis and machine intelligence, vol. 44, no. 11, pp. 7327–7347, 2021
2021
-
[12]
Energy-based generative adversarial network,
J. Zhao, M. Mathieu, and Y . LeCun, “Energy-based generative adversarial network,”arXiv preprint arXiv:1609.03126, 2016
Pith/arXiv arXiv 2016
-
[13]
Nice: Non- linear independent components estimation,
L. Dinh, D. Krueger, and Y . Bengio, “Nice: Non- linear independent components estimation,”arXiv preprint arXiv:1410.8516, 2014
Pith/arXiv arXiv 2014
-
[14]
Wavenet: A generative model for raw audio,
A. Van Den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, K. Kavukcuogluet al., “Wavenet: A generative model for raw audio,”arXiv preprint arXiv:1609.03499, vol. 12, no. 1, 2016
Pith/arXiv arXiv 2016
-
[15]
Unpaired image-to-image translation using cycle-consistent adversarial networks,
J.-Y . Zhu, T. Park, P. Isola, and A. A. Efros, “Unpaired image-to-image translation using cycle-consistent adversarial networks,” inProceedings of the IEEE international confer- ence on computer vision, 2017, pp. 2223–2232
2017
-
[16]
Diffusion models in vision: A survey,
F.-A. Croitoru, V . Hondru, R. T. Ionescu, and M. Shah, “Diffusion models in vision: A survey,”IEEE transactions on pattern analysis and machine intelligence, vol. 45, no. 9, pp. 10 850–10 869, 2023
2023
-
[17]
Mixture of global and local experts with diffusion transformer for controllable face generation,
X. Zou, S. Zhang, X. Fu, Y . Li, K. Li, Y . Cao, C. Lang, P. Tao, and J. Xing, “Mixture of global and local experts with diffusion transformer for controllable face generation,”IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026
2026
-
[18]
Remasking 25 discrete diffusion models with inference-time scaling,
G. Wang, Y . Schiff, S. Sahoo, and V . Kuleshov, “Remasking 25 discrete diffusion models with inference-time scaling,”Ad- vances in Neural Information Processing Systems, vol. 38, pp. 147 282–147 339, 2026
2026
-
[19]
Why diffu- sion models don’t memorize: The role of implicit dynamical regularization in training,
T. Bonnaire, R. Urfin, G. Biroli, and M. M ´ezard, “Why diffu- sion models don’t memorize: The role of implicit dynamical regularization in training,”Advances in Neural Information Processing Systems, vol. 38, pp. 141 266–141 286, 2026
2026
-
[20]
Guiding a diffusion model with a bad version of itself,
T. Karras, M. Aittala, T. Kynk ¨a¨anniemi, J. Lehtinen, T. Aila, and S. Laine, “Guiding a diffusion model with a bad version of itself,”Advances in Neural Information Processing Systems, vol. 37, pp. 52 996–53 021, 2024
2024
-
[21]
G. Tevet, S. Raab, B. Gordon, Y . Shafir, D. Cohen-Or, and A. H. Bermano, “Human motion diffusion model,”arXiv preprint arXiv:2209.14916, 2022
Pith/arXiv arXiv 2022
-
[22]
Diffusion models and representation learning: A survey,
M. Fuest, P. Ma, M. Gui, J. Schusterbauer, V . T. Hu, and B. Ommer, “Diffusion models and representation learning: A survey,”IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026
2026
-
[23]
Holod- iffusion: Training a 3d diffusion model using 2d images,
A. Karnewar, A. Vedaldi, D. Novotny, and N. J. Mitra, “Holod- iffusion: Training a 3d diffusion model using 2d images,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 18 423–18 433
2023
-
[24]
Klass: Kl-guided fast inference in masked diffusion models,
S. H. Kim, S. Hong, H. Jung, Y . Park, and S.-Y . Yun, “Klass: Kl-guided fast inference in masked diffusion models,” Advances in Neural Information Processing Systems, vol. 38, pp. 92 267–92 301, 2026
2026
-
[25]
Training-free constrained generation with stable diffusion models,
S. Zampini, J. K. Christopher, L. Oneto, D. Anguita, and F. Fioretto, “Training-free constrained generation with stable diffusion models,”Advances in Neural Information Processing Systems, vol. 38, pp. 27 285–27 316, 2026
2026
-
[26]
Do less, achieve more: Do we need every-step optimization for rl fine-tuning of diffusion models?
R. Yan, J. Cheng, S. Sun, Y . Sun, Y . Wu, W. Peng, Z. Wang, L. Liang, J. Xing, and Y . Cai, “Do less, achieve more: Do we need every-step optimization for rl fine-tuning of diffusion models?” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026, pp. 16 561– 16 571
2026
-
[27]
High-resolution image synthesis with latent diffusion models,
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Om- mer, “High-resolution image synthesis with latent diffusion models,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 10 684– 10 695
2022
-
[28]
Sdxl: Improving latent diffusion models for high-resolution image synthesis,
D. Podell, Z. English, K. Lacey, A. Blattmann, T. Dockhorn, J. M ¨uller, J. Penna, and R. Rombach, “Sdxl: Improving latent diffusion models for high-resolution image synthesis,” inInter- national Conference on Learning Representations, vol. 2024, 2024, pp. 1862–1874
2024
-
[29]
Align your latents: High-resolution video synthesis with latent diffusion models,
A. Blattmann, R. Rombach, H. Ling, T. Dockhorn, S. W. Kim, S. Fidler, and K. Kreis, “Align your latents: High-resolution video synthesis with latent diffusion models,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 22 563–22 575
2023
-
[30]
Stable video diffusion: Scaling latent video diffusion models to large datasets,
A. Blattmann, T. Dockhorn, S. Kulal, D. Mendelevitch, M. Kil- ian, D. Lorenz, Y . Levi, Z. English, V . V oleti, A. Lettset al., “Stable video diffusion: Scaling latent video diffusion models to large datasets,”arXiv preprint arXiv:2311.15127, 2023
Pith/arXiv arXiv 2023
-
[31]
Diffusion models already have a semantic latent space,
M. Kwon, J. Jeong, and Y . Uh, “Diffusion models already have a semantic latent space,”arXiv preprint arXiv:2210.10960, 2022
Pith/arXiv arXiv 2022
-
[32]
Latentpaint: Image inpainting in latent space with diffusion models,
C. Corneanu, R. Gadde, and A. M. Martinez, “Latentpaint: Image inpainting in latent space with diffusion models,” in Proceedings of the IEEE/CVF winter conference on applica- tions of computer vision, 2024, pp. 4334–4343
2024
-
[33]
Unconditional latent diffusion models memorize patient imaging data,
S. U. H. Dar, M. Seyfarth, I. Ayx, T. Papavassiliu, S. O. Schoenberg, R. M. Siepmann, F. C. Laqua, J. Kahmann, N. Frey, B. Baeßleret al., “Unconditional latent diffusion models memorize patient imaging data,”Nature biomedical engineering, vol. 10, no. 3, pp. 458–472, 2026
2026
-
[34]
Ladder variational autoencoders,
C. K. Sønderby, T. Raiko, L. Maaløe, S. K. Sønderby, and O. Winther, “Ladder variational autoencoders,”Advances in neural information processing systems, vol. 29, 2016
2016
-
[35]
Longitudinal variational autoencoder,
S. Ramchandran, G. Tikhonov, K. Kujanp ¨a¨a, M. Koskinen, and H. L¨ahdesm¨aki, “Longitudinal variational autoencoder,” inIn- ternational Conference on Artificial Intelligence and Statistics. PMLR, 2021, pp. 3898–3906
2021
-
[36]
Gram- mar variational autoencoder,
M. J. Kusner, B. Paige, and J. M. Hern ´andez-Lobato, “Gram- mar variational autoencoder,” inInternational conference on machine learning. PMLR, 2017, pp. 1945–1954
2017
-
[37]
The exploration-exploitation dilemma revisited: An entropy perspective,
R. Yan, Y . Gan, Y . Wu, L. Liang, J. Xing, Y . Cai, and R. Huang, “The exploration-exploitation dilemma revisited: An entropy perspective,”arXiv preprint arXiv:2408.09974, 2024
Pith/arXiv arXiv 2024
-
[38]
Mvae: Multimodal variational autoencoder for fake news detection,
D. Khattar, J. S. Goud, M. Gupta, and V . Varma, “Mvae: Multimodal variational autoencoder for fake news detection,” inThe world wide web conference, 2019, pp. 2915–2921
2019
-
[39]
Poisson variational autoencoder,
H. Vafaii, D. Galor, and J. L. Yates, “Poisson variational autoencoder,”Advances in Neural Information Processing Sys- tems, vol. 37, pp. 44 871–44 906, 2024
2024
-
[40]
Xidintfl-vae: Xgboost-based intrusion detection of imbalance network traffic via class-wise focal loss variational autoencoder,
O. H. Abdulganiyu, T. A. Tchakoucht, Y . K. Saheed, and H. A. Ahmed, “Xidintfl-vae: Xgboost-based intrusion detection of imbalance network traffic via class-wise focal loss variational autoencoder,”The Journal of Supercomputing, vol. 81, no. 1, p. 16, 2025
2025
-
[41]
Remaining useful life prediction based on interpretable serialized variational autoencoder: A drift- diffusion stochastic equation perspective,
J. Zhang, K. Chen, R. He, T. Huang, J. Tian, S. Wu, P. Yan, and Y . Cheng, “Remaining useful life prediction based on interpretable serialized variational autoencoder: A drift- diffusion stochastic equation perspective,”IEEE Transactions on Industrial Informatics, 2026
2026
-
[42]
Scalable diffusion models with transformers,
W. Peebles and S. Xie, “Scalable diffusion models with transformers,” inProceedings of the IEEE/CVF international conference on computer vision, 2023, pp. 4195–4205
2023
-
[43]
All are worth words: A vit backbone for diffusion models,
F. Bao, S. Nie, K. Xue, Y . Cao, C. Li, H. Su, and J. Zhu, “All are worth words: A vit backbone for diffusion models,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 22 669–22 679
2023
-
[44]
Sit: Exploring flow and diffusion-based generative models with scalable interpolant transformers,
N. Ma, M. Goldstein, M. S. Albergo, N. M. Boffi, E. Vanden- Eijnden, and S. Xie, “Sit: Exploring flow and diffusion-based generative models with scalable interpolant transformers,” in European Conference on Computer Vision. Springer, 2024, pp. 23–40
2024
-
[45]
Scaling rectified flow transformers for high-resolution image synthesis,
P. Esser, S. Kulal, A. Blattmann, R. Entezari, J. M ¨uller, H. Saini, Y . Levi, D. Lorenz, A. Sauer, F. Boeselet al., “Scaling rectified flow transformers for high-resolution image synthesis,” inForty-first international conference on machine learning, 2024
2024
-
[46]
Exploring diffusion transformer designs via grafting,
K. Chandrasegaran, M. Poli, D. Fu, D. Kim, L. M. Hadzic, M. Li, A. Gupta, S. Massaroli, A. Mirhoseini, J. C. Niebles et al., “Exploring diffusion transformer designs via grafting,” Advances in Neural Information Processing Systems, vol. 38, pp. 17 816–17 847, 2026
2026
-
[47]
Scalable high-resolution pixel- space image synthesis with hourglass diffusion transformers,
K. Crowson, S. A. Baumann, A. Birch, T. M. Abraham, D. Z. Kaplan, and E. Shippole, “Scalable high-resolution pixel- space image synthesis with hourglass diffusion transformers,” inForty-first International Conference on Machine Learning, 2024
2024
-
[48]
Edify image: High- quality image generation with pixel space laplacian diffusion models,
Y . Atzmon, M. Bala, Y . Balaji, T. Cai, Y . Cui, J. Fan, Y . Ge, S. Gururani, J. Huffman, R. Isaacet al., “Edify image: High- quality image generation with pixel space laplacian diffusion models,”arXiv preprint arXiv:2411.07126, 2024
Pith/arXiv arXiv 2024
-
[49]
Pixelgen: Pixel diffusion beats latent diffusion with perceptual loss,
Z. Ma, R. Xu, and S. Zhang, “Pixelgen: Pixel diffusion beats latent diffusion with perceptual loss,”arXiv preprint arXiv:2602.02493, 2026
Pith/arXiv arXiv 2026
-
[50]
Pixnerd: Pixel neural field diffusion,
S. Wang, Z. Gao, C. Zhu, W. Huang, and L. Wang, “Pixnerd: Pixel neural field diffusion,”arXiv preprint arXiv:2507.23268, 2025
Pith/arXiv arXiv 2025
-
[51]
Pixeldit: Pixel diffusion transformers for image generation,
Y . Yu, W. Xiong, W. Nie, Y . Sheng, S. Liu, and J. Luo, “Pixeldit: Pixel diffusion transformers for image generation,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026, pp. 14 273–14 282
2026
-
[52]
Deco: 26 Frequency-decoupled pixel diffusion for end-to-end image generation,
Z. Ma, L. Wei, S. Wang, S. Zhang, and Q. Tian, “Deco: 26 Frequency-decoupled pixel diffusion for end-to-end image generation,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026, pp. 43 600– 43 610
2026
-
[53]
Latent forcing: Reordering the diffusion trajectory for pixel-space image generation,
A. Baade, E. R. Chan, K. Sargent, C. Chen, J. Johnson, E. Adeli, and L. Fei-Fei, “Latent forcing: Reordering the diffusion trajectory for pixel-space image generation,”arXiv preprint arXiv:2602.11401, 2026
arXiv 2026
-
[54]
Your latent mask is wrong: Pixel-equivalent latent compositing for diffusion models,
R. Bradbury and D. Zhong, “Your latent mask is wrong: Pixel-equivalent latent compositing for diffusion models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026, pp. 18 630–18 639
2026
-
[55]
Scale space diffusion,
S. Mukhopadhyay, P. Udhayanan, and A. Shrivastava, “Scale space diffusion,” inProceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition, 2026, pp. 35 851–35 860
2026
-
[56]
Pixel motion diffusion is what we need for robot control,
E.-R. Nguyen, Y . Zhang, K. Ranasinghe, X. Li, and M. S. Ryoo, “Pixel motion diffusion is what we need for robot control,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026, pp. 23 663– 23 672
2026
-
[57]
Simpler diffusion: 1.5 fid on imagenet512 with pixel-space diffusion,
E. Hoogeboom, T. Mensink, J. Heek, K. Lamerigts, R. Gao, and T. Salimans, “Simpler diffusion: 1.5 fid on imagenet512 with pixel-space diffusion,” inProceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 18 062– 18 071
2025
-
[58]
Pixel-perfect depth with semantics-prompted diffusion transformers,
G. Xu, H. Lin, H. Luo, X. Wang, J. Yao, L. Zhu, Y . Pu, C. Chi , H. Sun, B. Wanget al., “Pixel-perfect depth with semantics-prompted diffusion transformers,”Advances in Neu- ral Information Processing Systems, vol. 38, pp. 174 731– 174 755, 2026
2026
-
[59]
Novel view synthesis with pixel-space diffu- sion models,
N. Elata, B. Kawar, Y . Ostrovsky-Berman, M. Farber, and R. Sokolovsky, “Novel view synthesis with pixel-space diffu- sion models,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025, pp. 26 756– 26 766
2025
-
[60]
Pixel-space post-training of latent diffusion models,
C. Zhang, S. Motwani, M. Yu, J. Hou, F. Juefei-Xu, S. Tsai, P. Vajda, Z. He, and J. Wang, “Pixel-space post-training of latent diffusion models,”arXiv preprint arXiv:2409.17565, 2024
Pith/arXiv arXiv 2024
-
[61]
Back to basics: Let denoising generative models denoise,
T. Li and K. He, “Back to basics: Let denoising generative models denoise,” inProceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition, 2026, pp. 36 115–36 125
2026
-
[62]
Q. Cai, J. Chen, C. Gao, Z. Gong, Y . Li, Y . Pan, Y . Peng, Z. Qiu, K. Yu, Y . Zhanget al., “Hidream-o1-image: A natively unified image generative foundation model with pixel-level unified transformer,”arXiv preprint arXiv:2605.11061, 2026
Pith/arXiv arXiv 2026
-
[63]
Show- o: One single transformer to unify multimodal understanding and generation,
J. Xie, W. Mao, Z. Bai, D. J. Zhang, W. Wang, K. Q. Lin, Y . Gu, Z. Chen, Z. Yang, and M. Z. Shou, “Show- o: One single transformer to unify multimodal understanding and generation,” inInternational Conference on Learning Representations, vol. 2025, 2025, pp. 28 240–28 264
2025
-
[64]
Transfusion: Predict the next token and diffuse images with one multi-modal model,
C. Zhou, L. Yu, A. Babu, K. Tirumala, M. Yasunaga, L. Shamis, J. Kahn, X. Ma, L. Zettlemoyer, and O. Levy, “Transfusion: Predict the next token and diffuse images with one multi-modal model,” inInternational Conference on Learning Representations, vol. 2025, 2025, pp. 6446–6469
2025
-
[65]
Learning transferable visual models from natural language supervision,
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clarket al., “Learning transferable visual models from natural language supervision,” inInternational conference on machine learning. PmLR, 2021, pp. 8748–8763
2021
-
[66]
Dinov2: Learning robust visual features without supervision,
M. Oquab, T. Darcet, T. Moutakanni, H. V o, M. Szafraniec, V . Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El- Noubyet al., “Dinov2: Learning robust visual features without supervision,”arXiv preprint arXiv:2304.07193, 2023
Pith/arXiv arXiv 2023
-
[67]
simple diffusion: End-to-end diffusion for high resolution images,
E. Hoogeboom, J. Heek, and T. Salimans, “simple diffusion: End-to-end diffusion for high resolution images,” inInterna- tional Conference on Machine Learning. PMLR, 2023, pp. 13 213–13 232
2023
-
[68]
Low light image enhancement challenge at ntire 2026,
G. Ciubotariu, A. Rehman, F. A. Dharejo, R. A. Naqvi, M. V . Conde, R. Timofte, Z. Jin, H. Wu, W. Zhang, C. Yeet al., “Low light image enhancement challenge at ntire 2026,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026, pp. 1741–1753
2026
-
[69]
Hyperspec- tral imaging,
D. Hong, C. Li, N. Yokoya, B. Zhang, X. Jia, A. Plaza, P. Gamba, J. A. Benediktsson, and J. Chanussot, “Hyperspec- tral imaging,”Nature Reviews Methods Primers, vol. 6, no. 1, p. 19, 2026
2026
-
[70]
Perception encoder: The best visual embeddings are not at the output of the network,
D. Bolya, P.-Y . Huang, P. Sun, J. H. Cho, A. Madotto, C. Wei, T. Ma, J. Zhi, J. Rajasegaran, H. Bangalathet al., “Perception encoder: The best visual embeddings are not at the output of the network,”Advances in Neural Information Processing Systems, vol. 38, pp. 60 884–60 937, 2026
2026
-
[71]
Ntire 2026 challenge on video saliency prediction: Methods and results,
A. Moskalenko, A. Bryncev, I. Kosmynin, K. Shilovskaya, M. Erofeev, D. Vatolin, R. Timofte, K. Wang, Y . Hu, Z. Li et al., “Ntire 2026 challenge on video saliency prediction: Methods and results,” inProceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, 2026, pp. 2193–2206
2026
-
[72]
Ntire 2026 challenge on robust ai-generated image detection in the wild,
A. Gushchin, K. Abud, E. Shumitskaya, A. Filippov, G. By- chkov, S. Lavrushkin, M. Erofeev, A. Antsiferova, C. Chen, S. Tanet al., “Ntire 2026 challenge on robust ai-generated image detection in the wild,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026, pp. 1895–1913
2026
-
[73]
One step diffusion via shortcut models,
K. Frans, D. Hafner, S. Levine, and P. Abbeel, “One step diffusion via shortcut models,” inInternational Conference on Learning Representations, vol. 2025, 2025, pp. 34 668–34 684
2025
-
[74]
Latent diffusion for language generation,
J. Lovelace, V . Kishore, C. Wan, E. Shekhtman, and K. Q. Weinberger, “Latent diffusion for language generation,”Ad- vances in Neural Information Processing Systems, vol. 36, pp. 56 998–57 025, 2023
2023
-
[75]
Auto-encoding variational bayes,
D. P. Kingma and M. Welling, “Auto-encoding variational bayes,”arXiv preprint arXiv:1312.6114, 2013
Pith/arXiv arXiv 2013
-
[76]
Turbo-vaed: Fast and stable transfer of video-vaes to mobile devices,
Y . Zou, J. Yao, S. Yu, S. Zhang, W. Liu, and X. Wang, “Turbo-vaed: Fast and stable transfer of video-vaes to mobile devices,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 40, no. 16, 2026, pp. 14 086–14 094
2026
-
[77]
Sparc3d: Sparse representation and construction for high-resolution 3d shapes modeling,
Z. Li, Y . Wang, H. Zheng, Y . Luo, and B. Wen, “Sparc3d: Sparse representation and construction for high-resolution 3d shapes modeling,”Advances in Neural Information Processing Systems, vol. 38, pp. 118 582–118 600, 2026
2026
-
[78]
Vision foundation models can be good tokenizers for latent diffusion models,
T. Bi, X. Zhang, Y . Lu, and N. Zheng, “Vision foundation models can be good tokenizers for latent diffusion models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026, pp. 43 310–43 319
2026
-
[79]
Softvq-vae: Efficient 1- dimensional continuous tokenizer,
H. Chen, Z. Wang, X. Li, X. Sun, F. Chen, J. Liu, J. Wang, B. Raj, Z. Liu, and E. Barsoum, “Softvq-vae: Efficient 1- dimensional continuous tokenizer,” inProceedings of the Com- puter Vision and Pattern Recognition Conference, 2025, pp. 28 358–28 370
2025
-
[80]
Mgvq: Could vq-vae beat vae? a generalizable tokenizer with multi-group quantization,
M. Jia, W. Yin, X. Hu, J. Guo, X. Guo, Q. Zhang, X.- X. Long, and P. Tan, “Mgvq: Could vq-vae beat vae? a generalizable tokenizer with multi-group quantization,”arXiv preprint arXiv:2507.07997, 2025
Pith/arXiv arXiv 2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.