REVIEW 3 major objections 1 minor 1 cited by
Diffusion Transformers benefit from register tokens even without ViT-style high-norm outliers, with larger gains in pixel space than latent space.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-12 16:42 UTC pith:674KEV4R
load-bearing objection We only have the abstract for the DiT/registers paper; the cached full text is an unrelated quant-ph manuscript, so none of the empirical claims can be checked. the 3 major comments →
Registers Matter for Pixel-Space Diffusion Transformers
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
DiTs lack the high-norm patch-token outliers characteristic of ViTs, yet they still benefit from the addition of register tokens. The benefit is stronger in pixel-space DiTs than in latent-space DiTs; registers yield cleaner intermediate feature maps at high noise levels, and recent high-performing pixel-space architectures already embed register-like mechanisms that may help explain their results.
What carries the argument
Register tokens (and the proposed Register Guidance method that amplifies their contribution). They clean high-noise intermediate feature maps and improve visual structure and coherence even though they are not needed to absorb outliers.
Load-bearing premise
The paper treats the cleaner high-noise feature maps produced by registers as a causal reason for better pixel-space generation, and treats the implicit register-like parts of recent architectures as sufficiently analogous to explain their performance, without a direct causal isolation of either link.
What would settle it
Train capacity-matched pixel-space DiTs with and without registers (or with the implicit register-like components ablated) and check whether the high-noise feature-map cleanliness gap appears or disappears exactly when generation quality improves or drops; if quality gains persist without cleaner maps, or maps clean without quality gains, the claimed mechanism fails.
If this is right
- Pixel-space DiT designs should include explicit registers or equivalent sink tokens rather than relying only on latent-space practice.
- Register Guidance can be applied at inference or fine-tuning time to boost structure and coherence by amplifying the register pathway.
- Implicit register-like components already present in recent pixel-space architectures partially account for their strong empirical results and should be preserved or strengthened.
- Inspection of intermediate feature maps across noise levels becomes a useful diagnostic for why certain DiT variants generate more coherent images.
Where Pith is reading between the lines
- Because registers help DiTs without absorbing outliers, their role in generative transformers is likely different from their role in discriminative ViTs and may be tied to the denoising trajectory itself.
- If high-noise map cleanliness is the operative factor, similar register or sink-token benefits should appear in other high-noise generative models, not only DiTs.
- Architectures could move beyond borrowed ViT registers toward tokens whose capacity or attention pattern is explicitly scheduled with the noise level.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The submission is presented as arXiv:2605.16147, a cs.CV paper claiming that Diffusion Transformers lack ViT-style high-norm patch-token outliers yet still benefit from register tokens (more so in pixel space than latent space), that registers yield cleaner high-noise intermediate feature maps, that recent pixel-space DiTs contain implicit register-like mechanisms, and that a proposed Register Guidance method improves structure and coherence. The supplied full manuscript body, however, is an unrelated quant-ph paper by Fabio Siringo on macroscopic superpositions, a stochastic reduction of the Schrödinger equation, and a dynamical derivation of the Born rule. No DiT experiments, ablations, feature-map analyses, or Register Guidance results appear in the body.
Significance. If the abstract’s claims were supported by a matching manuscript, the work would be of clear interest to the generative-modeling community: it would separate register benefits from the classic ViT outlier mechanism, give a pixel- vs latent-space comparison, and offer a practical guidance method. As submitted, that significance cannot be assessed. The body that is present is a conservative, unitary-dynamics account of macroscopic collapse; that is a different contribution in a different field and is not the paper under review.
major comments (3)
- Title/abstract vs. full text: the body is Siringo’s quant-ph manuscript (macroscopic ensembles, white-noise matrix elements W_nm, Itô reduction, Born rule), not a DiT/registers paper. Every load-bearing empirical claim in the abstract—no ViT-style outliers in DiTs, larger pixel-space gains, cleaner high-noise maps, implicit register-like mechanisms, Register Guidance efficacy—has no methods, figures, tables, or numbers that can be audited. The submission is not reviewable as the claimed work.
- Causal claim in the abstract (“registers produce cleaner feature maps at high noise levels, which may contribute to their effectiveness in pixel-space generation”) cannot be checked: there are no intermediate-representation analyses, noise-level sweeps, or ablations linking map cleanliness to sample quality. The interpretive step from correlation to contribution is therefore unsupported in the provided materials.
- The assertion that recent pixel-space DiT architectures “implicitly incorporate register-like mechanisms” is likewise unsubstantiated: no architectural mapping, token-role analysis, or controlled removal/insertion study is present in the body.
minor comments (1)
- Even as a standalone quant-ph text, the cached body has presentation issues (e.g., “must bepostu-lated”, “areducingequation”, “EQUA TION”, “STRA TONOVICH”) that would need copy-editing if that paper were under review; they are secondary to the identity mismatch.
Circularity Check
Mild definitional recovery of the Born rule from conserved p_n; the dynamical reduction chain from Schrödinger is not circular by construction.
specific steps
-
self definitional
[Sec. IV, after Eqs. (30)–(32); also setup Eqs. (2)–(5)]
"Then, the Born rule is recovered because, according to Eq.(12), in the Itô interpretation, the average of Eq.(9) gives d⟨p_m(t)⟩/dt=0 (there is no drift term) and the probability P_m =⟨p_m(t)⟩ is equal to the initial value p_m(0), as predicted by the Born rule."
p_n are defined from the start as the summed squared amplitudes over microstates in each macrostate (Eq. 2), i.e. the Born-rule quantities for that decomposition. Under the Itô reading the SDE has no drift, so ensemble averages of those same p_n are conserved; after reduction, outcome frequencies equal those averages. Matching the Born rule is therefore largely the identification of conserved initial p_n with collapse probabilities, not an independent derivation of a new probability law. The non-circular content is the reduction dynamics itself.
-
self definitional
[Sec. VII (Discussion), condition Eqs. (57)/(60)]
"the reduction condition ΔEΔX≫ℏc would set a limit to the maximum bin extension which can survive in a superposition, thus specifying what we actually mean by a forbidden macroscopic superposition."
The same inequality required for the Itô/reducing regime is used to define the class of superpositions that count as macroscopic (and therefore must collapse). Scope of the conclusion is partly fixed by the condition that makes the conclusion true, which is a mild definitional loop around applicability rather than a forced fit of a numerical prediction.
full rationale
The provided full text is Siringo’s quant-ph manuscript on macroscopic superpositions (not the DiT/registers paper named in the header). Auditing that derivation: the paper decouples macro/micro amplitudes, obtains an exact equation for p_n from the Schrödinger equation, approximates the off-diagonal matrix element as white noise from a large energy bandwidth, and argues the Itô (reducing) interpretation from causality (τ_R ≫ τ_c). Those steps cite external results (Pearle’s reducing SDEs, Kupferman et al. on Itô vs Stratonovich) and do not reduce to fitted parameters or author-only uniqueness theorems. The only mild circularity is identification: p_n are defined as the macroscopic squared amplitudes, the Itô reading has no drift so ⟨p_n⟩ is conserved, and collapse frequencies are then set equal to those conserved averages—recovering the Born rule largely by that identification rather than by an independent probability postulate. A secondary soft spot is using the reduction condition ΔEΔX ≫ ℏc to help specify what counts as a forbidden macroscopic superposition. Neither forces the dynamical claim by pure definition. Score 2.
Axiom & Free-Parameter Ledger
axioms (4)
- domain assumption ViT high-norm patch-token outliers degrade feature maps and are mitigated by register tokens (prior literature).
- domain assumption Pixel-space DiTs are sufficiently ViT-like that token-level interventions (registers) remain meaningful.
- ad hoc to paper Cleaner intermediate feature maps at high noise levels contribute to better pixel-space generation quality.
- ad hoc to paper Some recent pixel-space DiT components function as implicit register-like mechanisms.
invented entities (1)
-
Register Guidance
no independent evidence
read the original abstract
Vision Transformers (ViTs) are known to exhibit high-norm patch-token outliers that degrade feature map quality, a problem effectively mitigated by register tokens. As diffusion models increasingly adopt transformer architectures and move toward pixel-space training, they become closer in form to ViTs, raising the question of whether register tokens are also useful for Diffusion Transformers (DiTs). In this work, we show that DiTs differ from ViTs in a key respect: they do not exhibit patch-token outliers but still benefit from registers. Interestingly, registers are more effective in pixel-space DiTs than in latent-space DiTs. By analyzing intermediate representations, we find that register tokens produce cleaner feature maps at high noise levels, which may contribute to their effectiveness in pixel-space generation. We further observe that recent pixel-space DiT architectures implicitly incorporate register-like mechanisms, which may partially account for their strong empirical performance. Motivated by these observations, we propose Register Guidance, a technique that amplifies the contribution of register tokens responsible for improving visual structure and coherence.
Figures
Forward citations
Cited by 1 Pith paper
-
Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers
Chat-template tokens in LLM-conditioned DiTs act as implicit semantic registers: they absorb object identity from image latents and maintain it, while direct prompt-reading heads are causally inert.
Reference graph
Works this paper leans on
-
[1]
W. H. Zurek, Phys. Rev. D26, 1862 (1982); Rev. Mod. Phys.75, 715 (2003)
1982
-
[2]
G. C. Ghirardi, A. Rimini, T. Weber, Phys. Rev. D34, 470 (1986)
1986
-
[3]
Bassi and G
A. Bassi and G. C. Ghirardi, Phys. Rep.379, 257 (2003)
2003
-
[4]
Baldo, International Journal of Quantum Founda- tions10, 151 (2024)
M. Baldo, International Journal of Quantum Founda- tions10, 151 (2024)
2024
-
[5]
Baldo, International Journal of Quantum Founda- tions11, 346 (2025);12, 364 (2026)
M. Baldo, International Journal of Quantum Founda- tions11, 346 (2025);12, 364 (2026)
2025
-
[6]
Pearle, Int
P. Pearle, Int. J. Theor. Phys.48, 489 (1979)
1979
-
[7]
Pearle, Phys
P. Pearle, Phys. Rev. D33, 2240 (1986)
1986
-
[8]
Kupferman, G
R. Kupferman, G. A. Pavliotis and A. M. Stuart, Phys. Rev. E70, 036120 (2004)
2004
-
[9]
R. L. Stratonovich, SIAM J. Control4, 362 (1966)
1966
-
[10]
Itô, Proc
K. Itô, Proc. Imp. Acad.20, 519 (1944)
1944
-
[11]
B. J. West, A. R. Bulsara, K. Lindenberg, V. Seshadri and K. E. Shuler, Physica A97, 211 (1979)
1979
-
[12]
A. R. Bulsara, K. Lindenberg, V. Seshadri, K. E. Shuler 9 and B. J. West, Physica A97, 234 (1979)
1979
-
[13]
Hasegawa, M
H. Hasegawa, M. Mabuchi and T. Baba, Phys. Lett. A 79, 273 (1980)
1980
-
[14]
N. G. van Kampen, J. Stat. Phys.24, 175 (1981)
1981
-
[15]
Lieb, D.W
E.H. Lieb, D.W. Robinson, Commun. Math. Phys.28, 251 (1972)
1972
-
[16]
Chi-Fang Chen, Andrew Lucas, Chao Yin, Rep. Prog. Phys.86, 116001 (2023)
2023
-
[17]
97, 050401 (2006)
S.Bravyi, M.B.Hastings, F.Verstraete, Phys.Rev.Lett. 97, 050401 (2006)
2006
-
[18]
Rosenstein and M
B. Rosenstein and M. Usher, Phys. Rev. D36, 2381 (1987)
1987
-
[19]
Kowalski and J
K. Kowalski and J. Rembielinski, Phys. Rev. A84, 012108 (2011)
2011
-
[20]
Pedalino, B
S. Pedalino, B. E. Ramírez-Galindo, R. Ferstl, K. Horn- berger, M. Arndt, S. Gerlich, Nature649, 866 (2026)
2026
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.