Pith. sign in

REVIEW 3 major objections 1 minor 1 cited by

Diffusion Transformers benefit from register tokens even without ViT-style high-norm outliers, with larger gains in pixel space than latent space.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 16:42 UTC pith:674KEV4R

load-bearing objection We only have the abstract for the DiT/registers paper; the cached full text is an unrelated quant-ph manuscript, so none of the empirical claims can be checked. the 3 major comments →

arxiv 2605.16147 v2 pith:674KEV4R submitted 2026-05-15 cs.CV

Registers Matter for Pixel-Space Diffusion Transformers

classification cs.CV
keywords diffusion transformersregister tokenspixel-space diffusionvision transformersfeature mapsregister guidanceDiThigh-noise representations
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Vision Transformers use register tokens to soak up high-norm patch outliers that otherwise spoil feature maps. As diffusion models shift to transformer backbones and train directly in pixel space, they start to look more like ViTs, so the natural question is whether registers help them too. This paper shows that DiTs do not produce those high-norm outliers yet still improve when registers are added, and the improvement is larger for pixel-space models than for latent-space ones. Looking at intermediate activations, the authors find that registers produce cleaner feature maps especially at high noise levels, which may explain their usefulness for pixel-space generation. They also note that several recent strong pixel-space DiT designs already contain implicit register-like mechanisms, and they introduce Register Guidance to deliberately amplify the contribution of those tokens for better visual structure and coherence.

Core claim

DiTs lack the high-norm patch-token outliers characteristic of ViTs, yet they still benefit from the addition of register tokens. The benefit is stronger in pixel-space DiTs than in latent-space DiTs; registers yield cleaner intermediate feature maps at high noise levels, and recent high-performing pixel-space architectures already embed register-like mechanisms that may help explain their results.

What carries the argument

Register tokens (and the proposed Register Guidance method that amplifies their contribution). They clean high-noise intermediate feature maps and improve visual structure and coherence even though they are not needed to absorb outliers.

Load-bearing premise

The paper treats the cleaner high-noise feature maps produced by registers as a causal reason for better pixel-space generation, and treats the implicit register-like parts of recent architectures as sufficiently analogous to explain their performance, without a direct causal isolation of either link.

What would settle it

Train capacity-matched pixel-space DiTs with and without registers (or with the implicit register-like components ablated) and check whether the high-noise feature-map cleanliness gap appears or disappears exactly when generation quality improves or drops; if quality gains persist without cleaner maps, or maps clean without quality gains, the claimed mechanism fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Pixel-space DiT designs should include explicit registers or equivalent sink tokens rather than relying only on latent-space practice.
  • Register Guidance can be applied at inference or fine-tuning time to boost structure and coherence by amplifying the register pathway.
  • Implicit register-like components already present in recent pixel-space architectures partially account for their strong empirical results and should be preserved or strengthened.
  • Inspection of intermediate feature maps across noise levels becomes a useful diagnostic for why certain DiT variants generate more coherent images.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Because registers help DiTs without absorbing outliers, their role in generative transformers is likely different from their role in discriminative ViTs and may be tied to the denoising trajectory itself.
  • If high-noise map cleanliness is the operative factor, similar register or sink-token benefits should appear in other high-noise generative models, not only DiTs.
  • Architectures could move beyond borrowed ViT registers toward tokens whose capacity or attention pattern is explicitly scheduled with the noise level.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 1 minor

Summary. The submission is presented as arXiv:2605.16147, a cs.CV paper claiming that Diffusion Transformers lack ViT-style high-norm patch-token outliers yet still benefit from register tokens (more so in pixel space than latent space), that registers yield cleaner high-noise intermediate feature maps, that recent pixel-space DiTs contain implicit register-like mechanisms, and that a proposed Register Guidance method improves structure and coherence. The supplied full manuscript body, however, is an unrelated quant-ph paper by Fabio Siringo on macroscopic superpositions, a stochastic reduction of the Schrödinger equation, and a dynamical derivation of the Born rule. No DiT experiments, ablations, feature-map analyses, or Register Guidance results appear in the body.

Significance. If the abstract’s claims were supported by a matching manuscript, the work would be of clear interest to the generative-modeling community: it would separate register benefits from the classic ViT outlier mechanism, give a pixel- vs latent-space comparison, and offer a practical guidance method. As submitted, that significance cannot be assessed. The body that is present is a conservative, unitary-dynamics account of macroscopic collapse; that is a different contribution in a different field and is not the paper under review.

major comments (3)
  1. Title/abstract vs. full text: the body is Siringo’s quant-ph manuscript (macroscopic ensembles, white-noise matrix elements W_nm, Itô reduction, Born rule), not a DiT/registers paper. Every load-bearing empirical claim in the abstract—no ViT-style outliers in DiTs, larger pixel-space gains, cleaner high-noise maps, implicit register-like mechanisms, Register Guidance efficacy—has no methods, figures, tables, or numbers that can be audited. The submission is not reviewable as the claimed work.
  2. Causal claim in the abstract (“registers produce cleaner feature maps at high noise levels, which may contribute to their effectiveness in pixel-space generation”) cannot be checked: there are no intermediate-representation analyses, noise-level sweeps, or ablations linking map cleanliness to sample quality. The interpretive step from correlation to contribution is therefore unsupported in the provided materials.
  3. The assertion that recent pixel-space DiT architectures “implicitly incorporate register-like mechanisms” is likewise unsubstantiated: no architectural mapping, token-role analysis, or controlled removal/insertion study is present in the body.
minor comments (1)
  1. Even as a standalone quant-ph text, the cached body has presentation issues (e.g., “must bepostu-lated”, “areducingequation”, “EQUA TION”, “STRA TONOVICH”) that would need copy-editing if that paper were under review; they are secondary to the identity mismatch.

Circularity Check

2 steps flagged

Mild definitional recovery of the Born rule from conserved p_n; the dynamical reduction chain from Schrödinger is not circular by construction.

specific steps
  1. self definitional [Sec. IV, after Eqs. (30)–(32); also setup Eqs. (2)–(5)]
    "Then, the Born rule is recovered because, according to Eq.(12), in the Itô interpretation, the average of Eq.(9) gives d⟨p_m(t)⟩/dt=0 (there is no drift term) and the probability P_m =⟨p_m(t)⟩ is equal to the initial value p_m(0), as predicted by the Born rule."

    p_n are defined from the start as the summed squared amplitudes over microstates in each macrostate (Eq. 2), i.e. the Born-rule quantities for that decomposition. Under the Itô reading the SDE has no drift, so ensemble averages of those same p_n are conserved; after reduction, outcome frequencies equal those averages. Matching the Born rule is therefore largely the identification of conserved initial p_n with collapse probabilities, not an independent derivation of a new probability law. The non-circular content is the reduction dynamics itself.

  2. self definitional [Sec. VII (Discussion), condition Eqs. (57)/(60)]
    "the reduction condition ΔEΔX≫ℏc would set a limit to the maximum bin extension which can survive in a superposition, thus specifying what we actually mean by a forbidden macroscopic superposition."

    The same inequality required for the Itô/reducing regime is used to define the class of superpositions that count as macroscopic (and therefore must collapse). Scope of the conclusion is partly fixed by the condition that makes the conclusion true, which is a mild definitional loop around applicability rather than a forced fit of a numerical prediction.

full rationale

The provided full text is Siringo’s quant-ph manuscript on macroscopic superpositions (not the DiT/registers paper named in the header). Auditing that derivation: the paper decouples macro/micro amplitudes, obtains an exact equation for p_n from the Schrödinger equation, approximates the off-diagonal matrix element as white noise from a large energy bandwidth, and argues the Itô (reducing) interpretation from causality (τ_R ≫ τ_c). Those steps cite external results (Pearle’s reducing SDEs, Kupferman et al. on Itô vs Stratonovich) and do not reduce to fitted parameters or author-only uniqueness theorems. The only mild circularity is identification: p_n are defined as the macroscopic squared amplitudes, the Itô reading has no drift so ⟨p_n⟩ is conserved, and collapse frequencies are then set equal to those conserved averages—recovering the Born rule largely by that identification rather than by an independent probability postulate. A secondary soft spot is using the reduction condition ΔEΔX ≫ ℏc to help specify what counts as a forbidden macroscopic superposition. Neither forces the dynamical claim by pure definition. Score 2.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 1 invented entities

Abstract-only review of the DiT/registers paper. Load-bearing premises are domain assumptions from ViT/DiT practice plus interpretive claims about mechanism. No free parameters or invented physical entities are extractable from the abstract; Register Guidance is a proposed method, not a new ontological entity.

axioms (4)
  • domain assumption ViT high-norm patch-token outliers degrade feature maps and are mitigated by register tokens (prior literature).
    Used as the starting analogy that motivates testing registers in DiTs.
  • domain assumption Pixel-space DiTs are sufficiently ViT-like that token-level interventions (registers) remain meaningful.
    Implicit framing of the research question in the abstract.
  • ad hoc to paper Cleaner intermediate feature maps at high noise levels contribute to better pixel-space generation quality.
    Mechanistic bridge offered for why registers help more in pixel space; not independently established in the provided text.
  • ad hoc to paper Some recent pixel-space DiT components function as implicit register-like mechanisms.
    Explanatory claim used to account for strong empirical performance of prior architectures.
invented entities (1)
  • Register Guidance no independent evidence
    purpose: Amplify contribution of register tokens thought to improve visual structure and coherence during generation.
    Proposed technique in the abstract; independent evidence and exact formulation not available in provided materials.

pith-pipeline@v1.1.0-grok45 · 17561 in / 2379 out tokens · 32941 ms · 2026-07-12T16:42:30.435649+00:00 · methodology

0 comments
read the original abstract

Vision Transformers (ViTs) are known to exhibit high-norm patch-token outliers that degrade feature map quality, a problem effectively mitigated by register tokens. As diffusion models increasingly adopt transformer architectures and move toward pixel-space training, they become closer in form to ViTs, raising the question of whether register tokens are also useful for Diffusion Transformers (DiTs). In this work, we show that DiTs differ from ViTs in a key respect: they do not exhibit patch-token outliers but still benefit from registers. Interestingly, registers are more effective in pixel-space DiTs than in latent-space DiTs. By analyzing intermediate representations, we find that register tokens produce cleaner feature maps at high noise levels, which may contribute to their effectiveness in pixel-space generation. We further observe that recent pixel-space DiT architectures implicitly incorporate register-like mechanisms, which may partially account for their strong empirical performance. Motivated by these observations, we propose Register Guidance, a technique that amplifies the contribution of register tokens responsible for improving visual structure and coherence.

Figures

Figures reproduced from arXiv: 2605.16147 by Artem Babenko, Dmitry Baranchuk, Ilia Sudakov, Ilya Drobyshevskiy, Nikita Starodubcev.

Figure 1
Figure 1. Figure 1: Diffusion transformers do not exhibit attention-map outliers. Unlike ViTs, where attention-map anomalies typically appear in low-information regions (e.g., background), DiT attention remains focused on the main objects. Contributions. We find that, unlike ViTs, diffusion transformers in both latent and pixel spaces do not exhibit noticeable high-norm outliers among patch tokens. Instead, patch-token norms … view at source ↗
Figure 2
Figure 2. Figure 2: Without Registers. (a) In DINOv2, anomalies are localized to few image tokens, which exhibit significantly higher norms than others. (b) In contrast, no outliers are observed for pDiTs, suggesting that registers may be unnecessary in this case. (a) (b) 0 50 100 150 200 250 300 token index 0 1 2 3 4 5 6 Token norm ×104 pDiT-B/16, With Registers block 5 block 10 0 50 100 150 200 250 token index 0.0 0.6 1.2 1… view at source ↗
Figure 3
Figure 3. Figure 3: With Registers. (a) As expected, introducing register tokens in DINOv2 shifts high-norm outliers into these tokens. (b) Interestingly, pDiTs also exhibit high-norm tokens in the added registers, even though such outliers are absent without registers. pDiT-B/16, 131M pDiT-L/16, 459M pDiT-H/16, 953M Epoch w/o reg. w/ reg. w/ in-context w/o reg. w/ reg. w/ in-context w/o reg. w/ reg. w/ in-context 200 7.39 5.… view at source ↗
Figure 4
Figure 4. Figure 4: Register tokens consistently reduce feature norms across patch tokens. We measure feature norms for image tokens only (excluding registers) at three diffusion timesteps and observe a consistent reduction across all tokens when registers are used. Original image pDiT-B/16 with registers pDiT-B/16 without registers (a) (b) 0 1 2 3 4 5 6 7 8 9 10 avg Block 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1.0 Timestep … view at source ↗
Figure 5
Figure 5. Figure 5: Register tokens improve intermediate representations. (a) We compute the Total Variation (TV) of intermediate features for models with and without register tokens. We report the ratio (with / without registers), where lower values indicate that the model with registers produces smoother features. We find that registers improve feature smoothness at high noise levels (t ∈ [0, 0.2]). (b) We visualize feature… view at source ↗
Figure 6
Figure 6. Figure 6: Register tokens act as both global information carriers and norm sinks. Linear probing reveals that low-norm register tokens encode meaningful global semantics and achieve strong classification accuracy, whereas low-accuracy registers exhibit extremely large norms, suggesting that they primarily function as norm sinks that absorb magnitude from patch tokens. Input image Attention maps for register tokens (… view at source ↗
Figure 7
Figure 7. Figure 7: (a) Registers with high probing accuracy encode diverse semantic information about an image, whereas (b) low-accuracy norm sinks do not. We visualize attention maps for register tokens and observe that some attend to distinct semantic regions, such as foreground objects and background areas. In contrast, norm sinks with low probing accuracy do not exhibit meaningful semantic structure. First, we find that … view at source ↗
Figure 8
Figure 8. Figure 8: In-context class tokens act as registers. (a) Certain tokens acquire disproportionately high feature norms, functioning as norm sinks. (b) Some tokens encode broad global information, rather than purely class-specific features as originally intended. 2.5 Registers Are Effective in Deeper Layers Next, we ablate both the number of register tokens and the transformer blocks in which they are introduced. We co… view at source ↗
Figure 19
Figure 19. Figure 19: 15 [PITH_FULL_IMAGE:figures/full_fig_p015_19.png] view at source ↗
Figure 9
Figure 9. Figure 9: High-norm outliers consistently emerge within register tokens across timesteps. We visualize token-wise feature norms of pDiT-B/16 with registers for t = 0.0, 0.3, and 0.7, and observe the same behavior in all cases. 0 50 100 150 200 250 token index 1 2 3 4 5 Token norm ×103 DT-B/16, Without Registers block 2 block 5 block 8 block 10 0 50 100 150 200 250 300 token index 0 1 2 3 4 5 Token norm ×104 DT-B/16,… view at source ↗
Figure 10
Figure 10. Figure 10: Token-wise feature norms for pDiTs of varying scales on ImageNet 256 × 256, with and without registers. Without registers, patch-token norms remain uniform across scales. Introducing registers leads to the emergence of high-norm outliers within the register tokens. 18 [PITH_FULL_IMAGE:figures/full_fig_p018_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Token-wise feature norms for VAE-space SiTs of varying scales on ImageNet 256 × 256, with and without registers. SiTs without registers exhibit uniform patch-token norms across scales, while adding registers produces high-norm register tokens. 0 50 100 150 200 250 300 token index 0.5 1.0 1.5 2.0 Token norm ×102 RAE-S, With Registers block 0 block 2 block 5 block 8 block 10 0 50 100 150 200 250 token index… view at source ↗
Figure 12
Figure 12. Figure 12: Token-wise feature norms for DINOv2-space RAEs of varying scales on ImageNet 256 × 256, with and without registers. RAEs without registers exhibit uniform patch-token norms across scales, while adding registers produces high-norm register tokens. 19 [PITH_FULL_IMAGE:figures/full_fig_p019_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: Register tokens consistently reduce feature norms across patch tokens. We measure feature norms for image tokens only (excluding register tokens) at three diffusion timesteps for pDiT models of different scales, and observe a consistent reduction in feature norms across nearly all tokens when register tokens are used. 20 [PITH_FULL_IMAGE:figures/full_fig_p020_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: Register tokens make intermediate representations cleaner by reducing noise. We compute the Total Variation of intermediate features for models with and without register tokens. We report the ratio (with registers / without registers), where lower values indicate that models with registers produce smoother feature representations. We find that register tokens improve feature smoothness at high noise level… view at source ↗
Figure 15
Figure 15. Figure 15: Registers improve spatial organization at high noise levels. In addition to the TV ratio (left), we also analyze the correlation decay slope [45] (right), where lower values indicate stronger spatial organization. Both metrics show the same trend: register tokens improve internal representations at high noise levels, starting from block 4, where the registers are introduced. 21 [PITH_FULL_IMAGE:figures/f… view at source ↗
Figure 16
Figure 16. Figure 16: Linear probing of register tokens under different configurations. (Left) Standard register tokens introduced from the 4th layer; (Middle) Register tokens used as in-context class embeddings introduced from the 4th layer; (Right) Standard register tokens introduced from the 0th layer. Across different timesteps, we find that introducing registers from the earliest layers produces substantially less informa… view at source ↗
Figure 17
Figure 17. Figure 17: Pixel-space pDiTs have the highest feature norms across all tokens for different timesteps compared to latent-space counterparts. We compare token-wise feature-map norms for pDiT, SiT, and RAE models, all without register tokens. 0 4 9 13 17 21 26 avg Block 0.0 0.1 0.2 0.4 0.6 0.7 1.0 Timestep 0.12 0.59 2.05 1.19 0.53 0.18 0.40 0.72 0.23 0.83 1.32 1.19 0.47 0.10 0.18 0.62 0.33 1.10 1.20 0.96 0.44 0.11 0.2… view at source ↗
Figure 18
Figure 18. Figure 18: Pixel-space pDiTs exhibit substantially higher Total Variation (TV) values than latent-space counterparts. We compare the TV ratio of intermediate feature maps across timesteps and transformer blocks for pixel-space pDiTs (pDiT-H) and latent-space models (SiT-XL and RAE￾XL) without registers. Pixel-space pDiTs consistently produce noisier intermediate representations. 22 [PITH_FULL_IMAGE:figures/full_fig… view at source ↗
Figure 19
Figure 19. Figure 19: For SSL ViTs such as DINOv2, register tokens do not reduce patch-token feature norms, unlike in DiTs. We measure feature norms across all tokens for different blocks and model sizes of DINOv2, and observe that register tokens do not consistently reduce feature norms of patch tokens. 0 1000 2000 3000 4000 token index 0.0 0.5 1.0 1.5 2.0 token norm ×104 SD3.5. Norms of text (left) and image tokens (right) b… view at source ↗
Figure 20
Figure 20. Figure 20: Text sequences in text-to-image diffusion models exhibit behavior similar to register tokens in ImageNet-based DiTs: some tokens become high-norm outliers and potentially act as registers. We measure token-wise feature norms in SD3.5 (left) and FLUX (right) for both text and image tokens. We observe that the outliers primarily emerge within the text sequence. 23 [PITH_FULL_IMAGE:figures/full_fig_p023_20.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers

    cs.CV 2026-07 conditional novelty 7.0

    Chat-template tokens in LLM-conditioned DiTs act as implicit semantic registers: they absorb object identity from image latents and maintain it, while direct prompt-reading heads are causally inert.

Reference graph

Works this paper leans on

20 extracted references · cited by 1 Pith paper

  1. [1]

    W. H. Zurek, Phys. Rev. D26, 1862 (1982); Rev. Mod. Phys.75, 715 (2003)

  2. [2]

    G. C. Ghirardi, A. Rimini, T. Weber, Phys. Rev. D34, 470 (1986)

  3. [3]

    Bassi and G

    A. Bassi and G. C. Ghirardi, Phys. Rep.379, 257 (2003)

  4. [4]

    Baldo, International Journal of Quantum Founda- tions10, 151 (2024)

    M. Baldo, International Journal of Quantum Founda- tions10, 151 (2024)

  5. [5]

    Baldo, International Journal of Quantum Founda- tions11, 346 (2025);12, 364 (2026)

    M. Baldo, International Journal of Quantum Founda- tions11, 346 (2025);12, 364 (2026)

  6. [6]

    Pearle, Int

    P. Pearle, Int. J. Theor. Phys.48, 489 (1979)

  7. [7]

    Pearle, Phys

    P. Pearle, Phys. Rev. D33, 2240 (1986)

  8. [8]

    Kupferman, G

    R. Kupferman, G. A. Pavliotis and A. M. Stuart, Phys. Rev. E70, 036120 (2004)

  9. [9]

    R. L. Stratonovich, SIAM J. Control4, 362 (1966)

  10. [10]

    Itô, Proc

    K. Itô, Proc. Imp. Acad.20, 519 (1944)

  11. [11]

    B. J. West, A. R. Bulsara, K. Lindenberg, V. Seshadri and K. E. Shuler, Physica A97, 211 (1979)

  12. [12]

    A. R. Bulsara, K. Lindenberg, V. Seshadri, K. E. Shuler 9 and B. J. West, Physica A97, 234 (1979)

  13. [13]

    Hasegawa, M

    H. Hasegawa, M. Mabuchi and T. Baba, Phys. Lett. A 79, 273 (1980)

  14. [14]

    N. G. van Kampen, J. Stat. Phys.24, 175 (1981)

  15. [15]

    Lieb, D.W

    E.H. Lieb, D.W. Robinson, Commun. Math. Phys.28, 251 (1972)

  16. [16]

    Chi-Fang Chen, Andrew Lucas, Chao Yin, Rep. Prog. Phys.86, 116001 (2023)

  17. [17]

    97, 050401 (2006)

    S.Bravyi, M.B.Hastings, F.Verstraete, Phys.Rev.Lett. 97, 050401 (2006)

  18. [18]

    Rosenstein and M

    B. Rosenstein and M. Usher, Phys. Rev. D36, 2381 (1987)

  19. [19]

    Kowalski and J

    K. Kowalski and J. Rembielinski, Phys. Rev. A84, 012108 (2011)

  20. [20]

    Pedalino, B

    S. Pedalino, B. E. Ramírez-Galindo, R. Ferstl, K. Horn- berger, M. Arndt, S. Gerlich, Nature649, 866 (2026)