Pith. sign in

REVIEW 1 cited by

On Inductive Biases That Enable Generalization of Diffusion Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.21273 v1 pith:PHKFDQKJ submitted 2024-10-28 cs.CV

classification cs.CV
keywords generalizationattentioninductivebiasesdiffusionfindwindowsbases
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent work studying the generalization of diffusion models with UNet-based denoisers reveals inductive biases that can be expressed via geometry-adaptive harmonic bases. However, in practice, more recent denoising networks are often based on transformers, e.g., the diffusion transformer (DiT). This raises the question: do transformer-based denoising networks exhibit inductive biases that can also be expressed via geometry-adaptive harmonic bases? To our surprise, we find that this is not the case. This discrepancy motivates our search for the inductive bias that can lead to good generalization in DiT models. Investigating the pivotal attention modules of a DiT, we find that locality of attention maps are closely associated with generalization. To verify this finding, we modify the generalization of a DiT by restricting its attention windows. We inject local attention windows to a DiT and observe an improvement in generalization. Furthermore, we empirically find that both the placement and the effective attention size of these local attention windows are crucial factors. Experimental results on the CelebA, ImageNet, and LSUN datasets show that strengthening the inductive bias of a DiT can improve both generalization and generation quality when less training data is available. Source code will be released publicly upon paper publication. Project page: dit-generalization.github.io/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Latent-Reframe: Enabling Camera Control for Video Diffusion Model without Training

    cs.CV 2024-12 conditional novelty 5.0 of 10

    Latent-Reframe steers a pre-trained video diffusion model along a target camera trajectory by reframing halfway-denoised latents with time-aware 3D point clouds and then inpainting the resulting gaps, all without fine-tuning.

Pith tools