Zero-initializing the modulation weights is the dominant reason adaLN-Zero outperforms adaLN, and replacing it with a Gaussian initialization of std 0.001 improves FID at the same training steps.
Visionllama: A unified llama backbone for vision tasks,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Unveiling the Secret of AdaLN-Zero in Diffusion Transformer
Zero-initializing the modulation weights is the dominant reason adaLN-Zero outperforms adaLN, and replacing it with a Gaussian initialization of std 0.001 improves FID at the same training steps.