Diffusion models show grokking on modular addition by composing periodic operand representations in simple data regimes or by separating arithmetic computation from visual denoising across timesteps in varied regimes.
Vikrant Varma, Rohin Shah, Zachary Kenton, J ´anos Kram ´ar, and Ramana Kumar
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3verdicts
UNVERDICTED 3representative citing papers
A transcoder-based in-place replacement of the bottleneck layer enables selective concept removal in modern diffusion and autoregressive image models without degrading output quality.
The paper identifies a concept-layer topological alignment bottleneck in text-to-video diffusion models and introduces the CLEAR separability-driven optimization framework for targeted concept erasure.
citing papers explorer
-
Grokking of Diffusion Models: Case Study on Modular Addition
Diffusion models show grokking on modular addition by composing periodic operand representations in simple data regimes or by separating arithmetic computation from visual denoising across timesteps in varied regimes.
-
Concept Removal for Frontier Image Generative Models
A transcoder-based in-place replacement of the bottleneck layer enables selective concept removal in modern diffusion and autoregressive image models without degrading output quality.
-
Where Concept Erasure Should Occur: Concept-Layer Alignment in Text-to-Video Diffusion Models
The paper identifies a concept-layer topological alignment bottleneck in text-to-video diffusion models and introduces the CLEAR separability-driven optimization framework for targeted concept erasure.