Spherical soft-masking (Fréchet mean + SLERP with mask-norm restoration) improves MAUVE and generative perplexity over LERP feedback in a 169M-parameter masked diffusion language model.
At a typical mask- ing rate of 30-70%, this reduces FLOPs and intermediate memory by a factor proportional to the masking fraction
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Lost in Interpolation: Why Predictive Feedback Fails in Diffusion Language Models
Spherical soft-masking (Fréchet mean + SLERP with mask-norm restoration) improves MAUVE and generative perplexity over LERP feedback in a 169M-parameter masked diffusion language model.