Sumi is an openly released 7B parameter uniform diffusion language model pretrained from scratch on 1.5T tokens that matches autoregressive models on several benchmarks.
Scaling beyond masked diffusion language models
8 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 8roles
background 1polarities
background 1representative citing papers
The discrete diffusion NELBO equals data entropy plus an exact path KL to the oracle reverse process, and the denoiser, cavity, and score parameterizations are three interconvertible coordinates of the unique optimal reverse jump rate.
Recursive Masked Diffusion Models add recursive depth via repeated application of the same transformer to improve parameter efficiency and reduce inference steps in masked diffusion models.
TokenDrift refines discrete diffusion language models by applying anti-symmetric drifting to soft-token features during training, yielding large reductions in generation perplexity at low NFEs.
Diffusion language models and a CTC-USDM joint decoder improve ASR accuracy over standard approaches.
RePlaid achieves a 20x compute gap to autoregressive models, new SOTA PPL of 22.1 among continuous DLMs on OpenWebText, and competitive scaling laws by aligning architecture with modern discrete DLMs.
Bell-shaped time sampling — drawing the corruption level near t=0.5 — accelerates masked diffusion language model training by up to ~4× without changing the final loss.
ELF applies continuous-time flow matching in embedding space for language generation and reports outperforming prior discrete and continuous diffusion language models with fewer steps.
citing papers explorer
-
Sumi: Open Uniform Diffusion Language Model from Scratch
Sumi is an openly released 7B parameter uniform diffusion language model pretrained from scratch on 1.5T tokens that matches autoregressive models on several benchmarks.
-
What Does a Discrete Diffusion Model Learn?
The discrete diffusion NELBO equals data entropy plus an exact path KL to the oracle reverse process, and the denoiser, cavity, and score parameterizations are three interconvertible coordinates of the unique optimal reverse jump rate.
-
Recursive Scaling in Masked Diffusion Models
Recursive Masked Diffusion Models add recursive depth via repeated application of the same transformer to improve parameter efficiency and reduce inference steps in masked diffusion models.
-
Drifting Objectives for Refining Discrete Diffusion Language Models
TokenDrift refines discrete diffusion language models by applying anti-symmetric drifting to soft-token features during training, yielding large reductions in generation perplexity at low NFEs.
-
Diffusion Language Models for Speech Recognition
Diffusion language models and a CTC-USDM joint decoder improve ASR accuracy over standard approaches.
-
Continuous Diffusion Scales Competitively with Discrete Diffusion for Language
RePlaid achieves a 20x compute gap to autoregressive models, new SOTA PPL of 22.1 among continuous DLMs on OpenWebText, and competitive scaling laws by aligning architecture with modern discrete DLMs.
-
Understanding and Accelerating the Training of Masked Diffusion Language Models
Bell-shaped time sampling — drawing the corruption level near t=0.5 — accelerates masked diffusion language model training by up to ~4× without changing the final loss.
-
ELF: Embedded Language Flows
ELF applies continuous-time flow matching in embedding space for language generation and reports outperforming prior discrete and continuous diffusion language models with fewer steps.