Pith. sign in

Scaling beyond masked diffusion language models

8 Pith papers cite this work. Polarity classification is still indexing.

8 Pith papers citing it

citation-role summary

background 1

citation-polarity summary

fields

cs.CL 5 cs.LG 3

years

2026 8

roles

background 1

polarities

background 1

representative citing papers

What Does a Discrete Diffusion Model Learn?

cs.LG · 2026-07-06 · accept · novelty 7.0

The discrete diffusion NELBO equals data entropy plus an exact path KL to the oracle reverse process, and the denoiser, cavity, and score parameterizations are three interconvertible coordinates of the unique optimal reverse jump rate.

Recursive Scaling in Masked Diffusion Models

cs.LG · 2026-06-16 · unverdicted · novelty 7.0

Recursive Masked Diffusion Models add recursive depth via repeated application of the same transformer to improve parameter efficiency and reduce inference steps in masked diffusion models.

ELF: Embedded Language Flows

cs.CL · 2026-05-11 · unverdicted · novelty 6.0 · 2 refs

ELF applies continuous-time flow matching in embedding space for language generation and reports outperforming prior discrete and continuous diffusion language models with fewer steps.

citing papers explorer

Showing 8 of 8 citing papers.

  • Sumi: Open Uniform Diffusion Language Model from Scratch cs.CL · 2026-06-17 · unverdicted · none · ref 30

    Sumi is an openly released 7B parameter uniform diffusion language model pretrained from scratch on 1.5T tokens that matches autoregressive models on several benchmarks.

  • What Does a Discrete Diffusion Model Learn? cs.LG · 2026-07-06 · accept · none · ref 15

    The discrete diffusion NELBO equals data entropy plus an exact path KL to the oracle reverse process, and the denoiser, cavity, and score parameterizations are three interconvertible coordinates of the unique optimal reverse jump rate.

  • Recursive Scaling in Masked Diffusion Models cs.LG · 2026-06-16 · unverdicted · none · ref 34

    Recursive Masked Diffusion Models add recursive depth via repeated application of the same transformer to improve parameter efficiency and reduce inference steps in masked diffusion models.

  • Drifting Objectives for Refining Discrete Diffusion Language Models cs.CL · 2026-05-19 · unverdicted · none · ref 11

    TokenDrift refines discrete diffusion language models by applying anti-symmetric drifting to soft-token features during training, yielding large reductions in generation perplexity at low NFEs.

  • Diffusion Language Models for Speech Recognition cs.CL · 2026-04-15 · unverdicted · none · ref 31

    Diffusion language models and a CTC-USDM joint decoder improve ASR accuracy over standard approaches.

  • Continuous Diffusion Scales Competitively with Discrete Diffusion for Language cs.CL · 2026-05-18 · conditional · none · ref 56

    RePlaid achieves a 20x compute gap to autoregressive models, new SOTA PPL of 22.1 among continuous DLMs on OpenWebText, and competitive scaling laws by aligning architecture with modern discrete DLMs.

  • Understanding and Accelerating the Training of Masked Diffusion Language Models cs.LG · 2026-05-13 · conditional · none · ref 59

    Bell-shaped time sampling — drawing the corruption level near t=0.5 — accelerates masked diffusion language model training by up to ~4× without changing the final loss.

  • ELF: Embedded Language Flows cs.CL · 2026-05-11 · unverdicted · none · ref 58 · 2 links

    ELF applies continuous-time flow matching in embedding space for language generation and reports outperforming prior discrete and continuous diffusion language models with fewer steps.