A diffusion language model can be trained to reason via supervised fine-tuning plus a new policy-gradient RL method, diffu-GRPO, improving benchmark accuracy over the base model.
Arel’s sudoku generator
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement Learning
A diffusion language model can be trained to reason via supervised fine-tuning plus a new policy-gradient RL method, diffu-GRPO, improving benchmark accuracy over the base model.