A training-free, refinement-aware expert budget allocation scheme (REFLEX) reduces selected expert-token pairs by about 15% on MoE diffusion language models without hurting benchmark quality.
Dai and Simon Tong and Dmitry Lepikhin and Yuanzhong Xu and Maxim Krikun and Yanqi Zhou and Adams Wei Yu and Orhan Firat and Barret Zoph and Liam Fedus and Maarten P
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
REFLEX: Rethinking MoE Inference as Refinement-Aware Compute Allocation in Diffusion Language Models
A training-free, refinement-aware expert budget allocation scheme (REFLEX) reduces selected expert-token pairs by about 15% on MoE diffusion language models without hurting benchmark quality.