A training-free, refinement-aware expert budget allocation scheme (REFLEX) reduces selected expert-token pairs by about 15% on MoE diffusion language models without hurting benchmark quality.
CoRR , volume =
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
REFLEX: Rethinking MoE Inference as Refinement-Aware Compute Allocation in Diffusion Language Models
A training-free, refinement-aware expert budget allocation scheme (REFLEX) reduces selected expert-token pairs by about 15% on MoE diffusion language models without hurting benchmark quality.