REVIEW 3 cited by
EC-DIT: Scaling Diffusion Transformers with Adaptive Expert-Choice Routing
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Diffusion transformers have been widely adopted for text-to-image synthesis. While scaling these models up to billions of parameters shows promise, the effectiveness of scaling beyond current sizes remains underexplored and challenging. By explicitly exploiting the computational heterogeneity of image generations, we develop a new family of Mixture-of-Experts (MoE) models (EC-DIT) for diffusion transformers with expert-choice routing. EC-DIT learns to adaptively optimize the compute allocated to understand the input texts and generate the respective image patches, enabling heterogeneous computation aligned with varying text-image complexities. This heterogeneity provides an efficient way of scaling EC-DIT up to 97 billion parameters and achieving significant improvements in training convergence, text-to-image alignment, and overall generation quality over dense models and conventional MoE models. Through extensive ablations, we show that EC-DIT demonstrates superior scalability and adaptive compute allocation by recognizing varying textual importance through end-to-end training. Notably, in text-to-image alignment evaluation, our largest models achieve a state-of-the-art GenEval score of 71.68% and still maintain competitive inference speed with intuitive interpretability.
Forward citations
Cited by 3 Pith papers
-
Instruction-Following Pruning for Large Language Models
A small predictor reads each user instruction and dynamically chooses which feed-forward weights of a large LLM to activate, beating dense models of the same activated size.
-
Scaling Properties of Diffusion Models for Perceptual Tasks
Diffusion models for depth, optical flow, and amodal segmentation improve along power laws as training and test-time compute scale, and the fitted recipes match prior specialist models with less data.
-
Expert-Data Alignment Governs Generation Quality in Decentralized Diffusion Models
Sparse Top-2 routing beats full ensemble in decentralized diffusion models, and the paper attributes this to expert-data alignment rather than numerical stability — though much of the supporting evidence is circular.
Discussion (0). Continue with ORCID to comment.