GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs

· 2026 · cs.LG · arXiv 2605.23078

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

open full Pith review browse 1 citing papers arXiv PDF

abstract

Mixture-of-Experts Large Language Models (MoE-LLMs) achieve strong performance but incur substantial memory overhead due to massive expert parameters. Mixed-precision quantization mitigates this cost by allocating expert-wise bit-widths based on their importance, approaching the accuracy-memory Pareto frontier and enabling extreme low-bit quantization. However, existing methods rely on layer-wise importance estimation and overlook router shifts induced by quantization, resulting in suboptimal allocation and routing. In this work, we propose Global Expert-level Mixed-precision Quantization (GEMQ) to overcome these limitations via (1) a global linear-programming formulation that captures model-wide expert importance based on quantization error analysis, and (2) efficient router fine-tuning to adapt routing to quantized experts. These components are integrated into a progressive quantization framework that iteratively refines importance estimation and allocation. Experiments demonstrate that GEMQ significantly reduces memory and accelerates inference with minimal accuracy degradation. Source code is available at https://github.com/jndeng/GEMQ .

representative citing papers

Beyond Activation Alignment:The Alignment-Diversity Tradeoff in Task-Aware LLM Quantization

cs.LG · 2026-07-01 · conditional · novelty 7.0

TASA improves task-aware mixed-precision LLM quantization by searching calibration data mixtures via gradient-trace alignment and aggregating perplexity plus reasoning sensitivity signals, enabling 3.5-bit models to match or beat 4-bit baselines with over 20-point gains on GSM8K.

citing papers explorer

Showing 1 of 1 citing paper.

Beyond Activation Alignment:The Alignment-Diversity Tradeoff in Task-Aware LLM Quantization cs.LG · 2026-07-01 · conditional · none · ref 4 · internal anchor
TASA improves task-aware mixed-precision LLM quantization by searching calibration data mixtures via gradient-trace alignment and aggregating perplexity plus reasoning sensitivity signals, enabling 3.5-bit models to match or beat 4-bit baselines with over 20-point gains on GSM8K.

GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs

fields

years

verdicts

representative citing papers

citing papers explorer