GANQ minimizes layer-wise output error for lookup-table based non-uniform weight quantization, improving LLM perplexity at 3-4 bits and enabling up to 2.57x inference speedup.
Llm-qat: Data-free quantization aware training for large language models
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language Models
GANQ minimizes layer-wise output error for lookup-table based non-uniform weight quantization, improving LLM perplexity at 3-4 bits and enabling up to 2.57x inference speedup.