Opt-GPTQ ports grouped query attention, paging, and ALiBi into vLLM on Hygon DCU chips and measures small throughput gains, but lacks a GQA baseline, error bars, accuracy checks, and code.
Optimizing deep learning workloads on domestic AI accelerators,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.DC 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Opt-GPTQ: An Optimized GPTQ Combining Sparse Attention and Quantization Techniques
Opt-GPTQ ports grouped query attention, paging, and ALiBi into vLLM on Hygon DCU chips and measures small throughput gains, but lacks a GQA baseline, error bars, accuracy checks, and code.