Block-GTQ performs RoPE-aware greedy bit allocation on KV caches using per-block energy scores, cutting logit MAE 32-80% versus uniform TQ-MSE and lifting long-context task scores substantially at 2-3 bits per dimension.
SQuat: Subspace-orthogonal kv cache quantization.arXiv preprint arXiv:2503.24358, 2025
2 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.LG 2years
2026 2verdicts
UNVERDICTED 2representative citing papers
InnerQ delivers 1.3x average speedup over prior KV cache quantization and 2.7x over baseline by inner-dimension grouping, hybrid symmetric/asymmetric quantization, high-precision windows for recent and sink tokens, and prefold per-channel key normalization.
citing papers explorer
-
RoPE-Aware Bit Allocation for KV-Cache Quantization
Block-GTQ performs RoPE-aware greedy bit allocation on KV caches using per-block energy scores, cutting logit MAE 32-80% versus uniform TQ-MSE and lifting long-context task scores substantially at 2-3 bits per dimension.
-
InnerQ: Hardware-Aware Tuning-Free Quantization of KV Cache for Large Language Models
InnerQ delivers 1.3x average speedup over prior KV cache quantization and 2.7x over baseline by inner-dimension grouping, hybrid symmetric/asymmetric quantization, high-precision windows for recent and sink tokens, and prefold per-channel key normalization.