Block-GTQ performs RoPE-aware greedy bit allocation on KV caches using per-block energy scores, cutting logit MAE 32-80% versus uniform TQ-MSE and lifting long-context task scores substantially at 2-3 bits per dimension.
Query-aware mixed-precision kv cache quantization for long-context reasoning
2 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.LG 2years
2026 2representative citing papers
RateQuant uses per-quantizer distortion calibration and reverse waterfilling from rate-distortion theory to optimally allocate mixed-precision bits across KV cache heads, resolving a failure mode where mismatched distortion models invert allocation order.
citing papers explorer
-
RoPE-Aware Bit Allocation for KV-Cache Quantization
Block-GTQ performs RoPE-aware greedy bit allocation on KV caches using per-block energy scores, cutting logit MAE 32-80% versus uniform TQ-MSE and lifting long-context task scores substantially at 2-3 bits per dimension.
-
RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory
RateQuant uses per-quantizer distortion calibration and reverse waterfilling from rate-distortion theory to optimally allocate mixed-precision bits across KV cache heads, resolving a failure mode where mismatched distortion models invert allocation order.