Pith. sign in

Query-aware mixed-precision kv cache quantization for long-context reasoning

2 Pith papers cite this work. Polarity classification is still indexing.

2 Pith papers citing it

fields

cs.LG 2

years

2026 2

representative citing papers

RoPE-Aware Bit Allocation for KV-Cache Quantization

cs.LG · 2026-06-23 · unverdicted · novelty 7.0

Block-GTQ performs RoPE-aware greedy bit allocation on KV caches using per-block energy scores, cutting logit MAE 32-80% versus uniform TQ-MSE and lifting long-context task scores substantially at 2-3 bits per dimension.

citing papers explorer

Showing 2 of 2 citing papers.

  • RoPE-Aware Bit Allocation for KV-Cache Quantization cs.LG · 2026-06-23 · unverdicted · none · ref 44

    Block-GTQ performs RoPE-aware greedy bit allocation on KV caches using per-block energy scores, cutting logit MAE 32-80% versus uniform TQ-MSE and lifting long-context task scores substantially at 2-3 bits per dimension.

  • RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory cs.LG · 2026-04-22 · conditional · none · ref 27 · 2 links

    RateQuant uses per-quantizer distortion calibration and reverse waterfilling from rate-distortion theory to optimally allocate mixed-precision bits across KV cache heads, resolving a failure mode where mismatched distortion models invert allocation order.