CONF-KV dynamically allocates KV cache budget via model confidence, matching full-KV perplexity within 1.5-2.1 points at sliding-window memory cost and improving retrieval accuracy.
Scissorhands: Exploiting the persistence of importance hypothesis for LLM KV cache compression at test time
2 Pith papers cite this work. Polarity classification is still indexing.
years
2026 2verdicts
UNVERDICTED 2representative citing papers
FastOCR dynamically selects a small subset of visual tokens per decoding step using focal-guided pruning and cross-step reuse, retaining 98% accuracy on Qwen2.5-VL while attending to only 5% of tokens and cutting attention latency by 3x.
citing papers explorer
-
CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM
CONF-KV dynamically allocates KV cache budget via model confidence, matching full-KV perplexity within 1.5-2.1 points at sliding-window memory cost and improving retrieval accuracy.
-
FastOCR: Dynamic Visual Fixation via KV Cache Pruning for Efficient Document Parsing
FastOCR dynamically selects a small subset of visual tokens per decoding step using focal-guided pruning and cross-step reuse, retaining 98% accuracy on Qwen2.5-VL while attending to only 5% of tokens and cutting attention latency by 3x.