DepthWeave-KV achieves 8.3x KV cache memory reduction with near-full-cache task quality by factorizing key-value states across transformer layers using shared bases and token-adaptive residuals.
Title resolution pending
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3verdicts
CONDITIONAL 3representative citing papers
Frequency-guided inter-layer KV sharing with logit-aware head routing nearly matches full-cache long-context accuracy at about 3.9× lower peak KV memory.
A frozen video diffusion backbone augmented with low-rank temporal adapters and a recursive prompt bank outperforms prior long-video generation methods on six benchmarks while tuning only 3.8% of parameters.
citing papers explorer
-
DepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache Compression
DepthWeave-KV achieves 8.3x KV cache memory reduction with near-full-cache task quality by factorizing key-value states across transformer layers using shared bases and token-adaptive residuals.
-
FreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM Inference
Frequency-guided inter-layer KV sharing with logit-aware head routing nearly matches full-cache long-context accuracy at about 3.9× lower peak KV memory.
-
Prompt-Adapter Context Routing for Parameter-Efficient Multi-Shot Long Video Extrapolation
A frozen video diffusion backbone augmented with low-rank temporal adapters and a recursive prompt bank outperforms prior long-video generation methods on six benchmarks while tuning only 3.8% of parameters.