← back to paper
arxiv: 2607.01237 · 2 revisions
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression