CachePrune enables fine-grained, token-level KV cache reuse across LLM requests by masking sensitive segments, eliminating direct side-channel leakage while cutting TTFT by 4.5x and raising hit rates by 44% versus prior coarse-grained methods.
Title resolution pending
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2verdicts
UNVERDICTED 2representative citing papers
CacheRAG is a cache-augmented architecture for LLM KGQA using ISR parsing, hierarchical MMR-based retrieval, and bounded subgraph expansion, claiming +13.2% accuracy gains on CRAG.
citing papers explorer
-
CachePrune: Privacy-Aware and Fine-Grained KV Cache Sharing for Efficient LLM Inference
CachePrune enables fine-grained, token-level KV cache reuse across LLM requests by masking sensitive segments, eliminating direct side-channel leakage while cutting TTFT by 4.5x and raising hit rates by 44% versus prior coarse-grained methods.
-
CacheRAG: A Semantic Caching System for Retrieval-Augmented Generation in Knowledge Graph Question Answering
CacheRAG is a cache-augmented architecture for LLM KGQA using ISR parsing, hierarchical MMR-based retrieval, and bounded subgraph expansion, claiming +13.2% accuracy gains on CRAG.