Preloading a small knowledge base into a long-context LLM with a cached KV cache can beat traditional RAG on accuracy and latency, but the paper's evaluation gives CAG an unfair advantage by feeding it the exact answer-bearing document subset.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2024 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Don't Do RAG: When Cache-Augmented Generation is All You Need for Knowledge Tasks
Preloading a small knowledge base into a long-context LLM with a cached KV cache can beat traditional RAG on accuracy and latency, but the paper's evaluation gives CAG an unfair advantage by feeding it the exact answer-bearing document subset.