A cache that pre-generates answers to instructions a language model judges likely, achieving hit rates up to 2.3x the exact-repetition upper bound on WildChat and up to 50% faster token generation in vLLM.
GPTCache: An open-source semantic cache for LLM applications enabling faster answers and cost savings
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
InstCache: A Predictive Cache for LLM Serving
A cache that pre-generates answers to instructions a language model judges likely, achieving hit rates up to 2.3x the exact-repetition upper bound on WildChat and up to 50% faster token generation in vLLM.