Keeping only initial, separator (punctuation), and neighboring tokens in the KV cache preserves LLM performance while cutting cache size by more than half on GSM8K chain-of-thought.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator
Keeping only initial, separator (punctuation), and neighboring tokens in the KV cache preserves LLM performance while cutting cache size by more than half on GSM8K chain-of-thought.