Keeping only initial, separator (punctuation), and neighboring tokens in the KV cache preserves LLM performance while cutting cache size by more than half on GSM8K chain-of-thought.
Additionally, feed-forward networks with ReLU activation can effectively represent any piecewise linear function
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
method 1
citation-polarity summary
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1roles
method 1polarities
support 1representative citing papers
citing papers explorer
-
SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator
Keeping only initial, separator (punctuation), and neighboring tokens in the KV cache preserves LLM performance while cutting cache size by more than half on GSM8K chain-of-thought.