A co-designed chip and training method, Topkima-Former, selects the top-k attention scores inside an in-memory converter to speed up softmax by up to 15x with a 0.4-1.2% accuracy drop.
X-former: In- memory acceleration of transformers,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AR 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Topkima-Former: Low-energy, Low-Latency Inference for Transformers using top-k In-memory ADC
A co-designed chip and training method, Topkima-Former, selects the top-k attention scores inside an in-memory converter to speed up softmax by up to 15x with a 0.4-1.2% accuracy drop.