TARDIS folds two feed-forward weight matrices into one by linearly approximating activations in common input ranges, then recomputes outliers with a small predictor, claiming 80 percent FFN parameter reduction and up to 1.6x inference speedup.
[Online; accessed 2024- 11-01]
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.LG 1years
2025 1verdicts
REJECT 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Accelerating Large Language Models through Partially Linear Feed-Forward Network
TARDIS folds two feed-forward weight matrices into one by linearly approximating activations in common input ranges, then recomputes outliers with a small predictor, claiming 80 percent FFN parameter reduction and up to 1.6x inference speedup.