TruncFormer statically places truncations in private LLM inference so all nonlinear operations reduce to adds, multiplies, and truncations, cutting estimated truncation latency by up to about 1.92x versus PUMA without hurting accuracy.
Oblivious neu- ral network predictions via minionn transformations
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CR 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
TruncFormer: Private LLM Inference Using Only Truncations
TruncFormer statically places truncations in private LLM inference so all nonlinear operations reduce to adds, multiplies, and truncations, cutting estimated truncation latency by up to about 1.92x versus PUMA without hurting accuracy.