Tensor-train compression of LLM linear layers plus a systolic FPGA accelerator reduces first-token latency by 1.45x to 1.57x, with whole-network compression of 1.94x and 1.60x and modest accuracy loss.
A high performance multi-bit-width booth vector systolic accelerator for nas optimized deep learning neural networks,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.AR 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
A Tensor-Train Decomposition based Compression of LLMs on Group Vector Systolic Accelerator
Tensor-train compression of LLM linear layers plus a systolic FPGA accelerator reduces first-token latency by 1.45x to 1.57x, with whole-network compression of 1.94x and 1.60x and modest accuracy loss.