Tensor-train compression of LLM linear layers plus a systolic FPGA accelerator reduces first-token latency by 1.45x to 1.57x, with whole-network compression of 1.94x and 1.60x and modest accuracy loss.
A high throughput multi-bit-width 3d systolic accelerator for nas optimized deep neural networks on fpga,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.AR 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
A Tensor-Train Decomposition based Compression of LLMs on Group Vector Systolic Accelerator
Tensor-train compression of LLM linear layers plus a systolic FPGA accelerator reduces first-token latency by 1.45x to 1.57x, with whole-network compression of 1.94x and 1.60x and modest accuracy loss.