A tensor-compressed transformer training accelerator on FPGA that stores all parameters on-chip, claiming 20-51x memory reduction and up to 4x energy savings per epoch versus an RTX 3090.
EF-train: Enable efficient on-device CNN training on FPGA through data reshaping for online adaptation or personalization,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Ultra Memory-Efficient On-FPGA Training of Transformers via Tensor-Compressed Optimization
A tensor-compressed transformer training accelerator on FPGA that stores all parameters on-chip, claiming 20-51x memory reduction and up to 4x energy savings per epoch versus an RTX 3090.