Whisper dot-product kernel offloaded to IMAX CGLA with 32KB local memory and burst length 16 yields PDP of 11.58J for tiny Q8_0, 2.35x lower than Jetson AGX Orin.
Ftrans: Energy-efficient acceleration of transformers using fpga
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
citation-role summary
background 1
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
The paper surveys techniques to speed up and reduce the resource needs of LLM inference, organized by data-level, model-level, and system-level changes, with comparative experiments on representative methods.
citing papers explorer
-
Design and Evaluation of Energy-Efficient Whisper Dot-Product Kernel Offloading on a CGLA Architecture
Whisper dot-product kernel offloaded to IMAX CGLA with 32KB local memory and burst length 16 yields PDP of 11.58J for tiny Q8_0, 2.35x lower than Jetson AGX Orin.
-
A Survey on Efficient Inference for Large Language Models
The paper surveys techniques to speed up and reduce the resource needs of LLM inference, organized by data-level, model-level, and system-level changes, with comparative experiments on representative methods.