Extends ARM-CL with transformer kernels and CPU-GPU cooperative inference to achieve up to 3x faster inference and 15.72% additional latency reduction on ARM HMPSoCs.
Squeezebert: What can computer vision teach nlp about efficient neural networks?
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AR 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
Fast Transformer Inference on ARM-Based HMPSoCs
Extends ARM-CL with transformer kernels and CPU-GPU cooperative inference to achieve up to 3x faster inference and 15.72% additional latency reduction on ARM HMPSoCs.