Offloading frozen INT8 backbones to a Hailo-8L inference accelerator speeds up on-device head-only fine-tuning by up to 15.4x and cuts energy, but can cost 13-21 accuracy points on quantization-sensitive models.
et al.: Qft: Post-training quantization via fast joint finetuning of all degrees of freedom
1 Pith paper cite this work, alongside 4 external citations. Polarity classification is still indexing.
1
Pith paper citing it
4
external citations · OpenAlex
fields
cs.LG 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Empowering On-Device Model Adaptation with an Edge AI Inference Accelerator
Offloading frozen INT8 backbones to a Hailo-8L inference accelerator speeds up on-device head-only fine-tuning by up to 15.4x and cuts energy, but can cost 13-21 accuracy points on quantization-sensitive models.