Stage-level profiling on mobile SoC finds CPUs faster than NPUs in prefill (up to 1.6x) and only modest NPU gains in decode (1.05-1.2x), plus rising energy with greater NPU offload.
AWQ: Activation-aware weight quantiza- tion for on-device LLM compression and acceleration,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AR 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference
Stage-level profiling on mobile SoC finds CPUs faster than NPUs in prefill (up to 1.6x) and only modest NPU gains in decode (1.05-1.2x), plus rising energy with greater NPU offload.