Pith. sign in

Paper Citation Record · LEDGER

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference

As of 23 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2605.27435.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.27435 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T15:03:31.289211Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact7
  • verified fuzzy10
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 72c2fc1c-2789-47f1-a691-d88637bc627a · outbound

This paper cites A Survey of LLM Inference Systems.

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference A Survey of LLM Inference Systems

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:04:46.172024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T15:03:31.289211Z digest=sha256:46d473163cf51c25f1ee1db15bf90b0572a0f260a8d0a6204800fcf4df9865c0

Observation 4cd305fd-a733-450c-b296-270d1c1580d9 · outbound

This paper cites Empowering edge intelligence: A comprehensive survey on on-device AI models,.

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference Empowering edge intelligence: A comprehensive survey on on-device AI models,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T18:45:27.630511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T15:03:31.289211Z digest=sha256:05ee208b702d193e0b87b397bc6eebd9a645187e6b5189af201c916c54cef6fa

Observation 5474a78b-3d68-46c3-81d5-f5794b57b03b · outbound

This paper cites LLaMA 3.2: Vision and edge-optimized models for multi- modal and mobile AI,.

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference LLaMA 3.2: Vision and edge-optimized models for multi- modal and mobile AI,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T18:45:27.628414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T15:03:31.289211Z digest=sha256:ad656e7ed49c2627fc91a544ee9df1847a51c8c7f00510a42d2174d76261aedf

Observation 3a8cdc4c-220d-4b2d-95c9-f8ebc3c82caf · outbound

This paper cites Qwen3 Technical Report.

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference Qwen3 Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-30T15:04:46.185045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T15:03:31.289211Z digest=sha256:1815be172827264fe2582881d3b97521fa3419ad920b44266fa86f86e425a678

Observation 01f58251-7609-4ccc-bd02-7a78e73df3e5 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-30T15:04:46.174674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T15:03:31.289211Z digest=sha256:dc99edb17aef5b387230bd01ffbeaa386a89cb823e63ab7feb84a134cca6335d

Observation 729cc128-fa7a-4ba3-b2e9-575a32ef2bff · outbound

This paper cites AWQ: Activation-aware weight quantiza- tion for on-device LLM compression and acceleration,.

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference AWQ: Activation-aware weight quantiza- tion for on-device LLM compression and acceleration,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T18:45:27.633922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T15:03:31.289211Z digest=sha256:2cdd65fb5b4805eea5d351368703d64521b64a406c30013fcb41d0216babea39

Observation 889cde39-ecb6-414b-b0ff-bda5ba117faa · outbound

This paper cites SmoothQuant: Accurate and efficient post-training quantization for large language models,.

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference SmoothQuant: Accurate and efficient post-training quantization for large language models,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T18:45:27.625882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T15:03:31.289211Z digest=sha256:06bad304a9be01588ccfeb6aeb556de90609d82fddf04a3e265fafd0e1c1112f

Observation d56d4330-afff-4cf2-b55b-6864e4d2789b · outbound

This paper cites Scaling Laws for Precision.

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference Scaling Laws for Precision

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T15:04:46.177292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T15:03:31.289211Z digest=sha256:475fe268ea6336b69abbd3579b961f07ef6a5157475760463211ec25d5f34044

Observation 21935cb9-172e-44a0-b139-cc05bb947124 · outbound

This paper cites Efficient memory management for large language model serving with PagedAttention,.

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference Efficient memory management for large language model serving with PagedAttention,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T18:45:27.622158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T15:03:31.289211Z digest=sha256:bd8b2918574bc1bd9a559635ff3a8cc4e989980b63b23204c7f366a5f1091485

Observation 16c7ed9f-6b6e-4285-a022-2b2a5177799d · outbound

This paper cites MLC-LLM: Universal llm deployment engine with ml com- pilation,.

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference MLC-LLM: Universal llm deployment engine with ml com- pilation,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T18:45:27.639236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T15:03:31.289211Z digest=sha256:02cf332b28fc4a3b2f009b866b27434befba38616fa1814249975e8ba9bf641a

Observation 193de604-d52b-485a-9bf1-930be28d0665 · outbound

This paper cites Fast on-device LLM inference with NPUs,.

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference Fast on-device LLM inference with NPUs,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T18:45:27.620264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T15:03:31.289211Z digest=sha256:e12fced1b068900a4c3dbb6da1f76300e875712013c8e1c2e4758ca323f6b096

Observation e6e60fec-694c-40db-a75a-e7b687591a5c · outbound

This paper cites PowerInfer-2: Fast Large Language Model Inference on a Smartphone.

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:04:46.182427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T15:03:31.289211Z digest=sha256:5542a9455c41f74a73b1c6061afd7d85198c184cae2e2dc6e0eabc3fc2b76a16

Observation f1bbf492-22d1-4b26-89be-f67b6f343da2 · outbound

This paper cites Flightllm: Efficient large language model inference with a complete mapping flow on FPGAs,.

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference Flightllm: Efficient large language model inference with a complete mapping flow on FPGAs,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T18:45:27.624019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T15:03:31.289211Z digest=sha256:fb919d5228e43d09f3f4060d7bb3c49f42ce30293135d9c27e7d8b176607d386

Observation f0ba2ee1-1d48-4343-abd1-db478141dbaf · outbound

This paper cites MLPerf Mobile Inference Benchmark.

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference MLPerf Mobile Inference Benchmark

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:04:46.179771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T15:03:31.289211Z digest=sha256:649b6f52d3b1e5f5f188a131e9770eb1b38c86d00e8c8781c47966071da7771d

Observation 5d1dd023-4a81-4d4f-8038-6656acde83c4 · outbound

This paper cites Characterizing mobile SoC for accelerating heterogeneous LLM inference,.

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference Characterizing mobile SoC for accelerating heterogeneous LLM inference,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T18:45:27.642901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T15:03:31.289211Z digest=sha256:fd7b31de9c00ea035c3280014b61c2649154a9f0ac60cbdd82894fc7ae5dee9a

Observation 579fe8a0-dfef-46b8-ac71-1e6ab1fd8449 · outbound

This paper cites Scaling llm test-time compute with mobile npu on smart- phones.

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference Scaling llm test-time compute with mobile npu on smart- phones

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:04:46.187806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T15:03:31.289211Z digest=sha256:8c368fdd0a94c8bc2b68cdc9f90057501a6a95a46449ea4837da5c6dca4ef29d

Observation 6613b2e4-7add-4248-9952-8a083720d293 · outbound

This paper cites Challenging GPU Dominance: When CPUs Outperform for On-Device LLM Inference.

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference Challenging GPU Dominance: When CPUs Outperform for On-Device LLM Inference

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:04:46.190798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T15:03:31.289211Z digest=sha256:d379f48e99000c48ef7d8834ee86917bb4061294e157ba7d4c599e8f6c2cdd3d

Observation 6df9f1df-40f1-4179-8587-ec2d85c833c9 · outbound

This paper cites llama.cpp: Llm inference in c/c++,.

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference llama.cpp: Llm inference in c/c++,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T18:45:27.641047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T15:03:31.289211Z digest=sha256:2e1b82d25c0427b34f259a02a9d911b7e6cc78bebd9c4b6103853a7f76e1e374

Pith citing papers

No inbound Pith citation observations are available.