Pith. sign in

Paper Citation Record · LEDGER

PowerInfer-2: Fast Large Language Model Inference on a Smartphone

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 42 inbound Pith citation observations for arXiv:2406.06282.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.06282 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 42 of 42 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:09:55.031898Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

5
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 456d174e-001a-4ade-b2fd-533489973d47 · inbound

BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices cites this paper.

BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 132

Resolution
unresolved
no resolver link, observed 2026-08-12T19:33:01.033890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:33:01.033890Z digest=sha256:e0b4fb2a88a7bf4e22213d2633d72a1723e059c686a183a092eacd814621ce8e

Observation 8c7d2330-15c8-4845-b7b3-f9906a84d1b9 · inbound

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking cites this paper.

Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T04:30:36.119613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:30:36.119613Z digest=sha256:5c7d43c71239b34576cf2481a27930de5c1abd5495a483d47802256a4ca12fc1

Observation f0aac06a-5260-43f4-bdb5-24caff9e8358 · inbound

Densing Law of LLMs cites this paper.

Densing Law of LLMs PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T21:40:05.602148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:40:05.602148Z digest=sha256:c7410b1c941d2ec130bf0b77360c21cbbcb746fbf597c34d2ac1e0968e4ee460

Observation 7c569c36-6cf0-4a3b-8347-14970b0e2415 · inbound

IC-Cache: Efficient Large Language Model Serving via In-context Caching cites this paper.

IC-Cache: Efficient Large Language Model Serving via In-context Caching PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T16:59:01.428571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:59:01.428571Z digest=sha256:b71bfc0bbc35397f160c06699362af4c9a3cb176d839869298e2436dacab750a

Observation 73483b14-f398-4fdc-ae5e-94611d77db4d · inbound

iServe: An Intent-based Serving System for LLMs cites this paper.

iServe: An Intent-based Serving System for LLMs PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 120

Resolution
unresolved
no resolver link, observed 2026-08-10T21:37:14.534676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:37:14.534676Z digest=sha256:50975a6112b100636cad25a950af42e16557935ee8ed006a71f1be531e2d6b8f

Observation 339db099-dd19-46ef-be92-964ead0a53a3 · inbound

Scaling On-Device GPU Inference for Large Generative Models cites this paper.

Scaling On-Device GPU Inference for Large Generative Models PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T04:51:31.537714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:51:31.537714Z digest=sha256:95b1cfa2810d11b3307136897722dbc3415b82b5c72d75577826f57fd2b34d3c

Observation 4b03021e-085b-46b0-b834-2a12b430ecde · inbound

Responsive DNN Adaptation for Video Analytics against Environment Shift via Hierarchical Mobile-Cloud Collaborations cites this paper.

Responsive DNN Adaptation for Video Analytics against Environment Shift via Hierarchical Mobile-Cloud Collaborations PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-16T05:09:55.031898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:09:55.031898Z digest=sha256:8303e77cb4a33d9eabf8c438fdb1bfa995aa82b33caffa2b08d919fd29563b9e

Observation 46a15fe9-ec8e-41b3-98e1-6d898f5b7c49 · inbound

RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference cites this paper.

RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 111

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:35:24.624547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T15:59:04.724780Z digest=sha256:8b3834adea5df082e57eb108f99e46517330db51201d911877827e6351e93886

Observation fc9a2071-2c3c-4899-ad54-d6cbf007e740 · inbound

FloE: On-the-Fly MoE Inference on Memory-constrained GPU cites this paper.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.153988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.153988Z digest=sha256:dc60e8c4cc41247252bcd718494b87448210ade8ea93743754da6b59392b7b0c

Observation 0f5c1124-bece-4e9b-a235-fd754a1aaa5e · inbound

CoreMatching: A Co-adaptive Sparse Inference Framework with Token and Neuron Pruning for Comprehensive Acceleration of Vision-Language Models cites this paper.

CoreMatching: A Co-adaptive Sparse Inference Framework with Token and Neuron Pruning for Comprehensive Acceleration of Vision-Language Models PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:17.250081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:17.250081Z digest=sha256:723fba61959a7b7137503ecdcd2b0b8bfecfedd395109724e4db5da7a28b9aa6

Observation 9f3dcdbe-99c7-4bea-9254-fdd19144a1cb · inbound

TensorShield: Safeguarding On-Device Inference by Shielding Critical DNN Tensors with TEE cites this paper.

TensorShield: Safeguarding On-Device Inference by Shielding Critical DNN Tensors with TEE PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:01.749848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:01.749848Z digest=sha256:28e18c654dbaec54682dee9688d214be229da58860651179e6a85e1e2bfcd112

Observation 6491a2a5-891d-4668-ac7c-12812760b38b · inbound

Small Language Models are the Future of Agentic AI cites this paper.

Small Language Models are the Future of Agentic AI PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:55:51.043385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T11:55:50.897500Z digest=sha256:af7d4f0edf175479e668985639d818a27c9bb6f83f87fdee6f98d3bad3fd4f04

Observation feb5f522-870d-4ed0-a219-fdf23f4847e2 · inbound

MoQAE: Mixed-Precision Quantization for Long-Context LLM Inference via Mixture of Quantization-Aware Experts cites this paper.

MoQAE: Mixed-Precision Quantization for Long-Context LLM Inference via Mixture of Quantization-Aware Experts PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:52.808116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:37:52.808116Z digest=sha256:940a3c13181f939bf96f8fb1d6a9e774f7bd34362dbf766e659af5b23d0d197c

Observation 6af74b8a-00e2-4fe4-9dec-a48a499e844f · inbound

DIVE into MoE: Diversity-Enhanced Reconstruction of Large Language Models from Dense into Mixture-of-Experts cites this paper.

DIVE into MoE: Diversity-Enhanced Reconstruction of Large Language Models from Dense into Mixture-of-Experts PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T04:59:18.087086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:59:18.087086Z digest=sha256:d231b71bb3b3e387deaa6851995755cee045d7b7a7f20b0c18434eeaa182f3d1

Observation dfe72331-8322-4198-ad01-cb3db0095e9a · inbound

MNN-LLM: A Generic Inference Engine for Fast Large Language Model Deployment on Mobile Devices cites this paper.

MNN-LLM: A Generic Inference Engine for Fast Large Language Model Deployment on Mobile Devices PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:41.493324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:41.493324Z digest=sha256:afab3949b6af706195bdf8d50450c2b464fb5cd674c14315f98a406d5d3d8dad

Observation 04ba6675-87b1-4c23-b053-d03152eb5e06 · inbound

SparseLoRA: Accelerating LLM Fine-Tuning with Contextual Sparsity cites this paper.

SparseLoRA: Accelerating LLM Fine-Tuning with Contextual Sparsity PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T19:30:31.172365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:30:31.172365Z digest=sha256:d623a96611232aa2c19a45d50ed98989c12e03e3c4c590f803ac2d8ed4182ca8

Observation 53e72b54-46d4-42fd-8b28-627f600ef4c0 · inbound

MNN-AECS: Energy Optimization for LLM Decoding on Mobile Devices via Adaptive Core Selection cites this paper.

MNN-AECS: Energy Optimization for LLM Decoding on Mobile Devices via Adaptive Core Selection PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T18:40:01.498224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:40:01.498224Z digest=sha256:d4be76f1d21431f8191f4b076ad3345af6818165e26894e70e371caf9c99fba6

Observation 53927732-1194-40e3-ba29-008579fe1cf7 · inbound

Dissecting the Impact of Mobile DVFS Governors on LLM Inference Performance and Energy Efficiency cites this paper.

Dissecting the Impact of Mobile DVFS Governors on LLM Inference Performance and Energy Efficiency PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T20:44:17.762246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:44:17.762246Z digest=sha256:0594c916439108611875db5e03a1419cca9feb7e1ba8b4dad4ab4ebef3c8abca

Observation 0b428182-fc7c-427b-a729-c682db3579a5 · inbound

Efficient Deployment of Vision-Language Models on Mobile Devices: A Case Study on OnePlus 13R cites this paper.

Efficient Deployment of Vision-Language Models on Mobile Devices: A Case Study on OnePlus 13R PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:53.546644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:20:53.546644Z digest=sha256:53cb2dff7ac42fd1ce956ede688ff561345749fea1fd3a18875f44da3d393e27

Observation c3bd469b-6ec3-40fb-a33e-85f97cc595f6 · inbound

BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity cites this paper.

BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:00.457313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:20:00.457313Z digest=sha256:6e3f84876271151b9fd0d9036aeb546acfdab14bbcd127fab2599bcdae793a23

Observation dbb96f30-34c4-438f-addc-dc9902738ff9 · inbound

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning cites this paper.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:21.155515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:21.155515Z digest=sha256:2d4d3693db5cd89c4b2dbc23aee9c62498a3a0e266d4a6a7fe6511832921b97b

Observation ab19481e-dffc-40bc-96bd-7084443c15c4 · inbound

ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference cites this paper.

ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:06:52.141041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T22:03:10.316005Z digest=sha256:d00175d9f7c1ab87b3e5f00cb9449ef629da4a409f29f7bb15faa8d465523371

Observation 8ad5e15e-bae3-48a4-968e-90f9a3b25f4d · inbound

WISP: Waste- and Interference-Suppressed Distributed Speculative LLM Serving at the Edge via Dynamic Drafting and SLO-Aware Batching cites this paper.

WISP: Waste- and Interference-Suppressed Distributed Speculative LLM Serving at the Edge via Dynamic Drafting and SLO-Aware Batching PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:12:58.650668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T14:12:06.034679Z digest=sha256:1cf14f730cfb47233aa9e5cc0ba1fbd2604e29218cd014c61088afbb866d34c8

Observation 61e88517-3136-4dab-8730-5c9978488b2c · inbound

Understanding User Privacy Perceptions of GenAI Smartphones cites this paper.

Understanding User Privacy Perceptions of GenAI Smartphones PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:10:51.598346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T19:18:29.955515Z digest=sha256:8081c6eefa2a6ce7eef3dc510427f32ac9f0b98fe66c7bfee9a8cf32be700021

Observation a3e0deb1-2816-4836-8266-7be16572cec3 · inbound

MCAP: Deployment-Time Layer Profiling for Memory-Constrained LLM Inference cites this paper.

MCAP: Deployment-Time Layer Profiling for Memory-Constrained LLM Inference PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:54:48.401477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T00:54:33.897112Z digest=sha256:84cf219619292c99b8899514aedc3aba8702527c78e5a6acb4ac70f38c8d4146

Observation ac7eb187-db2c-43fb-85a5-6fe3d9f45801 · inbound

EdgeFM: Efficient Edge Inference for Vision-Language Models cites this paper.

EdgeFM: Efficient Edge Inference for Vision-Language Models PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:01:27.717820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-07T08:30:35.264804Z digest=sha256:530605b013d674fd8159ed1d05b165d4792efaed3d092de9585bd029de1aaeaf

Observation 5caa659d-b24b-4210-823e-f444f7ed121d · inbound

EdgeFM: Efficient Edge Inference for Vision-Language Models cites this paper.

EdgeFM: Efficient Edge Inference for Vision-Language Models PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:55:34.742443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-01T08:54:41.845820Z digest=sha256:6415822b8df329b2e0bd350f81112ff94d9e78837bb0b94f8e63e0e3dd7d0caa

Observation 1d99e954-1e6b-495a-b3a5-5857d15d1b5d · inbound

Lever: Speculative LLM Inference on Smartphones cites this paper.

Lever: Speculative LLM Inference on Smartphones PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-19T21:47:48.483707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:3518052b2f23c7ec92066dd3038e93c9464679e4abe509a48d5e1da444b78cd4

Observation e6e60fec-694c-40db-a75a-e7b687591a5c · inbound

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference cites this paper.

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:04:46.182427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T15:03:31.289211Z digest=sha256:bbe08803c65a76e08691c223a39731ac42300f12461c88a37f2a5875d9510004

Observation af2b36e5-e4b7-4408-bc44-cfb1e0fbb080 · inbound

TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUs cites this paper.

TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUs PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T07:57:44.941264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T11:26:56.230283Z digest=sha256:5bad027371ca2ff9700b827d6db541e268fff59493589a909f2fe253964c331f

Observation 1a444fdc-ed0e-44c7-a1b4-b1cc05198f42 · inbound

TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUs cites this paper.

TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUs PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:39:02.905070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-03T23:37:25.391853Z digest=sha256:030eb33dc360dfe514008ddd7608bbd62d254ddeaef42804e2533e14a9a38d01

Observation 071dec5c-0a4d-49ed-b77c-d6b98206327d · inbound

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study cites this paper.

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:39.607441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T12:30:55.628115Z digest=sha256:1eaa84d4b13af95dcc658193942b000282f3d806c8c51b58ea747d88f1c0192c

Observation 38bbb818-13f7-46c1-afa4-c6412830500d · inbound

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study cites this paper.

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-12T13:05:17.273287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:05:17.273287Z digest=sha256:5570af44701727b9db13c0e6ed843a1f52570bc700acd5650394d3732191bce1

Observation a4708dba-eaae-4573-9c1b-af9c4fd2dc6b · inbound

Enabling Cloud-Level Accuracy in Edge AI through IoT Data Preprocessing cites this paper.

Enabling Cloud-Level Accuracy in Edge AI through IoT Data Preprocessing PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:49:44.091015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T09:34:00.058213Z digest=sha256:e820aa567425666426e380a9fa9033091b1abb791bf536b51670bf4df498fc41

Observation 02212a14-d014-4c2e-9473-8d7696156391 · inbound

EnerInfer: Energy-Aware On-Device LLM Inference cites this paper.

EnerInfer: Energy-Aware On-Device LLM Inference PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:29:50.714252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T07:58:29.150228Z digest=sha256:bf70451d54f0081d36d00357e0c71eec5aafc04fcd1e09210975ef1586d7d387

Observation 9cd432d8-7f33-45ec-a4fb-8fb12919a0e3 · inbound

Is Your NPU Ready for LLMs? Dissecting the Hidden Efficiency Bottlenecks in Mobile LLM Inference cites this paper.

Is Your NPU Ready for LLMs? Dissecting the Hidden Efficiency Bottlenecks in Mobile LLM Inference PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-11T12:14:57.342551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T12:14:57.342551Z digest=sha256:bb0121d7b2b74939b2ea6581a7d4e4edebb8d9d864eca15dd38cc334899d2ce2

Observation 868c65d1-9304-493d-af0e-5d65aefd8baf · inbound

Voltron: Enabling Elastic Multi-Device Execution of LLM Inference for Empowered Edge Intelligence cites this paper.

Voltron: Enabling Elastic Multi-Device Execution of LLM Inference for Empowered Edge Intelligence PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-07-09T21:16:34.319320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T21:08:24.293077Z digest=sha256:cbf0d14ce927d4a6e9b82bd18f086fc2113cc84c348aa708f82a6e13c3d6d8ae

Observation 6db40602-1fd7-4e89-b1d1-26327047031e · inbound

HeteroMosaic: Exposing and Exploiting Heterogeneous Execution Opportunities for Energy-Efficient Edge LLM Inference cites this paper.

HeteroMosaic: Exposing and Exploiting Heterogeneous Execution Opportunities for Energy-Efficient Edge LLM Inference PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-02T06:22:22.488014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:22:22.488014Z digest=sha256:9c4a4d447182dfd2c9500e8bc473893b49c47ad9b29d1e5d595ccd1b4286acab

Observation 2c38e300-665a-4aef-a38d-3b57963224b7 · inbound

Transition-Aware Backend Dispatch for Edge LLM Inference cites this paper.

Transition-Aware Backend Dispatch for Edge LLM Inference PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:57.453503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:57.453503Z digest=sha256:f04d16aa3e825c8f4e115c303bc8403c5202a7f18d3179b406289d84c8dbb711

Observation 88c1eab1-7bb5-4783-88e3-211bbbeb280d · inbound

SelectInfer: Selective Neuron Loading and Computation for On-Device LLMs cites this paper.

SelectInfer: Selective Neuron Loading and Computation for On-Device LLMs PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T16:14:37.007544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:14:37.007544Z digest=sha256:0c55731b5454fee63722671531ca4f077e2bdff3fa0eb24298b14987faa7915f

Observation 57cc3a2f-510a-4154-936f-42009b093336 · inbound

SparSEEty: Extracting Tokens from Sparsity-Exploiting LLM Serving Systems via Deterministic Side Channels cites this paper.

SparSEEty: Extracting Tokens from Sparsity-Exploiting LLM Serving Systems via Deterministic Side Channels PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T04:21:30.070796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:21:30.070796Z digest=sha256:14a915e43965fa781b830db3dcaac2d84f27d6afe469be91172462e4dde4424b

Observation 38c75e83-06e6-482d-a605-704dc4492a13 · inbound

The Ingestion Tax: Adopting File-Backed Weights in Tensor Frameworks cites this paper.

The Ingestion Tax: Adopting File-Backed Weights in Tensor Frameworks PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T00:22:44.065594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:22:44.065594Z digest=sha256:437f950dd230adea50fe3c22ab3ea7e12799e0a1d7433623dd3cda8ca6f8b861