Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 42 inbound Pith citation observations for arXiv:2406.06282.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T05:09:55.031898Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
5
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 456d174e-001a-4ade-b2fd-533489973d47 · inbound
BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 132
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c7d2330-15c8-4845-b7b3-f9906a84d1b9 · inbound
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0aac06a-5260-43f4-bdb5-24caff9e8358 · inbound
Densing Law of LLMs PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c569c36-6cf0-4a3b-8347-14970b0e2415 · inbound
IC-Cache: Efficient Large Language Model Serving via In-context Caching PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73483b14-f398-4fdc-ae5e-94611d77db4d · inbound
iServe: An Intent-based Serving System for LLMs PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 120
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 339db099-dd19-46ef-be92-964ead0a53a3 · inbound
Scaling On-Device GPU Inference for Large Generative Models PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b03021e-085b-46b0-b834-2a12b430ecde · inbound
Responsive DNN Adaptation for Video Analytics against Environment Shift via Hierarchical Mobile-Cloud Collaborations PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46a15fe9-ec8e-41b3-98e1-6d898f5b7c49 · inbound
RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 111
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fc9a2071-2c3c-4899-ad54-d6cbf007e740 · inbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f5c1124-bece-4e9b-a235-fd754a1aaa5e · inbound
CoreMatching: A Co-adaptive Sparse Inference Framework with Token and Neuron Pruning for Comprehensive Acceleration of Vision-Language Models PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f3dcdbe-99c7-4bea-9254-fdd19144a1cb · inbound
TensorShield: Safeguarding On-Device Inference by Shielding Critical DNN Tensors with TEE PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6491a2a5-891d-4668-ac7c-12812760b38b · inbound
Small Language Models are the Future of Agentic AI PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation feb5f522-870d-4ed0-a219-fdf23f4847e2 · inbound
MoQAE: Mixed-Precision Quantization for Long-Context LLM Inference via Mixture of Quantization-Aware Experts PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6af74b8a-00e2-4fe4-9dec-a48a499e844f · inbound
DIVE into MoE: Diversity-Enhanced Reconstruction of Large Language Models from Dense into Mixture-of-Experts PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfe72331-8322-4198-ad01-cb3db0095e9a · inbound
MNN-LLM: A Generic Inference Engine for Fast Large Language Model Deployment on Mobile Devices PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04ba6675-87b1-4c23-b053-d03152eb5e06 · inbound
SparseLoRA: Accelerating LLM Fine-Tuning with Contextual Sparsity PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53e72b54-46d4-42fd-8b28-627f600ef4c0 · inbound
MNN-AECS: Energy Optimization for LLM Decoding on Mobile Devices via Adaptive Core Selection PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53927732-1194-40e3-ba29-008579fe1cf7 · inbound
Dissecting the Impact of Mobile DVFS Governors on LLM Inference Performance and Energy Efficiency PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b428182-fc7c-427b-a729-c682db3579a5 · inbound
Efficient Deployment of Vision-Language Models on Mobile Devices: A Case Study on OnePlus 13R PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3bd469b-6ec3-40fb-a33e-85f97cc595f6 · inbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbb96f30-34c4-438f-addc-dc9902738ff9 · inbound
MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab19481e-dffc-40bc-96bd-7084443c15c4 · inbound
ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8ad5e15e-bae3-48a4-968e-90f9a3b25f4d · inbound
WISP: Waste- and Interference-Suppressed Distributed Speculative LLM Serving at the Edge via Dynamic Drafting and SLO-Aware Batching PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 61e88517-3136-4dab-8730-5c9978488b2c · inbound
Understanding User Privacy Perceptions of GenAI Smartphones PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a3e0deb1-2816-4836-8266-7be16572cec3 · inbound
MCAP: Deployment-Time Layer Profiling for Memory-Constrained LLM Inference PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ac7eb187-db2c-43fb-85a5-6fe3d9f45801 · inbound
EdgeFM: Efficient Edge Inference for Vision-Language Models PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5caa659d-b24b-4210-823e-f444f7ed121d · inbound
EdgeFM: Efficient Edge Inference for Vision-Language Models PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1d99e954-1e6b-495a-b3a5-5857d15d1b5d · inbound
Lever: Speculative LLM Inference on Smartphones PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e6e60fec-694c-40db-a75a-e7b687591a5c · inbound
When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation af2b36e5-e4b7-4408-bc44-cfb1e0fbb080 · inbound
TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUs PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1a444fdc-ed0e-44c7-a1b4-b1cc05198f42 · inbound
TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUs PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 071dec5c-0a4d-49ed-b77c-d6b98206327d · inbound
Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 38bbb818-13f7-46c1-afa4-c6412830500d · inbound
Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4708dba-eaae-4573-9c1b-af9c4fd2dc6b · inbound
Enabling Cloud-Level Accuracy in Edge AI through IoT Data Preprocessing PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 02212a14-d014-4c2e-9473-8d7696156391 · inbound
EnerInfer: Energy-Aware On-Device LLM Inference PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9cd432d8-7f33-45ec-a4fb-8fb12919a0e3 · inbound
Is Your NPU Ready for LLMs? Dissecting the Hidden Efficiency Bottlenecks in Mobile LLM Inference PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 868c65d1-9304-493d-af0e-5d65aefd8baf · inbound
Voltron: Enabling Elastic Multi-Device Execution of LLM Inference for Empowered Edge Intelligence PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6db40602-1fd7-4e89-b1d1-26327047031e · inbound
HeteroMosaic: Exposing and Exploiting Heterogeneous Execution Opportunities for Energy-Efficient Edge LLM Inference PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c38e300-665a-4aef-a38d-3b57963224b7 · inbound
Transition-Aware Backend Dispatch for Edge LLM Inference PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88c1eab1-7bb5-4783-88e3-211bbbeb280d · inbound
SelectInfer: Selective Neuron Loading and Computation for On-Device LLMs PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57cc3a2f-510a-4154-936f-42009b093336 · inbound
SparSEEty: Extracting Tokens from Sparsity-Exploiting LLM Serving Systems via Deterministic Side Channels PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38c75e83-06e6-482d-a605-704dc4492a13 · inbound
The Ingestion Tax: Adopting File-Backed Weights in Tensor Frameworks PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.