Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T04:51:31.541760Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 1 inbound Pith citation observation for arXiv:2505.00232.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T04:51:31.541760Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:57:30.372775Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-15T20:57:30.733875Z
59 of 59 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8a6ea8cf-7f83-451e-b584-6b4e92d8fb08 · outbound
Scaling On-Device GPU Inference for Large Generative Models The Khronos Group Inc., 2019
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 84050aef-ddba-42d5-b04b-ecb0c3593d90 · outbound
Scaling On-Device GPU Inference for Large Generative Models The Khronos Group Inc., 2025
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 64e918cc-7901-4b36-ac95-2d2349573ec9 · outbound
Scaling On-Device GPU Inference for Large Generative Models AMD ROCm Soft- ware
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 35619c9e-5e6b-4bf3-80c4-cf0862e6fb69 · outbound
Scaling On-Device GPU Inference for Large Generative Models LLM in a flash: Effi- cient Large Language Model Inference with Limited Mem- ory, 2024
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 50f83bf4-3208-4146-9e8f-0d8a1022d446 · outbound
Scaling On-Device GPU Inference for Large Generative Models Core ML Stable Diffusion
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1600da46-e464-4788-8bf6-58925191c0e6 · outbound
Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 03b334a5-6335-4ae8-a223-c2118fe97de6 · outbound
Scaling On-Device GPU Inference for Large Generative Models Compute Library
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3a2aac5d-634a-4b7f-9068-6a346dbb38e6 · outbound
Scaling On-Device GPU Inference for Large Generative Models TVM: An automated End-to-End optimizing com- piler for deep learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ea88d8f3-31d9-4065-b0a5-4af7c03d8324 · outbound
Scaling On-Device GPU Inference for Large Generative Models Speed Is All You Need: On-Device Acceleration of Large Diffusion Models via GPU-Aware Optimizations, 2023
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1bb45ed1-17a1-458f-8cd6-cb1221f0cc8f · outbound
Scaling On-Device GPU Inference for Large Generative Models LMDeploy: A Toolkit for Com- pressing, Deploying, and Serving LLM
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a02154a5-1f7c-46cc-85ef-d54b8eef9de5 · outbound
Scaling On-Device GPU Inference for Large Generative Models Fu, Stefano Ermon, Atri Rudra, and Christopher R´e
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d1287604-d0ae-4cfc-aa76-73f52a5097b6 · outbound
Scaling On-Device GPU Inference for Large Generative Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 677ff3df-9209-451b-a302-346cbdb2a438 · outbound
Scaling On-Device GPU Inference for Large Generative Models GPTQ: Accurate Post-Training Quantization for Gen- erative Pre-trained Transformers, 2023
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 02d036d3-af8f-4eee-b43d-c27a107cd2d7 · outbound
Scaling On-Device GPU Inference for Large Generative Models Gemma 2: Improving Open Language Mod- els at a Practical Size, 2024
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c7015efd-80bd-4314-9f32-01401939d2b7 · outbound
Scaling On-Device GPU Inference for Large Generative Models Gemma: Open Models Based on Gemini Research and Technology, 2024
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0aea38be-4609-4085-89ad-90ae866ed3ee · outbound
Scaling On-Device GPU Inference for Large Generative Models llama.cpp
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6e8b1186-d337-49c2-8e95-ddfdba6a5788 · outbound
Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation be793fe5-21cb-4c21-858d-104d7958d003 · outbound
Scaling On-Device GPU Inference for Large Generative Models LiteRT Overview
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4f96f6ab-f0fe-43a4-a47c-bdc7b9ba331f · outbound
Scaling On-Device GPU Inference for Large Generative Models Huawei HiAI
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4c372e43-57e6-4f94-a62a-e31759ed6572 · outbound
Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a88a2194-250a-488e-9509-80693c5284fe · outbound
Scaling On-Device GPU Inference for Large Generative Models OpenVINO
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9ac07930-d5b7-4318-8149-b9713f0d4c83 · outbound
Scaling On-Device GPU Inference for Large Generative Models Intel Core Ultra Series 2 Media Deck
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 64a14253-3f9a-4fb7-ba60-b5f0216bb40d · outbound
Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5171003c-2e67-4fdb-8567-d34c505fe892 · outbound
Scaling On-Device GPU Inference for Large Generative Models MNN: A Universal and Efficient Inference Engine
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f28ca99f-b57a-461b-8dde-39ac0ab86c1f · outbound
Scaling On-Device GPU Inference for Large Generative Models Gonzalez, Hao Zhang, and Ion Stoica
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 01410587-29d5-48ee-a2fe-8319d24df45f · outbound
Scaling On-Device GPU Inference for Large Generative Models On-Device Neu- ral Net Inference with Mobile GPUs
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6ac4ea48-9622-4f7b-92d0-760a9b383d46 · outbound
Scaling On-Device GPU Inference for Large Generative Models OpenGL ES Version 3.1
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b7b87274-5c91-466f-82f3-535ef56f05ce · outbound
Scaling On-Device GPU Inference for Large Generative Models Fast In- ference from Transformers via Speculative Decoding, 2023
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 48103cb2-cd68-40b7-8954-9ca3ef371738 · outbound
Scaling On-Device GPU Inference for Large Generative Models AWQ: Activation-aware Weight Quantization for LLM Compression and Accelera- tion
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d5a778d6-56d8-4ef9-88ce-5d00b8bba3a9 · outbound
Scaling On-Device GPU Inference for Large Generative Models The Llama 3 Herd of Models, 2024
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 39867c08-cad7-418b-bd15-3f6fba21ecca · outbound
Scaling On-Device GPU Inference for Large Generative Models DeepMon: Mobile GPU-based Deep Learning Framework for Continuous Vision Applications
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bee3c6e9-0c36-4cf2-8a12-31c50c1489ac · outbound
Scaling On-Device GPU Inference for Large Generative Models NeuroPilot
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c3ef5adf-bf8c-4145-8251-dab241f281a6 · outbound
Scaling On-Device GPU Inference for Large Generative Models ExecuTorch
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ac7b1861-7567-46ab-9388-0652468d864f · outbound
Scaling On-Device GPU Inference for Large Generative Models Llama 3.2: Revolutionizing edge AI and vision with open, customizable models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c0838d42-45f4-4b9e-a3a1-48effb96e14d · outbound
Scaling On-Device GPU Inference for Large Generative Models DirectML Overview
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation af0ad225-015b-4b94-b423-7c10975a72c6 · outbound
Scaling On-Device GPU Inference for Large Generative Models Get started with ONNX Run- time Mobile
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d0268f61-6ea7-4c56-9c13-91e0b67c767b · outbound
Scaling On-Device GPU Inference for Large Generative Models Stable Diffusion Op- timization with DirectML
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bfa471c6-016f-43ce-96da-0bb3acfc7576 · outbound
Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 67262328-6b4d-4379-a936-70c921ad655a · outbound
Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 96c5db1d-d03b-4196-b9ad-5167ee423445 · outbound
Scaling On-Device GPU Inference for Large Generative Models NVIDIA Tensor Cores
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f4b38987-3aeb-4e27-a84d-b1d2b3938805 · outbound
Scaling On-Device GPU Inference for Large Generative Models TensorRT
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 17bc6ceb-1dde-4cc4-9ee9-81596a10cf25 · outbound
Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2572a02f-e45d-4571-9175-ec554570753e · outbound
Scaling On-Device GPU Inference for Large Generative Models Efficient Memory Manage- ment for Deep Neural Net Inference
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c1f17b16-a16d-4582-99ac-55d254b36740 · outbound
Scaling On-Device GPU Inference for Large Generative Models Snapdragon Neural Processing Engine SDK
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation cfda3ff7-68cc-4da5-914f-db7745bd842e · outbound
Scaling On-Device GPU Inference for Large Generative Models QualComm AI Hub Llama-v3.2-3B-Chat
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5e70183f-a43b-4b95-81a6-02a97d5caadb · outbound
Scaling On-Device GPU Inference for Large Generative Models World’s first on-device demonstration of Stable Diffusion on an Android phone
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9ed90683-f398-487f-a7d3-b88063eda807 · outbound
Scaling On-Device GPU Inference for Large Generative Models Ruan, Yucheng Qin, Xun Zhou, Ruihang Lai, Hongyi Jin, Yixin Dong, Bohan Hou, Meng-Shiun Yu, Yiyan Zhai, Sudeep Agarwal, Hangrui Cao, Siyuan Feng, and Tianqi Chen
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f547c7bc-5890-481e-b1d7-8b8c3decf77e · outbound
Scaling On-Device GPU Inference for Large Generative Models XLA: Compiling Machine Learning for Peak Performance, 2020
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 10aef569-a863-4b62-a5e4-f42a03f003c5 · outbound
Scaling On-Device GPU Inference for Large Generative Models Introducing Stable Diffusion 3.5
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7fa999f2-5162-44ad-a18e-57d302ee4093 · outbound
Scaling On-Device GPU Inference for Large Generative Models Introducing torchchat: Accelerating Local LLM Inference on Laptop, Desktop and Mobile
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8524c447-e6d9-4edc-9d07-816c3c206f37 · outbound
Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6adc8945-0d67-42f6-8bbd-27b2fb14f743 · outbound
Scaling On-Device GPU Inference for Large Generative Models Dawn, a WebGPU implemen- tation
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 050b6cd3-07d8-4802-aa26-7d3b9057637f · outbound
Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3b55a910-0d4c-4677-8cf2-78534dd35c7e · outbound
Scaling On-Device GPU Inference for Large Generative Models SmoothQuant: Accurate and Effi- cient Post-Training Quantization for Large Language Mod- els
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4ebfc85a-c516-44d0-8f83-2390397b2c97 · outbound
Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4c0350f0-bd82-43d1-a93e-54649577265e · outbound
Scaling On-Device GPU Inference for Large Generative Models LLMCad: Fast and Scalable On-device Large Language Model Inference, 2023
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d73f38fa-cffa-4fa2-b165-fbec16e9cc78 · outbound
Scaling On-Device GPU Inference for Large Generative Models Fast On-device LLM Inference with NPUs
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 339db099-dd19-46ef-be92-964ead0a53a3 · outbound
Scaling On-Device GPU Inference for Large Generative Models PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 984377af-19f7-4a48-9783-9de59c1de3fc · outbound
Scaling On-Device GPU Inference for Large Generative Models Gonzalez, Clark Barrett, and Ying Sheng
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 035f8d65-6761-4107-896c-35502c5af71f · inbound
EdgeWisePersona: A Dataset for On-Device User Profiling from Natural Language Interactions Scaling On-Device GPU Inference for Large Generative Models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.