Pith. sign in

Paper Citation Record · LEDGER

Scaling On-Device GPU Inference for Large Generative Models

As of 17 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 1 inbound Pith citation observation for arXiv:2505.00232.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.00232 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:51:31.541760Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:57:30.372775Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T20:57:30.733875Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy46
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8a6ea8cf-7f83-451e-b584-6b4e92d8fb08 · outbound

This paper cites The Khronos Group Inc., 2019.

Scaling On-Device GPU Inference for Large Generative Models The Khronos Group Inc., 2019

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.226757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.326199Z digest=sha256:fb3dee9563137ff4520806e870cc47ed0fa8594df5b3f97669d39d784fe5ce11

Observation 84050aef-ddba-42d5-b04b-ecb0c3593d90 · outbound

This paper cites The Khronos Group Inc., 2025.

Scaling On-Device GPU Inference for Large Generative Models The Khronos Group Inc., 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.216060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.330309Z digest=sha256:ecc9a69e8977f1a286d3a121d3ea2669fbaad86b7a66957dc69e2e4b570db587

Observation 64e918cc-7901-4b36-ac95-2d2349573ec9 · outbound

This paper cites AMD ROCm Soft- ware.

Scaling On-Device GPU Inference for Large Generative Models AMD ROCm Soft- ware

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.205450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.333808Z digest=sha256:60881da2147a5136632805633f61003955a9e14b93dfb8306fe5d7b99098c7c2

Observation 35619c9e-5e6b-4bf3-80c4-cf0862e6fb69 · outbound

This paper cites LLM in a flash: Effi- cient Large Language Model Inference with Limited Mem- ory, 2024.

Scaling On-Device GPU Inference for Large Generative Models LLM in a flash: Effi- cient Large Language Model Inference with Limited Mem- ory, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.195166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.337188Z digest=sha256:79b11c4bfb0800436e9f8f520e0d00e1ce28e6fdf023ce8ecf0438d0e57c40ea

Observation 50f83bf4-3208-4146-9e8f-0d8a1022d446 · outbound

This paper cites Core ML Stable Diffusion.

Scaling On-Device GPU Inference for Large Generative Models Core ML Stable Diffusion

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.184617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.340853Z digest=sha256:4f0f200f0bcf06e61e4f70a8a8b9c4399418024f59969f51321ece8af232a4f1

Observation 1600da46-e464-4788-8bf6-58925191c0e6 · outbound

This paper cites an unresolved cited work.

Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:51:32.174494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.343987Z digest=sha256:1435ddb5ac7c5178ce935cefd4c17778d0510f99d6aae4de19d9721019d6731d

Observation 03b334a5-6335-4ae8-a223-c2118fe97de6 · outbound

This paper cites Compute Library.

Scaling On-Device GPU Inference for Large Generative Models Compute Library

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.163667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.347636Z digest=sha256:874657483b6efa1f4bf66c7035096fef681725147af2c32ada8ca17f5787023c

Observation 3a2aac5d-634a-4b7f-9068-6a346dbb38e6 · outbound

This paper cites TVM: An automated End-to-End optimizing com- piler for deep learning.

Scaling On-Device GPU Inference for Large Generative Models TVM: An automated End-to-End optimizing com- piler for deep learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.153440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.350756Z digest=sha256:b7297c36d307a1c6f22583bfe06f9e7f91e05d70d59027201a42a250b7bd4d91

Observation ea88d8f3-31d9-4065-b0a5-4af7c03d8324 · outbound

This paper cites Speed Is All You Need: On-Device Acceleration of Large Diffusion Models via GPU-Aware Optimizations, 2023.

Scaling On-Device GPU Inference for Large Generative Models Speed Is All You Need: On-Device Acceleration of Large Diffusion Models via GPU-Aware Optimizations, 2023

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.143000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.353895Z digest=sha256:190a8cf43162dfdf5705923298bec2a79572c1646f5e9d76bd21b4d0604d2bd6

Observation 1bb45ed1-17a1-458f-8cd6-cb1221f0cc8f · outbound

This paper cites LMDeploy: A Toolkit for Com- pressing, Deploying, and Serving LLM.

Scaling On-Device GPU Inference for Large Generative Models LMDeploy: A Toolkit for Com- pressing, Deploying, and Serving LLM

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.132514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.356758Z digest=sha256:841e50d5c7121c5c1d266ef535ee15c88edee63ce6bc5d2f5f5b22bd3e92c69d

Observation a02154a5-1f7c-46cc-85ef-d54b8eef9de5 · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher R´e.

Scaling On-Device GPU Inference for Large Generative Models Fu, Stefano Ermon, Atri Rudra, and Christopher R´e

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.122766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.359871Z digest=sha256:8069ff74b335ded78c25605c2d29a892101b666926b8ce801a8558a4abea9f55

Observation d1287604-d0ae-4cfc-aa76-73f52a5097b6 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Scaling On-Device GPU Inference for Large Generative Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T04:51:31.363141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:51:31.363141Z digest=sha256:1c9687a7fde35da599c9f59410a56174a8ed796748f4cc636ca10b7c4e51454c

Observation 677ff3df-9209-451b-a302-346cbdb2a438 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Gen- erative Pre-trained Transformers, 2023.

Scaling On-Device GPU Inference for Large Generative Models GPTQ: Accurate Post-Training Quantization for Gen- erative Pre-trained Transformers, 2023

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.111893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.367086Z digest=sha256:ee30bd280c0f66e0f7f2d54ce10ecdcb38f9068a0580017092a326bc4b7f2461

Observation 02d036d3-af8f-4eee-b43d-c27a107cd2d7 · outbound

This paper cites Gemma 2: Improving Open Language Mod- els at a Practical Size, 2024.

Scaling On-Device GPU Inference for Large Generative Models Gemma 2: Improving Open Language Mod- els at a Practical Size, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.099149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.370650Z digest=sha256:9c72469edb55a6fc0f9f92836e456f20efd32b7b8383833507eefaa3f94f592f

Observation c7015efd-80bd-4314-9f32-01401939d2b7 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology, 2024.

Scaling On-Device GPU Inference for Large Generative Models Gemma: Open Models Based on Gemini Research and Technology, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.085958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.374881Z digest=sha256:4eade257f55bd2a2500401c69a31801a5b37bfb455ed1f7e715c36aaeedb5e9c

Observation 0aea38be-4609-4085-89ad-90ae866ed3ee · outbound

This paper cites llama.cpp.

Scaling On-Device GPU Inference for Large Generative Models llama.cpp

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.075012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.379167Z digest=sha256:80cd84be5fe3854312c6c5418e7b7f3f7efc6e3b06077fc7ca995d92ca890d6c

Observation 6e8b1186-d337-49c2-8e95-ddfdba6a5788 · outbound

This paper cites an unresolved cited work.

Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:51:32.064032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.383046Z digest=sha256:b3a25b7f83069a5661b2592f2216695dfb76ebce46101f3e3fd63a939210384a

Observation be793fe5-21cb-4c21-858d-104d7958d003 · outbound

This paper cites LiteRT Overview.

Scaling On-Device GPU Inference for Large Generative Models LiteRT Overview

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.052536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.387319Z digest=sha256:adb63c6dc778701e4ef024a61668f453d122f22d484aa297ddfb0323af5f695e

Observation 4f96f6ab-f0fe-43a4-a47c-bdc7b9ba331f · outbound

This paper cites Huawei HiAI.

Scaling On-Device GPU Inference for Large Generative Models Huawei HiAI

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.038995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.392090Z digest=sha256:336ab7b2466b7d4dd6e178ad56c3981d43a61492870baa891898029270082b86

Observation 4c372e43-57e6-4f94-a62a-e31759ed6572 · outbound

This paper cites an unresolved cited work.

Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:51:32.029100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.396993Z digest=sha256:e0d5a084852ef227bc2c0e58da1ca4451de8f76b4c9c7925297ee3b6f05f514e

Observation a88a2194-250a-488e-9509-80693c5284fe · outbound

This paper cites OpenVINO.

Scaling On-Device GPU Inference for Large Generative Models OpenVINO

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.018800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.401664Z digest=sha256:a1638a172792c10f40b13d3b5d74dfb3115a8dbcbdddf990a1bfb708ef830b41

Observation 9ac07930-d5b7-4318-8149-b9713f0d4c83 · outbound

This paper cites Intel Core Ultra Series 2 Media Deck.

Scaling On-Device GPU Inference for Large Generative Models Intel Core Ultra Series 2 Media Deck

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.007819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.405781Z digest=sha256:3eeb8f6e8a3c24f1a27b721f8f047974878e1d7045fd9556b81aec739b589e44

Observation 64a14253-3f9a-4fb7-ba60-b5f0216bb40d · outbound

This paper cites an unresolved cited work.

Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:51:31.996993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.411970Z digest=sha256:18dc1e80a9e5d8ccc58abc65181e9c4d984082e3f39c3638e51807b8b2a615e4

Observation 5171003c-2e67-4fdb-8567-d34c505fe892 · outbound

This paper cites MNN: A Universal and Efficient Inference Engine.

Scaling On-Device GPU Inference for Large Generative Models MNN: A Universal and Efficient Inference Engine

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T04:51:31.415946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:51:31.415946Z digest=sha256:e5aaebf974d7fd9a767de6bedfd70208b914e77af4d9994aff4f854c5a199c4f

Observation f28ca99f-b57a-461b-8dde-39ac0ab86c1f · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Scaling On-Device GPU Inference for Large Generative Models Gonzalez, Hao Zhang, and Ion Stoica

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.985623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.420135Z digest=sha256:6dbfadfa638da89e6c8e06a56b7f97a852c682c290dda8a398d8c6a4d5e4977b

Observation 01410587-29d5-48ee-a2fe-8319d24df45f · outbound

This paper cites On-Device Neu- ral Net Inference with Mobile GPUs.

Scaling On-Device GPU Inference for Large Generative Models On-Device Neu- ral Net Inference with Mobile GPUs

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.974737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.424276Z digest=sha256:f4e299b2670920721164da517b2947792b6920e2603f9492d338d5239ca5038a

Observation 6ac4ea48-9622-4f7b-92d0-760a9b383d46 · outbound

This paper cites OpenGL ES Version 3.1.

Scaling On-Device GPU Inference for Large Generative Models OpenGL ES Version 3.1

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.963044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.427827Z digest=sha256:de1c54d0a35852d109c7d768f819d5436b265797ffed1c92a5ffbd06be6dd76e

Observation b7b87274-5c91-466f-82f3-535ef56f05ce · outbound

This paper cites Fast In- ference from Transformers via Speculative Decoding, 2023.

Scaling On-Device GPU Inference for Large Generative Models Fast In- ference from Transformers via Speculative Decoding, 2023

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.951919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.431461Z digest=sha256:0e458cf1fbe1d115cd55b29609fdf1d07f5e25ea7e73bf7d53134ae9d469e6f7

Observation 48103cb2-cd68-40b7-8954-9ca3ef371738 · outbound

This paper cites AWQ: Activation-aware Weight Quantization for LLM Compression and Accelera- tion.

Scaling On-Device GPU Inference for Large Generative Models AWQ: Activation-aware Weight Quantization for LLM Compression and Accelera- tion

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.942289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.435413Z digest=sha256:d746cd92352eddd65758c7abebc69b90a762ead349b0328f1f7a04c703b269f9

Observation d5a778d6-56d8-4ef9-88ce-5d00b8bba3a9 · outbound

This paper cites The Llama 3 Herd of Models, 2024.

Scaling On-Device GPU Inference for Large Generative Models The Llama 3 Herd of Models, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.932892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.438744Z digest=sha256:1b973c7f7e84af193a6dbd74e97648c0d956c8aae41dafea7c9e2c0ac26eef79

Observation 39867c08-cad7-418b-bd15-3f6fba21ecca · outbound

This paper cites DeepMon: Mobile GPU-based Deep Learning Framework for Continuous Vision Applications.

Scaling On-Device GPU Inference for Large Generative Models DeepMon: Mobile GPU-based Deep Learning Framework for Continuous Vision Applications

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.921260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.442284Z digest=sha256:e089866e4f9840d74c7e526f1aa1b93ab7f52dcbd75b1541e546128763a58cd9

Observation bee3c6e9-0c36-4cf2-8a12-31c50c1489ac · outbound

This paper cites NeuroPilot.

Scaling On-Device GPU Inference for Large Generative Models NeuroPilot

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.909923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.445790Z digest=sha256:70d26caff7c99f7119ea1ed14e4cdda1ce528db1d3e84dfa80a137cc5a7ba98c

Observation c3ef5adf-bf8c-4145-8251-dab241f281a6 · outbound

This paper cites ExecuTorch.

Scaling On-Device GPU Inference for Large Generative Models ExecuTorch

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.898276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.449046Z digest=sha256:5d0ce33bf512f3a63e2806826c2f96832ce4348fe5f0c5ab0418c9be9338bbc8

Observation ac7b1861-7567-46ab-9388-0652468d864f · outbound

This paper cites Llama 3.2: Revolutionizing edge AI and vision with open, customizable models.

Scaling On-Device GPU Inference for Large Generative Models Llama 3.2: Revolutionizing edge AI and vision with open, customizable models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.887304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.452475Z digest=sha256:e1af933726fe57dc3d4806f4b93a557ccd5d178c4513506d5475ace4e2574987

Observation c0838d42-45f4-4b9e-a3a1-48effb96e14d · outbound

This paper cites DirectML Overview.

Scaling On-Device GPU Inference for Large Generative Models DirectML Overview

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.876047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.456091Z digest=sha256:f0267f2e9af851360a86569bb7c3aaf79207c55eb6452bfdbd78ea7b21d0627e

Observation af0ad225-015b-4b94-b423-7c10975a72c6 · outbound

This paper cites Get started with ONNX Run- time Mobile.

Scaling On-Device GPU Inference for Large Generative Models Get started with ONNX Run- time Mobile

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.864756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.459461Z digest=sha256:6ceaf10470dbc1406b37e7bb240cb7b8e724938591328699b99c1aafd1893398

Observation d0268f61-6ea7-4c56-9c13-91e0b67c767b · outbound

This paper cites Stable Diffusion Op- timization with DirectML.

Scaling On-Device GPU Inference for Large Generative Models Stable Diffusion Op- timization with DirectML

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.854183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.463215Z digest=sha256:fe429a5bf5b2a32acc8e682c35faba1a95809fd8b56b876d648a194f01a8c86f

Observation bfa471c6-016f-43ce-96da-0bb3acfc7576 · outbound

This paper cites an unresolved cited work.

Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:51:31.842022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.466325Z digest=sha256:40ac25090030df083a371adbef85a054806881b5fca6fc05d0729eca69dd2548

Observation 67262328-6b4d-4379-a936-70c921ad655a · outbound

This paper cites an unresolved cited work.

Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:51:31.828786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.469475Z digest=sha256:53c5d21ec9d7cbeb25fe9b884dad13ce33ae838ffe1f2ab6996898f9fd903bf4

Observation 96c5db1d-d03b-4196-b9ad-5167ee423445 · outbound

This paper cites NVIDIA Tensor Cores.

Scaling On-Device GPU Inference for Large Generative Models NVIDIA Tensor Cores

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.816678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.473074Z digest=sha256:2b1bff9942ca41a800eaf37d1b3fcd42bdea01576d6b8accf0d1e853d28d0403

Observation f4b38987-3aeb-4e27-a84d-b1d2b3938805 · outbound

This paper cites TensorRT.

Scaling On-Device GPU Inference for Large Generative Models TensorRT

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.803472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.475989Z digest=sha256:76271d77d66af87f45797a7e3013a6f5edc1ae31732d4408dff9cd09cc97c2d6

Observation 17bc6ceb-1dde-4cc4-9ee9-81596a10cf25 · outbound

This paper cites an unresolved cited work.

Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:51:31.792711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.479341Z digest=sha256:655667e18d9b57442c68a8078b2ec5759a6e37017d48d0c15a806c942489e86c

Observation 2572a02f-e45d-4571-9175-ec554570753e · outbound

This paper cites Efficient Memory Manage- ment for Deep Neural Net Inference.

Scaling On-Device GPU Inference for Large Generative Models Efficient Memory Manage- ment for Deep Neural Net Inference

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.782707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.482345Z digest=sha256:30e7568884de3483e4c8e05a81742dba7784b928290871d511e765d22d9e6d03

Observation c1f17b16-a16d-4582-99ac-55d254b36740 · outbound

This paper cites Snapdragon Neural Processing Engine SDK.

Scaling On-Device GPU Inference for Large Generative Models Snapdragon Neural Processing Engine SDK

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.771094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.485334Z digest=sha256:9363385dc4eed011ee5ba628f5cf73469e6a43aff98d24386195607ff997a51d

Observation cfda3ff7-68cc-4da5-914f-db7745bd842e · outbound

This paper cites QualComm AI Hub Llama-v3.2-3B-Chat.

Scaling On-Device GPU Inference for Large Generative Models QualComm AI Hub Llama-v3.2-3B-Chat

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.760222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.488945Z digest=sha256:2136925ecec90c50945e9f0ea5cd711123205c3e08ec3194af8d70410c918c86

Observation 5e70183f-a43b-4b95-81a6-02a97d5caadb · outbound

This paper cites World’s first on-device demonstration of Stable Diffusion on an Android phone.

Scaling On-Device GPU Inference for Large Generative Models World’s first on-device demonstration of Stable Diffusion on an Android phone

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.749703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.492697Z digest=sha256:7d0c5903457c01f39863c8dcad5095a4c0185b71309cc1701bbd8664accd9571

Observation 9ed90683-f398-487f-a7d3-b88063eda807 · outbound

This paper cites Ruan, Yucheng Qin, Xun Zhou, Ruihang Lai, Hongyi Jin, Yixin Dong, Bohan Hou, Meng-Shiun Yu, Yiyan Zhai, Sudeep Agarwal, Hangrui Cao, Siyuan Feng, and Tianqi Chen.

Scaling On-Device GPU Inference for Large Generative Models Ruan, Yucheng Qin, Xun Zhou, Ruihang Lai, Hongyi Jin, Yixin Dong, Bohan Hou, Meng-Shiun Yu, Yiyan Zhai, Sudeep Agarwal, Hangrui Cao, Siyuan Feng, and Tianqi Chen

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.737091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.496210Z digest=sha256:b56c83399ab046b11d44c2973193a96645f553b57190c9a97ea9a1628533b5d1

Observation f547c7bc-5890-481e-b1d7-8b8c3decf77e · outbound

This paper cites XLA: Compiling Machine Learning for Peak Performance, 2020.

Scaling On-Device GPU Inference for Large Generative Models XLA: Compiling Machine Learning for Peak Performance, 2020

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.725717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.499934Z digest=sha256:53d937f901c27fa80a7d1fcd92c472c4f1212a05fbd9cbb2a4fcb177e1bc0094

Observation 10aef569-a863-4b62-a5e4-f42a03f003c5 · outbound

This paper cites Introducing Stable Diffusion 3.5.

Scaling On-Device GPU Inference for Large Generative Models Introducing Stable Diffusion 3.5

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.713687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.503769Z digest=sha256:0a5906db88482ace12726698994f9a5ebf1d0cc222b0c492abbdf82c777c64db

Observation 7fa999f2-5162-44ad-a18e-57d302ee4093 · outbound

This paper cites Introducing torchchat: Accelerating Local LLM Inference on Laptop, Desktop and Mobile.

Scaling On-Device GPU Inference for Large Generative Models Introducing torchchat: Accelerating Local LLM Inference on Laptop, Desktop and Mobile

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.703132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.507604Z digest=sha256:1d2396f91ef35d25c1caa8af44e92e1f5cf24fa67d593de5cad8850e9bd6afb1

Observation 8524c447-e6d9-4edc-9d07-816c3c206f37 · outbound

This paper cites an unresolved cited work.

Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:51:31.691885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.511990Z digest=sha256:e5f93dbd70494db06d8204b5429b12d90d6c2bd28cf8cc082ffba71b15206180

Observation 6adc8945-0d67-42f6-8bbd-27b2fb14f743 · outbound

This paper cites Dawn, a WebGPU implemen- tation.

Scaling On-Device GPU Inference for Large Generative Models Dawn, a WebGPU implemen- tation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.680108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.515581Z digest=sha256:d872037495c95d1099278a724771c98ef11b4d95a4a79f5d3a196374e8196e13

Observation 050b6cd3-07d8-4802-aa26-7d3b9057637f · outbound

This paper cites an unresolved cited work.

Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:51:31.668575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.519818Z digest=sha256:d73a9720a3e9c84b9f208267d3a5442f4bdde93310cb056e5151115f28688081

Observation 3b55a910-0d4c-4677-8cf2-78534dd35c7e · outbound

This paper cites SmoothQuant: Accurate and Effi- cient Post-Training Quantization for Large Language Mod- els.

Scaling On-Device GPU Inference for Large Generative Models SmoothQuant: Accurate and Effi- cient Post-Training Quantization for Large Language Mod- els

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.655869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.523550Z digest=sha256:211265fdefde2fa0227e2120cf4c3b0a83d934fdad4548fb30ce5f51c687ef05

Observation 4ebfc85a-c516-44d0-8f83-2390397b2c97 · outbound

This paper cites an unresolved cited work.

Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:51:31.643405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.527151Z digest=sha256:ba9dd8e39ea4e9cadc67a48d230257abc9d690d55462ae002973465419b122a8

Observation 4c0350f0-bd82-43d1-a93e-54649577265e · outbound

This paper cites LLMCad: Fast and Scalable On-device Large Language Model Inference, 2023.

Scaling On-Device GPU Inference for Large Generative Models LLMCad: Fast and Scalable On-device Large Language Model Inference, 2023

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.632711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.530744Z digest=sha256:5774d7a67d00125b3a081bab2d3505b1b0da0aceeb4f442b70cdf5c6fc4c08a1

Observation d73f38fa-cffa-4fa2-b165-fbec16e9cc78 · outbound

This paper cites Fast On-device LLM Inference with NPUs.

Scaling On-Device GPU Inference for Large Generative Models Fast On-device LLM Inference with NPUs

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.621573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.534059Z digest=sha256:64721b180d9e459975973863d391861cb6d4cd6b3938871cd9209d3eb411fb37

Observation 339db099-dd19-46ef-be92-964ead0a53a3 · outbound

This paper cites PowerInfer-2: Fast Large Language Model Inference on a Smartphone.

Scaling On-Device GPU Inference for Large Generative Models PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T04:51:31.537714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:51:31.537714Z digest=sha256:95b1cfa2810d11b3307136897722dbc3415b82b5c72d75577826f57fd2b34d3c

Observation 984377af-19f7-4a48-9783-9de59c1de3fc · outbound

This paper cites Gonzalez, Clark Barrett, and Ying Sheng.

Scaling On-Device GPU Inference for Large Generative Models Gonzalez, Clark Barrett, and Ying Sheng

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.609190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:51:31.541760Z digest=sha256:7e8bffe1d5487ec1c5b8836cb66d80687afe114ccd1ae55419123b0e2940115c

Pith citing papers

Observation 035f8d65-6761-4107-896c-35502c5af71f · inbound

EdgeWisePersona: A Dataset for On-Device User Profiling from Natural Language Interactions cites this paper.

EdgeWisePersona: A Dataset for On-Device User Profiling from Natural Language Interactions Scaling On-Device GPU Inference for Large Generative Models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:57:30.740928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:57:30.372775Z digest=sha256:7ee1e848bd22f1507923fe72960233f561bd39b939118308173a307212dc652e