Pith. sign in

Paper Citation Record · LEDGER

Scaling On-Device GPU Inference for Large Generative Models

As of 17 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 1 inbound Pith citation observation for arXiv:2505.00232.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.00232 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:51:31.541760Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:57:30.372775Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T20:57:30.733875Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy46
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8a6ea8cf-7f83-451e-b584-6b4e92d8fb08 · outbound

This paper cites The Khronos Group Inc., 2019.

Scaling On-Device GPU Inference for Large Generative Models The Khronos Group Inc., 2019

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.226757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.326199Z digest=sha256:cb78cbac1bb0fa836eec41caa7d02d95e380b388a0dd525679d0a63f6e26b6e8

Observation 84050aef-ddba-42d5-b04b-ecb0c3593d90 · outbound

This paper cites The Khronos Group Inc., 2025.

Scaling On-Device GPU Inference for Large Generative Models The Khronos Group Inc., 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.216060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.330309Z digest=sha256:d86adfe8a672e6953d0d356f8f6fc4b807591adb642dcad38071eda5d5a7e520

Observation 64e918cc-7901-4b36-ac95-2d2349573ec9 · outbound

This paper cites AMD ROCm Soft- ware.

Scaling On-Device GPU Inference for Large Generative Models AMD ROCm Soft- ware

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.205450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.333808Z digest=sha256:d08d63eef8323890bb61e00095850ff0145327c266a7f1764fa90ef19093e1bd

Observation 35619c9e-5e6b-4bf3-80c4-cf0862e6fb69 · outbound

This paper cites LLM in a flash: Effi- cient Large Language Model Inference with Limited Mem- ory, 2024.

Scaling On-Device GPU Inference for Large Generative Models LLM in a flash: Effi- cient Large Language Model Inference with Limited Mem- ory, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.195166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.337188Z digest=sha256:684f0a2c586c33a3aac5c32a0e812ac7f5d9213df457cb4d73a9c3aaa566895e

Observation 50f83bf4-3208-4146-9e8f-0d8a1022d446 · outbound

This paper cites Core ML Stable Diffusion.

Scaling On-Device GPU Inference for Large Generative Models Core ML Stable Diffusion

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.184617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.340853Z digest=sha256:80d8025e87a0d6645d054cec5063604d8910f59fe2ee3d803fbe1fdca8d658b3

Observation 1600da46-e464-4788-8bf6-58925191c0e6 · outbound

This paper cites an unresolved cited work.

Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:51:32.174494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.343987Z digest=sha256:19ce092558fd4229fae6fedea8925f35c8fa0235895d2b9723bef0814f50ff0f

Observation 03b334a5-6335-4ae8-a223-c2118fe97de6 · outbound

This paper cites Compute Library.

Scaling On-Device GPU Inference for Large Generative Models Compute Library

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.163667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.347636Z digest=sha256:82f5eb3c87237fb0c2b7bef8e85d79fa794c1f884f0853aa28307603b19e7410

Observation 3a2aac5d-634a-4b7f-9068-6a346dbb38e6 · outbound

This paper cites TVM: An automated End-to-End optimizing com- piler for deep learning.

Scaling On-Device GPU Inference for Large Generative Models TVM: An automated End-to-End optimizing com- piler for deep learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.153440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.350756Z digest=sha256:0735fdbd90b8e83917d4e3cd2ac572cc9aa98b7577f6def42f752c637be6f6fb

Observation ea88d8f3-31d9-4065-b0a5-4af7c03d8324 · outbound

This paper cites Speed Is All You Need: On-Device Acceleration of Large Diffusion Models via GPU-Aware Optimizations, 2023.

Scaling On-Device GPU Inference for Large Generative Models Speed Is All You Need: On-Device Acceleration of Large Diffusion Models via GPU-Aware Optimizations, 2023

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.143000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.353895Z digest=sha256:94e560e6447424105d5ba90d43e121b361091dd29796aa75ee721cb66b099ee5

Observation 1bb45ed1-17a1-458f-8cd6-cb1221f0cc8f · outbound

This paper cites LMDeploy: A Toolkit for Com- pressing, Deploying, and Serving LLM.

Scaling On-Device GPU Inference for Large Generative Models LMDeploy: A Toolkit for Com- pressing, Deploying, and Serving LLM

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.132514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.356758Z digest=sha256:5eafbc0c42a0fb305b93f31454cd4fc065d8583d7814051c031f8d223e3fc1a8

Observation a02154a5-1f7c-46cc-85ef-d54b8eef9de5 · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher R´e.

Scaling On-Device GPU Inference for Large Generative Models Fu, Stefano Ermon, Atri Rudra, and Christopher R´e

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.122766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.359871Z digest=sha256:bc249fc6b20ab6028802e802fb80bcca5468051ac59fec596505284433d4241c

Observation d1287604-d0ae-4cfc-aa76-73f52a5097b6 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Scaling On-Device GPU Inference for Large Generative Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T04:51:31.363141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:51:31.363141Z digest=sha256:1c9687a7fde35da599c9f59410a56174a8ed796748f4cc636ca10b7c4e51454c

Observation 677ff3df-9209-451b-a302-346cbdb2a438 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Gen- erative Pre-trained Transformers, 2023.

Scaling On-Device GPU Inference for Large Generative Models GPTQ: Accurate Post-Training Quantization for Gen- erative Pre-trained Transformers, 2023

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.111893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.367086Z digest=sha256:483ed23b5e6fc014a5f10dc823b7c02b56884adff170b6b7df290cbda89fbf4b

Observation 02d036d3-af8f-4eee-b43d-c27a107cd2d7 · outbound

This paper cites Gemma 2: Improving Open Language Mod- els at a Practical Size, 2024.

Scaling On-Device GPU Inference for Large Generative Models Gemma 2: Improving Open Language Mod- els at a Practical Size, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.099149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.370650Z digest=sha256:9616dce055a03c0ce8242f4e24320496306b80c91d4e585cadc4243429f0057c

Observation c7015efd-80bd-4314-9f32-01401939d2b7 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology, 2024.

Scaling On-Device GPU Inference for Large Generative Models Gemma: Open Models Based on Gemini Research and Technology, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.085958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.374881Z digest=sha256:1530743c23d86ed034044fe97f8390d9f7e92efa3ae2633010c42eec46b225bf

Observation 0aea38be-4609-4085-89ad-90ae866ed3ee · outbound

This paper cites llama.cpp.

Scaling On-Device GPU Inference for Large Generative Models llama.cpp

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.075012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.379167Z digest=sha256:d217335503248264842a794909fcca097caab5fe285b1f0f82aeb01e6f0c70c1

Observation 6e8b1186-d337-49c2-8e95-ddfdba6a5788 · outbound

This paper cites an unresolved cited work.

Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:51:32.064032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.383046Z digest=sha256:1039bccc16c14a69ef70899834e3b2c4db8491100d7099229e02d9e7a98ede9f

Observation be793fe5-21cb-4c21-858d-104d7958d003 · outbound

This paper cites LiteRT Overview.

Scaling On-Device GPU Inference for Large Generative Models LiteRT Overview

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.052536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.387319Z digest=sha256:ab1f222b6fe97b2487ab79ecc4c190631b5be8552c1425c1c03621cace8ff0f7

Observation 4f96f6ab-f0fe-43a4-a47c-bdc7b9ba331f · outbound

This paper cites Huawei HiAI.

Scaling On-Device GPU Inference for Large Generative Models Huawei HiAI

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.038995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.392090Z digest=sha256:5de0b91ac0a4c69df79ae114898c661190ff47eba5427f5b3e670a6b3eaa13e8

Observation 4c372e43-57e6-4f94-a62a-e31759ed6572 · outbound

This paper cites an unresolved cited work.

Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:51:32.029100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.396993Z digest=sha256:1cb49bbf304419cbf9da46997288d7b2e4d091f7262c72a68e20e1afa3d52fcf

Observation a88a2194-250a-488e-9509-80693c5284fe · outbound

This paper cites OpenVINO.

Scaling On-Device GPU Inference for Large Generative Models OpenVINO

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.018800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.401664Z digest=sha256:33cea029157361e8e51f2bcee89be12f251713f277445eac16c80a419ea7478a

Observation 9ac07930-d5b7-4318-8149-b9713f0d4c83 · outbound

This paper cites Intel Core Ultra Series 2 Media Deck.

Scaling On-Device GPU Inference for Large Generative Models Intel Core Ultra Series 2 Media Deck

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.007819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.405781Z digest=sha256:f08990f53dd31aed5039ba3c38bf618d3692290ff4e83be399dfa3b326bbf61e

Observation 64a14253-3f9a-4fb7-ba60-b5f0216bb40d · outbound

This paper cites an unresolved cited work.

Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:51:31.996993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.411970Z digest=sha256:3166c420f565b4ae399c9ad1e1d82e3765f2b885a080160672a37a3e5d450e4f

Observation 5171003c-2e67-4fdb-8567-d34c505fe892 · outbound

This paper cites MNN: A Universal and Efficient Inference Engine.

Scaling On-Device GPU Inference for Large Generative Models MNN: A Universal and Efficient Inference Engine

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T04:51:31.415946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:51:31.415946Z digest=sha256:e5aaebf974d7fd9a767de6bedfd70208b914e77af4d9994aff4f854c5a199c4f

Observation f28ca99f-b57a-461b-8dde-39ac0ab86c1f · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Scaling On-Device GPU Inference for Large Generative Models Gonzalez, Hao Zhang, and Ion Stoica

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.985623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.420135Z digest=sha256:6907c4c39019d71a1e80e4cb2881731667db79ffb14dce3086c1b2095cb5b9ab

Observation 01410587-29d5-48ee-a2fe-8319d24df45f · outbound

This paper cites On-Device Neu- ral Net Inference with Mobile GPUs.

Scaling On-Device GPU Inference for Large Generative Models On-Device Neu- ral Net Inference with Mobile GPUs

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.974737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.424276Z digest=sha256:3bd20a4504bd2f55c224570785c4125e166f950248e2767d460cd0221e2508b7

Observation 6ac4ea48-9622-4f7b-92d0-760a9b383d46 · outbound

This paper cites OpenGL ES Version 3.1.

Scaling On-Device GPU Inference for Large Generative Models OpenGL ES Version 3.1

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.963044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.427827Z digest=sha256:1acde51d708f03c0f49a67f7853cf59e8b2d32dabe83e83faeb5b3fe5775e467

Observation b7b87274-5c91-466f-82f3-535ef56f05ce · outbound

This paper cites Fast In- ference from Transformers via Speculative Decoding, 2023.

Scaling On-Device GPU Inference for Large Generative Models Fast In- ference from Transformers via Speculative Decoding, 2023

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.951919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.431461Z digest=sha256:c9f17d8b85dc5c58176de77121c8ef3024e1a9899ffc6ff5001777a01917f599

Observation 48103cb2-cd68-40b7-8954-9ca3ef371738 · outbound

This paper cites AWQ: Activation-aware Weight Quantization for LLM Compression and Accelera- tion.

Scaling On-Device GPU Inference for Large Generative Models AWQ: Activation-aware Weight Quantization for LLM Compression and Accelera- tion

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.942289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.435413Z digest=sha256:af4ed6acb6321fd212985f4af65966afbf38d29bdf63a1ffcbb478eeedbf8cbd

Observation d5a778d6-56d8-4ef9-88ce-5d00b8bba3a9 · outbound

This paper cites The Llama 3 Herd of Models, 2024.

Scaling On-Device GPU Inference for Large Generative Models The Llama 3 Herd of Models, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.932892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.438744Z digest=sha256:9be7296621296fe0acaee49e7fbe70a1ec4808a32eb009bf4249383f79fd65ce

Observation 39867c08-cad7-418b-bd15-3f6fba21ecca · outbound

This paper cites DeepMon: Mobile GPU-based Deep Learning Framework for Continuous Vision Applications.

Scaling On-Device GPU Inference for Large Generative Models DeepMon: Mobile GPU-based Deep Learning Framework for Continuous Vision Applications

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.921260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.442284Z digest=sha256:4e457473aaf0fc11b150978c69b3ece83000b4f75e16ac4ad84f3bd6d8126895

Observation bee3c6e9-0c36-4cf2-8a12-31c50c1489ac · outbound

This paper cites NeuroPilot.

Scaling On-Device GPU Inference for Large Generative Models NeuroPilot

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.909923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.445790Z digest=sha256:be6cd2d56ab10ad0dd6277de1f812b71aa36505e99439bf968d2d316c68e9fd8

Observation c3ef5adf-bf8c-4145-8251-dab241f281a6 · outbound

This paper cites ExecuTorch.

Scaling On-Device GPU Inference for Large Generative Models ExecuTorch

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.898276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.449046Z digest=sha256:0fe5c41e9925a18c794f9d0ef9b52cc2b4eb80c77ee5f7455cc59b8310707921

Observation ac7b1861-7567-46ab-9388-0652468d864f · outbound

This paper cites Llama 3.2: Revolutionizing edge AI and vision with open, customizable models.

Scaling On-Device GPU Inference for Large Generative Models Llama 3.2: Revolutionizing edge AI and vision with open, customizable models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.887304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.452475Z digest=sha256:08b329a058367ca3a40ba08a8d1e8dee9c60b646b00fef634d69960fdc54232a

Observation c0838d42-45f4-4b9e-a3a1-48effb96e14d · outbound

This paper cites DirectML Overview.

Scaling On-Device GPU Inference for Large Generative Models DirectML Overview

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.876047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.456091Z digest=sha256:1efb03371e19283db99f484d80a4e99db4734dfdff3fce6c8f2de5e732e7b7ef

Observation af0ad225-015b-4b94-b423-7c10975a72c6 · outbound

This paper cites Get started with ONNX Run- time Mobile.

Scaling On-Device GPU Inference for Large Generative Models Get started with ONNX Run- time Mobile

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.864756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.459461Z digest=sha256:fd0db7f04619a4526ba224fdfe311e4f5b6a69e613d8f6785c11bff2906be872

Observation d0268f61-6ea7-4c56-9c13-91e0b67c767b · outbound

This paper cites Stable Diffusion Op- timization with DirectML.

Scaling On-Device GPU Inference for Large Generative Models Stable Diffusion Op- timization with DirectML

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.854183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.463215Z digest=sha256:f86368f1ac7ea87e3861ff1d6482a4b3b17c461fe82796f60d6fcde17bdcdbc7

Observation bfa471c6-016f-43ce-96da-0bb3acfc7576 · outbound

This paper cites an unresolved cited work.

Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:51:31.842022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.466325Z digest=sha256:603aa8350211058beff8bf47d5f8e6951d0bfbf5c9e597b0c828fab41ab58fa6

Observation 67262328-6b4d-4379-a936-70c921ad655a · outbound

This paper cites an unresolved cited work.

Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:51:31.828786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.469475Z digest=sha256:738fb7c6eb6beaa3a19e76ab4bf066a9be269ccb90bae5d432fca3dd1a2d068c

Observation 96c5db1d-d03b-4196-b9ad-5167ee423445 · outbound

This paper cites NVIDIA Tensor Cores.

Scaling On-Device GPU Inference for Large Generative Models NVIDIA Tensor Cores

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.816678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.473074Z digest=sha256:2109958ef6c22e6d3096111086d87bd174d00fd10ecc27f382e316f972a38fd7

Observation f4b38987-3aeb-4e27-a84d-b1d2b3938805 · outbound

This paper cites TensorRT.

Scaling On-Device GPU Inference for Large Generative Models TensorRT

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.803472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.475989Z digest=sha256:a275d87cbe70dec3f09718c27ff77a8b89032a5dca9c4699b346223882ffe94c

Observation 17bc6ceb-1dde-4cc4-9ee9-81596a10cf25 · outbound

This paper cites an unresolved cited work.

Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:51:31.792711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.479341Z digest=sha256:15fba2852f5c7977ab49567cb23bfa5eac64f91fe9dd1f459513df7b3bd3bfe8

Observation 2572a02f-e45d-4571-9175-ec554570753e · outbound

This paper cites Efficient Memory Manage- ment for Deep Neural Net Inference.

Scaling On-Device GPU Inference for Large Generative Models Efficient Memory Manage- ment for Deep Neural Net Inference

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.782707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.482345Z digest=sha256:24a5ae1f6585ff1a887ef2c68f9260c088b29dd59be4f8dda1097f1086c0549d

Observation c1f17b16-a16d-4582-99ac-55d254b36740 · outbound

This paper cites Snapdragon Neural Processing Engine SDK.

Scaling On-Device GPU Inference for Large Generative Models Snapdragon Neural Processing Engine SDK

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.771094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.485334Z digest=sha256:4809ad98cea722426f465227a653ba576121959bda22bd9fd438b01f328d12ae

Observation cfda3ff7-68cc-4da5-914f-db7745bd842e · outbound

This paper cites QualComm AI Hub Llama-v3.2-3B-Chat.

Scaling On-Device GPU Inference for Large Generative Models QualComm AI Hub Llama-v3.2-3B-Chat

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.760222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.488945Z digest=sha256:bac0b8b6f1ee6469be5b1a6c9946d544a4df7e585a4d67161ee4efe42978cb63

Observation 5e70183f-a43b-4b95-81a6-02a97d5caadb · outbound

This paper cites World’s first on-device demonstration of Stable Diffusion on an Android phone.

Scaling On-Device GPU Inference for Large Generative Models World’s first on-device demonstration of Stable Diffusion on an Android phone

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.749703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.492697Z digest=sha256:3eed1dea589dde937c6faf42a3442fe1a1d13326420520561b2c6519867d0387

Observation 9ed90683-f398-487f-a7d3-b88063eda807 · outbound

This paper cites Ruan, Yucheng Qin, Xun Zhou, Ruihang Lai, Hongyi Jin, Yixin Dong, Bohan Hou, Meng-Shiun Yu, Yiyan Zhai, Sudeep Agarwal, Hangrui Cao, Siyuan Feng, and Tianqi Chen.

Scaling On-Device GPU Inference for Large Generative Models Ruan, Yucheng Qin, Xun Zhou, Ruihang Lai, Hongyi Jin, Yixin Dong, Bohan Hou, Meng-Shiun Yu, Yiyan Zhai, Sudeep Agarwal, Hangrui Cao, Siyuan Feng, and Tianqi Chen

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.737091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.496210Z digest=sha256:38e1f746fa637322f5fd780eff39557eaf8f4bed84955bd6e09431411f07b9a1

Observation f547c7bc-5890-481e-b1d7-8b8c3decf77e · outbound

This paper cites XLA: Compiling Machine Learning for Peak Performance, 2020.

Scaling On-Device GPU Inference for Large Generative Models XLA: Compiling Machine Learning for Peak Performance, 2020

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.725717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.499934Z digest=sha256:028e744c68db50520d728b75c39052b83dc7429d10b991e4165b66a79a35c1f2

Observation 10aef569-a863-4b62-a5e4-f42a03f003c5 · outbound

This paper cites Introducing Stable Diffusion 3.5.

Scaling On-Device GPU Inference for Large Generative Models Introducing Stable Diffusion 3.5

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.713687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.503769Z digest=sha256:578e1b0d9ceee4b2328aef640a13d414f07485085289544f6d52c8e298c38c86

Observation 7fa999f2-5162-44ad-a18e-57d302ee4093 · outbound

This paper cites Introducing torchchat: Accelerating Local LLM Inference on Laptop, Desktop and Mobile.

Scaling On-Device GPU Inference for Large Generative Models Introducing torchchat: Accelerating Local LLM Inference on Laptop, Desktop and Mobile

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.703132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.507604Z digest=sha256:af8279008323381db0b7f4cc6e9446ac97acc0ab9069b408f831df8e9b257fe3

Observation 8524c447-e6d9-4edc-9d07-816c3c206f37 · outbound

This paper cites an unresolved cited work.

Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:51:31.691885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.511990Z digest=sha256:4d53a478f429306accb35f6dc13072442e549c07e795380f5b353328efdd0c89

Observation 6adc8945-0d67-42f6-8bbd-27b2fb14f743 · outbound

This paper cites Dawn, a WebGPU implemen- tation.

Scaling On-Device GPU Inference for Large Generative Models Dawn, a WebGPU implemen- tation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.680108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.515581Z digest=sha256:5942962112db37b04e1a52a8d08fc0ad0651903c3538816fdd1afff11558eaff

Observation 050b6cd3-07d8-4802-aa26-7d3b9057637f · outbound

This paper cites an unresolved cited work.

Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:51:31.668575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.519818Z digest=sha256:e99025b8177dad37a855260ab2b94ea262fa63f45504c8f7b15ebf306ad5a051

Observation 3b55a910-0d4c-4677-8cf2-78534dd35c7e · outbound

This paper cites SmoothQuant: Accurate and Effi- cient Post-Training Quantization for Large Language Mod- els.

Scaling On-Device GPU Inference for Large Generative Models SmoothQuant: Accurate and Effi- cient Post-Training Quantization for Large Language Mod- els

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.655869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.523550Z digest=sha256:e0b6593ffce4fe24fbdc66af87787e20d45ff1ed5d6d8ad6064f3e5f1c9a2586

Observation 4ebfc85a-c516-44d0-8f83-2390397b2c97 · outbound

This paper cites an unresolved cited work.

Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:51:31.643405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.527151Z digest=sha256:d6ec05085b564c97abcabe0f882223ddd3f9eba3abcccf7845493bbcabc31c77

Observation 4c0350f0-bd82-43d1-a93e-54649577265e · outbound

This paper cites LLMCad: Fast and Scalable On-device Large Language Model Inference, 2023.

Scaling On-Device GPU Inference for Large Generative Models LLMCad: Fast and Scalable On-device Large Language Model Inference, 2023

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.632711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.530744Z digest=sha256:45d70d6f0185b422b7b320d5beacbe268c760bb0e642e19a0d31fef34a10fdba

Observation d73f38fa-cffa-4fa2-b165-fbec16e9cc78 · outbound

This paper cites Fast On-device LLM Inference with NPUs.

Scaling On-Device GPU Inference for Large Generative Models Fast On-device LLM Inference with NPUs

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.621573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.534059Z digest=sha256:0a4f9688ca9827735a7df4691b3e97f4ab84151b516393e809805e9032cd9e1b

Observation 339db099-dd19-46ef-be92-964ead0a53a3 · outbound

This paper cites PowerInfer-2: Fast Large Language Model Inference on a Smartphone.

Scaling On-Device GPU Inference for Large Generative Models PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T04:51:31.537714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:51:31.537714Z digest=sha256:95b1cfa2810d11b3307136897722dbc3415b82b5c72d75577826f57fd2b34d3c

Observation 984377af-19f7-4a48-9783-9de59c1de3fc · outbound

This paper cites Gonzalez, Clark Barrett, and Ying Sheng.

Scaling On-Device GPU Inference for Large Generative Models Gonzalez, Clark Barrett, and Ying Sheng

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.609190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:51:31.541760Z digest=sha256:3b624bbb565e719e487711fd798fa9fa66b22a7e3f1e2a800f6bde3c1e031413

Pith citing papers

Observation 035f8d65-6761-4107-896c-35502c5af71f · inbound

EdgeWisePersona: A Dataset for On-Device User Profiling from Natural Language Interactions cites this paper.

EdgeWisePersona: A Dataset for On-Device User Profiling from Natural Language Interactions Scaling On-Device GPU Inference for Large Generative Models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:57:30.740928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:57:30.372775Z digest=sha256:116f123c5a9b928891af718da7f5d1ef6af5f57ed67ab88ca09eaffcd1312038