Pith. sign in

Paper Citation Record · LEDGER

LLMCad: Fast and Scalable On-device Large Language Model Inference

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2309.04255.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.04255 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:55:06.634716Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

13
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 31ee100c-7934-4068-9c2a-4b4055b25329 · inbound

Deploying Foundation Model Powered Agent Services: A Survey cites this paper.

Deploying Foundation Model Powered Agent Services: A Survey LLMCad: Fast and Scalable On-device Large Language Model Inference

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.214235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.214235Z digest=sha256:d75258746e8a9c3c76467141ad2776d74b78e3e868540fc99e018dc56fd59ff4

Observation 207e6959-3378-4bc1-a55d-42618d093e59 · inbound

Large Language Models on Small Resource-Constrained Systems: Performance Characterization, Analysis and Trade-offs cites this paper.

Large Language Models on Small Resource-Constrained Systems: Performance Characterization, Analysis and Trade-offs LLMCad: Fast and Scalable On-device Large Language Model Inference

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T11:33:29.609712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:33:29.609712Z digest=sha256:70205783780c45338dadc9e1ea823142cd34200bc2e0ef4c139a6bff6cef3c3d

Observation 86f04886-6772-4780-8891-1abc5f0de0fb · inbound

PICE: A Semantic-Driven Progressive Inference System for LLM Serving in Cloud-Edge Networks cites this paper.

PICE: A Semantic-Driven Progressive Inference System for LLM Serving in Cloud-Edge Networks LLMCad: Fast and Scalable On-device Large Language Model Inference

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T20:13:42.873701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:13:42.873701Z digest=sha256:95205720e878dfec0c55160a9da6e73dee8440f616d33566b35fb32cdb572ee6

Observation b84fa4dd-a1af-40a9-a267-37d9d60a920a · inbound

The MoE-Empowered Edge LLMs Deployment: Architecture, Challenges, and Opportunities cites this paper.

The MoE-Empowered Edge LLMs Deployment: Architecture, Challenges, and Opportunities LLMCad: Fast and Scalable On-device Large Language Model Inference

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T05:19:35.751961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:19:35.751961Z digest=sha256:e3f814328dd9706878982330b3d20abd8a25720ede6c9b5fa6e88a1f2c9bb458

Observation 1d47bec5-bbfc-414c-a56c-79b985314e41 · inbound

Communication-Efficient Hybrid Language Model via Uncertainty-Aware Opportunistic and Compressed Transmission cites this paper.

Communication-Efficient Hybrid Language Model via Uncertainty-Aware Opportunistic and Compressed Transmission LLMCad: Fast and Scalable On-device Large Language Model Inference

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:55:06.634716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:55:06.634716Z digest=sha256:71a896350f0c4d0c3f7f4314a126c0b11e6418f3457ef9e9f0126f7b76c8015a

Observation 991996f3-5e6e-4046-889d-6806864ff512 · inbound

Edge-First Language Model Inference: Models, Metrics, and Tradeoffs cites this paper.

Edge-First Language Model Inference: Models, Metrics, and Tradeoffs LLMCad: Fast and Scalable On-device Large Language Model Inference

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:08.669376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:03:08.669376Z digest=sha256:c09587b3fff4f2d8d16792bb1157af8ab3285f6848c3862331075597d7bf6deb

Observation 4012f9de-80bc-48e5-bbbf-2660dc0bc53f · inbound

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism cites this paper.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism LLMCad: Fast and Scalable On-device Large Language Model Inference

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:10.198594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:10.198594Z digest=sha256:758a9506b9ba2996fc6488e07cf4b10ddcd6d925a5b8a252e7dff430acd94376

Observation f0112c21-da6c-4356-a3b6-c4449348fa7e · inbound

SoK: The Privacy Paradox of Large Language Models: Advancements, Privacy Risks, and Mitigation cites this paper.

SoK: The Privacy Paradox of Large Language Models: Advancements, Privacy Risks, and Mitigation LLMCad: Fast and Scalable On-device Large Language Model Inference

Reference 129

Resolution
unresolved
no resolver link, observed 2026-08-07T00:46:10.431567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:46:10.431567Z digest=sha256:dc965f6d4828a0c3fe4d3e14a9a157d3af3cef4e3526c016463cf4835c99d9cb

Observation 84b9f65e-e87a-4695-8004-395b4d9fd9e0 · inbound

Toward Edge General Intelligence with Multiple-Large Language Model (Multi-LLM): Architecture, Trust, and Orchestration cites this paper.

Toward Edge General Intelligence with Multiple-Large Language Model (Multi-LLM): Architecture, Trust, and Orchestration LLMCad: Fast and Scalable On-device Large Language Model Inference

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:14.815374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:14.815374Z digest=sha256:65bcb6929601403c62b6b52754b18e843af58be4393de6671fe510d337e7c824

Observation 946a9f0f-bebf-4248-9d2f-047f2fac3976 · inbound

Dissecting the Impact of Mobile DVFS Governors on LLM Inference Performance and Energy Efficiency cites this paper.

Dissecting the Impact of Mobile DVFS Governors on LLM Inference Performance and Energy Efficiency LLMCad: Fast and Scalable On-device Large Language Model Inference

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T20:44:17.577855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:44:17.577855Z digest=sha256:d2bebe7877a4fc39a5cc453966171c455ba7f42593713613e0b0662c1f05ddc4

Observation c27dfe0c-4edb-423c-8fa7-d61f33313141 · inbound

Collaborative Inference and Learning between Edge SLMs and Cloud LLMs: A Survey of Algorithms, Execution, and Open Challenges cites this paper.

Collaborative Inference and Learning between Edge SLMs and Cloud LLMs: A Survey of Algorithms, Execution, and Open Challenges LLMCad: Fast and Scalable On-device Large Language Model Inference

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:47.908155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:47.908155Z digest=sha256:f97c021a8a43167092744a6c712e56696db5c0ea5c91443a08065652dde801b7

Observation 3d8da4af-b8d6-4510-972e-098368068528 · inbound

A Unified Model and Document Representation for On-Device Retrieval-Augmented Generation cites this paper.

A Unified Model and Document Representation for On-Device Retrieval-Augmented Generation LLMCad: Fast and Scalable On-device Large Language Model Inference

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T11:55:20.092471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-10T11:54:09.047134Z digest=sha256:0c3c6a21becca1ddcb9c4916d8e5835ae099f406479151b524a841f4ebfd2c46

Observation 26ed021b-5237-4a1b-825a-665c2fa70f3d · inbound

Litespark Inference For CPUs: Ultra-Fast SIMD Framework for Ternary (1.58-bit) Language Models cites this paper.

Litespark Inference For CPUs: Ultra-Fast SIMD Framework for Ternary (1.58-bit) Language Models LLMCad: Fast and Scalable On-device Large Language Model Inference

Reference 30

Resolution
malformed identifier
arxiv_id, observed 2026-05-11T20:11:08.765744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T10:12:36.972813Z digest=sha256:8f2c852709774a815ac73898433248d60a01e64b436396a8233963493620a889

Observation 99247455-449b-4bbd-a21a-6313359a4049 · inbound

Litespark Inference For CPUs: Ultra-Fast SIMD Framework for Ternary (1.58-bit) Language Models cites this paper.

Litespark Inference For CPUs: Ultra-Fast SIMD Framework for Ternary (1.58-bit) Language Models LLMCad: Fast and Scalable On-device Large Language Model Inference

Reference 30

Resolution
malformed identifier
arxiv_id, observed 2026-07-01T13:25:46.082791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T23:12:52.038253Z digest=sha256:8c69e9850bcadb86ad40784ecd3f28f2ef47851959400d4a63a35b6dddd50732

Observation 7803a98a-0f7d-4895-b61c-0ebdf468abeb · inbound

AsymSpec: Efficient Cloud-Edge Speculative Decoding over Asymmetric Networks cites this paper.

AsymSpec: Efficient Cloud-Edge Speculative Decoding over Asymmetric Networks LLMCad: Fast and Scalable On-device Large Language Model Inference

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T12:36:18.965771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:36:18.965771Z digest=sha256:0d59ed621db64cca0431e035b34de0bac1b0a4665751611e497e3fbd841e68dd