Pith. sign in

Paper Citation Record · LEDGER

Lever: Speculative LLM Inference on Smartphones

As of 7 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2605.16786.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.16786 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-19T21:44:11.735963Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact12
  • verified fuzzy28
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4a55a98f-0f50-4dfc-8666-f184dcc78f19 · outbound

This paper cites Llm in a flash: Efficient large language model inference with limited memory.

Lever: Speculative LLM Inference on Smartphones Llm in a flash: Efficient large language model inference with limited memory

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:53:14.777804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:9eef8959d99dc245e7422d8fba0385fda9af7f61020b395a21541ddc8fef67ad

Observation 7ab4b8d0-a37f-4544-9947-f5bc2cfc998b · outbound

This paper cites Hydra: Sequentially-dependent draft heads for medusa decoding.

Lever: Speculative LLM Inference on Smartphones Hydra: Sequentially-dependent draft heads for medusa decoding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:53:14.759875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:b3f84736603a6fff09a5b83189350e6e38c39772997286c50d8cf8563ab433b4

Observation c806b47d-9bf4-41f2-8609-b9e8e2630af3 · outbound

This paper cites Program Synthesis with Large Language Models.

Lever: Speculative LLM Inference on Smartphones Program Synthesis with Large Language Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-19T21:47:48.480677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:5e1a0d3e73e538324fd4630b0ec1718c90e67c5bd7ca187174ef98a060175f9b

Observation f20d8a99-77b4-4a30-b5f4-534c8b326dcc · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

Lever: Speculative LLM Inference on Smartphones Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-19T21:47:48.477935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:48d5dc3e703b409a5a49db4880a5eb37d3e51a2fb3c4760618f468f8c45a2bb4

Observation 47ac8862-a568-4d7e-ac26-d6cc41891747 · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

Lever: Speculative LLM Inference on Smartphones Accelerating Large Language Model Decoding with Speculative Sampling

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-19T21:47:48.468798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:9ef368b80e9dcc45a429b83c6bc5d56ccae027887f5fb55f6d17869b8a7c2ced

Observation f5320cb9-4837-416c-8ebe-3085916cd432 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Lever: Speculative LLM Inference on Smartphones Evaluating Large Language Models Trained on Code

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-19T21:47:48.474946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:6df62df3d20c5df14be5b09ccd272b1c11aa1ab5e93b72daec79b04017e569c1

Observation 2b9bf438-a235-4596-a35f-05d1cbc2a091 · outbound

This paper cites Sequoia: Scalable and robust speculative decoding.

Lever: Speculative LLM Inference on Smartphones Sequoia: Scalable and robust speculative decoding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:53:14.768038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:b141aae5ed952d7c52aed2eb1b1cb8148d1ba3fa0d948b19ffafa0f4923cb80b

Observation 5583c18f-e631-4dc9-a91a-306e7e3061d1 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Lever: Speculative LLM Inference on Smartphones Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-19T21:47:48.492084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:a5dbba534ed71d3ffc98f2ee34a945d70f654dcb7755e3b2fa18990b4e03920c

Observation fd0e143f-51f9-4f33-8f11-8b2fc3925402 · outbound

This paper cites LayerSkip: Enabling early exit inference and self-speculative decoding.

Lever: Speculative LLM Inference on Smartphones LayerSkip: Enabling early exit inference and self-speculative decoding

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:53:14.753660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:64956d3d990598edd5065347c65ed6bdd06b42aee1b161155f1e8433ea3c3dfb

Observation 929d8eb6-4ac7-428f-acae-2fcdd069c586 · outbound

This paper cites GPTQ: Accurate post-training quantization for generative pre-trained transformers.

Lever: Speculative LLM Inference on Smartphones GPTQ: Accurate post-training quantization for generative pre-trained transformers

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:53:14.751773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:7177af87547ca4b8065a8b89b1d433c24155c5b09f71855ca60a3bbe9af3694a

Observation 8480fb3a-0047-4799-89d4-b0aac697cf86 · outbound

This paper cites Break the se- quential dependency of LLM inference using lookahead decoding.

Lever: Speculative LLM Inference on Smartphones Break the se- quential dependency of LLM inference using lookahead decoding

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:53:14.783523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:6639c75f5acd5e060abc1ba73bbc34680aae6bff78d81b9dcfb63dffba5b4e54

Observation b142e5f8-67fb-4ed2-955b-b1203d0c774c · outbound

This paper cites CE-CoLLM: Efficient and Adaptive Large Language Models Through Cloud-Edge Collaboration.

Lever: Speculative LLM Inference on Smartphones CE-CoLLM: Efficient and Adaptive Large Language Models Through Cloud-Edge Collaboration

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-19T21:47:48.472228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:982abe23869f569abe2ff972b756adf96f97775e5a0e92ca7f3397326aefe028

Observation 0e4091ae-1c7c-4a78-a8d0-ed50d983fe06 · outbound

This paper cites Fast inference from transformers via speculative decoding.

Lever: Speculative LLM Inference on Smartphones Fast inference from transformers via speculative decoding

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:53:14.763054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:9f0cc3d461ed4056053dc0fa420bd4dc452744257960edb2af2a40936af97d87

Observation 0109dac0-2377-4400-aa27-26221df8c522 · outbound

This paper cites EAGLE-2: Faster inference of language models with dynamic draft trees.

Lever: Speculative LLM Inference on Smartphones EAGLE-2: Faster inference of language models with dynamic draft trees

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:53:14.781676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:729511613830b6cd43001c0d96f263b35ce55f9d241f09cde2d92ce68c2f3ad0

Observation 43fa5db7-7f3c-4c18-828b-8390a1667344 · outbound

This paper cites EA- GLE: Speculative sampling requires rethinking feature uncertainty.

Lever: Speculative LLM Inference on Smartphones EA- GLE: Speculative sampling requires rethinking feature uncertainty

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:53:14.787736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:f9c4d0323856a816b45c146734631e7976a4b5a6bea20ce2bab85b71f67c63e6

Observation 471a3735-ebdb-49f2-8f51-f3cbaebe3b6e · outbound

This paper cites EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test.

Lever: Speculative LLM Inference on Smartphones EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-19T21:47:48.489421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:bc299a5e2127603fc9dace472a8f971b56cb365b336b3b2be603227550eb3607

Observation d8c7925e-1f15-489b-a754-914c6a258623 · outbound

This paper cites Awq: Activation-aware weight quantization for on-device llm compression and acceleration.Proceedings of machine learning and systems, 6:87–100.

Lever: Speculative LLM Inference on Smartphones Awq: Activation-aware weight quantization for on-device llm compression and acceleration.Proceedings of machine learning and systems, 6:87–100

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:53:14.779676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:91b691ff9d2929689f72ba5b68c7968cf6aedfa6a1b5163257815811918f7c1b

Observation da23b84a-0465-454a-b284-4f10f56fc600 · outbound

This paper cites FastBERT: a self-distilling BERT with adaptive inference time.

Lever: Speculative LLM Inference on Smartphones FastBERT: a self-distilling BERT with adaptive inference time

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:53:14.747754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:53dd27fb1bbb874ab42867977ea2206cf3e3fcd99fd81178bea0d3c8f5b562d5

Observation 9cbfb044-abdd-4fa9-b2af-c53ab62e940d · outbound

This paper cites MobileLLM: Optimizing sub-billion parameter language models for on-device use cases.

Lever: Speculative LLM Inference on Smartphones MobileLLM: Optimizing sub-billion parameter language models for on-device use cases

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:53:14.750145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:e5d2b2dd678ccec7b9048cd1748361e46d62734ca224218dfba3f86d88c595c0

Observation e134dff8-012a-410d-9661-19996f42d2be · outbound

This paper cites Deja vu: Contextual sparsity for efficient llms at inference time.

Lever: Speculative LLM Inference on Smartphones Deja vu: Contextual sparsity for efficient llms at inference time

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:53:14.736382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:62d0ef179510e58566b98715ec75e2dc0507647859c028e0a41b4087d28e0d66

Observation d3bf30dc-e7e5-4930-a343-054dc590740b · outbound

This paper cites Llm-pruner: On the structural pruning of large language models.Advances in neural information processing systems, 36:21702–21720.

Lever: Speculative LLM Inference on Smartphones Llm-pruner: On the structural pruning of large language models.Advances in neural information processing systems, 36:21702–21720

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:53:14.765787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:726d434f100a664a460112cd244e6e82e44ca8d02811b039613c8e55aaa28dd7

Observation b742bc4f-9bd9-4f55-82bf-18906314c159 · outbound

This paper cites Specinfer: Accelerating large language model serving with tree-based speculative inference and verification.

Lever: Speculative LLM Inference on Smartphones Specinfer: Accelerating large language model serving with tree-based speculative inference and verification

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:53:14.758095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:579c279f3cc704a79d97bb754340ecfbc1d5d493a77a229f967a81d3a68147c8

Observation e71a7583-867b-4fab-a4c8-66420eab21f3 · outbound

This paper cites Fu, Zhiqiang Xie, Beidi Chen, Clark Barrett, Joseph E.

Lever: Speculative LLM Inference on Smartphones Fu, Zhiqiang Xie, Beidi Chen, Clark Barrett, Joseph E

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:53:14.772721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:aacd4c121fd2348a71cd885a06df4b3546c793b9f0008a4d6843476181c4f811

Observation 349c404d-c796-40dd-91a8-5f9ea9d9a8a2 · outbound

This paper cites Blockwise parallel decoding for deep autoregressive models.

Lever: Speculative LLM Inference on Smartphones Blockwise parallel decoding for deep autoregressive models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:53:14.775372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:9f43dcfd300e903d625317195fcb720ffcd645bebf99f897e51ab1e1ded9d9ef

Observation e5b61e7b-8026-4c81-9064-08b3d4a162c9 · outbound

This paper cites BitNet: Scaling 1-bit Transformers for Large Language Models.

Lever: Speculative LLM Inference on Smartphones BitNet: Scaling 1-bit Transformers for Large Language Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-19T21:47:48.486580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:7ca772b559da8bda2b9f16c0c1feb441085a2039e71015941eded93ed6f58364

Observation a6c2edc2-a107-41ca-a1b9-f6615c9ff43d · outbound

This paper cites OPT-tree: Speculative decoding with adaptive draft tree structure.Transactions of the Association for Computational Linguistics, 13:188–199.

Lever: Speculative LLM Inference on Smartphones OPT-tree: Speculative decoding with adaptive draft tree structure.Transactions of the Association for Computational Linguistics, 13:188–199

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:53:14.734189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:4cc8f71d061f3761e5ddb4f4e9bbb1588b38f49d161b254e24825655e224ee53

Observation b9334896-7057-4559-98d9-fa463ccd9d48 · outbound

This paper cites JENGA: Enhancing LLM Long-Context fine-tuning with con- textual token sparsity.

Lever: Speculative LLM Inference on Smartphones JENGA: Enhancing LLM Long-Context fine-tuning with con- textual token sparsity

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:53:14.770773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:0319cc3bed2da1cf39aae029688f4d714b3af6da2552fa5bc4ac905e4ae9eb38

Observation fd04f795-bb45-447c-b4e5-cff42508daa7 · outbound

This paper cites SWARM: Co- activation aware KVCache offloading across multiple SSDs.

Lever: Speculative LLM Inference on Smartphones SWARM: Co- activation aware KVCache offloading across multiple SSDs

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-19T21:47:48.498196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:0e4b7602dd201eb6f803b70e9d29466ad9657b8d5a70eb9729995f2c917ecffc

Observation 2d7bdc98-a329-4bdd-ba93-813a0514f57d · outbound

This paper cites Neuralink: Fast on-device llm inference with neuron co-activation linking.

Lever: Speculative LLM Inference on Smartphones Neuralink: Fast on-device llm inference with neuron co-activation linking

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:53:14.768293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:c2cb346e75a3c8050f57cdefd4f6590945bb0972d809396bc210a1a956dcc6b2

Observation 4c6efb37-e32a-4645-ac99-d614070de065 · outbound

This paper cites DynaKV: Enabling accurate and efficient long-sequence LLM decoding on smartphones.

Lever: Speculative LLM Inference on Smartphones DynaKV: Enabling accurate and efficient long-sequence LLM decoding on smartphones

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-19T21:47:48.501015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:4cb105d4905cf21e619b778d22482137a4a1ff58aafbadc058e70576de1494e7

Observation 1bb69f58-0f30-4c85-9323-3af3c30743fb · outbound

This paper cites Long Exposure: Accelerating parameter- efficient fine-tuning for LLMs under shadowy sparsity.

Lever: Speculative LLM Inference on Smartphones Long Exposure: Accelerating parameter- efficient fine-tuning for LLMs under shadowy sparsity

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:53:14.738404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:713a613d1b7c4f801b68f0f65e946b1b673950b23bf541e9ca306e2ece43bfa9

Observation 50b94e6d-310f-4e86-ae6a-7479332d9588 · outbound

This paper cites Mosaic: Cross-Modal Clustering for Efficient Video Understanding.

Lever: Speculative LLM Inference on Smartphones Mosaic: Cross-Modal Clustering for Efficient Video Understanding

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-19T21:47:48.495034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:19bbe18b6b184b77aa0d89142c09329a40b5b825ea099e6fd31266f96d37bb21

Observation 7aad47aa-567c-4c32-ad94-b0939b439caa · outbound

This paper cites SmoothQuant: Accurate and efficient post-training quantization for large language models.

Lever: Speculative LLM Inference on Smartphones SmoothQuant: Accurate and efficient post-training quantization for large language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:53:14.785339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:7c70b292c75dc2e91815aee6a8843150126397a41673ee170733031ed8fa1a69

Observation 119da033-5b1c-4ad6-b959-cd18aa96eebf · outbound

This paper cites Dee- BERT: Dynamic early exiting for accelerating BERT inference.

Lever: Speculative LLM Inference on Smartphones Dee- BERT: Dynamic early exiting for accelerating BERT inference

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:53:14.759602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:42780f941ed941ff2937a28f3e5f35bccd0e949379b10bf5c346764b4c24f5e4

Observation d07a455a-f962-4c5a-8548-1a192cfde327 · outbound

This paper cites Edgellm: Fast on-device llm inference with speculative decoding.IEEE Transactions on Mobile Computing, 24(4):3256–3273.

Lever: Speculative LLM Inference on Smartphones Edgellm: Fast on-device llm inference with speculative decoding.IEEE Transactions on Mobile Computing, 24(4):3256–3273

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:53:14.727483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:246170c38bf213beff18bd3fb7b68ff565a8332ad6fd1af359710c272d48d224

Observation 1d99e954-1e6b-495a-b3a5-5857d15d1b5d · outbound

This paper cites PowerInfer-2: Fast Large Language Model Inference on a Smartphone.

Lever: Speculative LLM Inference on Smartphones PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-19T21:47:48.483707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:06bca5026cb2abf56ee8ff951e74426979b3971d0bc4fa25530ba616587b5f79

Observation d232f587-7c61-4aaa-9f75-57d0fc04566f · outbound

This paper cites A first look at efficient and secure on-device LLM inference against KV leakage.

Lever: Speculative LLM Inference on Smartphones A first look at efficient and secure on-device LLM inference against KV leakage

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:53:14.730334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:dde76af52422d3320706532b619b8b644cacd1d8f97c66419d0c29e79cc20738

Observation 5a8811b9-0cac-434a-90a4-67032ffba475 · outbound

This paper cites Prism: Privacy-aware routing for adaptive cloud–edge llm inference via se- mantic sketch collaboration.

Lever: Speculative LLM Inference on Smartphones Prism: Privacy-aware routing for adaptive cloud–edge llm inference via se- mantic sketch collaboration

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:53:14.732253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:5f220df14baf76aa4bca46b5e77f1e26ec25173fec344d3c5caba1cdaea3cf60

Observation 619a6814-d07b-4818-bdf3-617c75127f91 · outbound

This paper cites Edgeshard: Efficient llm inference via collaborative edge com- puting.IEEE Internet of Things Journal, 12(10):13119–13131.

Lever: Speculative LLM Inference on Smartphones Edgeshard: Efficient llm inference via collaborative edge com- puting.IEEE Internet of Things Journal, 12(10):13119–13131

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:53:14.757555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:bae9df56145596e84036e934a8984f5c86d0a8fde2242b725c178d49c1d758c4

Observation 2e58a536-11af-4852-a4ef-ee19f4b32f63 · outbound

This paper cites Xing, Hao Zhang, Joseph E.

Lever: Speculative LLM Inference on Smartphones Xing, Hao Zhang, Joseph E

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T21:53:14.751969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:44:11.735963Z digest=sha256:f28673939e688277f0c4f4c40c1f0fe8609bfae04af42564154d0d40fab0b900

Pith citing papers

No inbound Pith citation observations are available.