Pith. sign in

Paper Citation Record · LEDGER

Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 40 inbound Pith citation observations for arXiv:2401.07851.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.07851 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 40 of 40 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:49:10.203035Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T18:50:04.384621Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e7c8b869-cec5-4ecb-ad09-032e8276c2cf · inbound

DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models cites this paper.

DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-24T03:23:49.645933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-24T03:23:18.827351Z digest=sha256:70ed4f905e60cd3bc38c4cf90ad4e65ea703fdfae3b5f690bbe7b4df565b5446

Observation 3d69fc85-6b0c-442c-b0d4-c8660aa56cb9 · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 269

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:39:33.456272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:e1a49879bd390a7e19baa7da4fb2aa38fc9d46275ec6eb1db0cfee77b8a86864

Observation 33c00f9f-f156-45ba-9bf7-de68cdd0c05c · inbound

SAM Decoding: Speculative Decoding via Suffix Automaton cites this paper.

SAM Decoding: Speculative Decoding via Suffix Automaton Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T19:32:52.275286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:32:52.275286Z digest=sha256:51a39d054b22e62accb3397593038eb37837c24035db0e249706dc49c47812e7

Observation 917008b4-235a-437f-8a8f-eeb94b1ac049 · inbound

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding cites this paper.

Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:22.096162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:49:22.096162Z digest=sha256:cf666b4d13835ec55d21a2d0f6d961dea470c34b759b7aa8aeb881a2665495ef

Observation d116a5fe-8de5-42a1-9943-686b508e1d1d · inbound

Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration cites this paper.

Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T13:15:07.274460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:15:07.274460Z digest=sha256:5a4351c45887f53f32bf083d808ba74a568699ee5bf94f72175921c621344a21

Observation 1bfa167f-cdf0-4eb9-9c36-a736132907b6 · inbound

PLD+: Accelerating LLM inference by leveraging Language Model Artifacts cites this paper.

PLD+: Accelerating LLM inference by leveraging Language Model Artifacts Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T04:26:59.618955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:26:59.618955Z digest=sha256:08999ca33a152949626f871dde504fd2749f8363cf90d9a1179d972862374721

Observation 6a36a5c1-318b-4044-b772-3322b34f97d5 · inbound

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree cites this paper.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.772067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.772067Z digest=sha256:578db9a21ddab01ab10636eaa832ae4a5c382e1a47571ce2820f133dd30ee7ae

Observation 0130fc75-4620-40e0-a89c-b6e3cd17348b · inbound

A Survey on Large Language Model Acceleration based on KV Cache Management cites this paper.

A Survey on Large Language Model Acceleration based on KV Cache Management Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T00:38:46.726696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:38:46.726696Z digest=sha256:01c6f65d80f296f63f159b5d67aa61c7461a6ef473ae43c3ce5d7956809e2db5

Observation e84f72ca-173d-40ef-ae74-e116df59bb97 · inbound

Small Language Models (SLMs) Can Still Pack a Punch: A survey (updated 2026) cites this paper.

Small Language Models (SLMs) Can Still Pack a Punch: A survey (updated 2026) Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 137

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:52:37.555244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-23T05:47:48.488826Z digest=sha256:88228f55b1fd902a90ff46e198ceea032a6b6e4d557c0deebe2c2bc927d42229

Observation d4d370f4-c46f-424b-98f5-db247d902e32 · inbound

Reward-Guided Speculative Decoding for Efficient LLM Reasoning cites this paper.

Reward-Guided Speculative Decoding for Efficient LLM Reasoning Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-09T20:42:44.390745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:42:44.390745Z digest=sha256:d5d9a0751c0f412d5ea2463507b1bfe448bb8448089132e552c6bd757db95934

Observation 771287dc-540e-4263-a3ca-95401869bafc · inbound

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation cites this paper.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.795746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.795746Z digest=sha256:2339fd82bbe698eacf52870d9d9fdcf11da89c79c534f56f11b5e8996b0c5b45

Observation a87a8082-5fe0-4a6c-a082-c99ae2dc746a · inbound

Reflection-Window Decoding: Text Generation with Selective Refinement cites this paper.

Reflection-Window Decoding: Text Generation with Selective Refinement Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T04:13:44.424825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:13:44.424825Z digest=sha256:662ff3d0680b70a9733b5120cd2f291b910b1e2104b326903ba0ec887f0a18d1

Observation 8f738c5d-bf30-4c14-9d3d-1108f790f33a · inbound

Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE cites this paper.

Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T16:12:28.993285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:12:28.993285Z digest=sha256:1d2175028bc8fd5fd728ab3aa583c06db04eac25fc94879b7d3374baa70cdf65

Observation 4b6a643e-a163-4897-818a-5763be994129 · inbound

Speculate, then Collaborate: Fusing Knowledge of Language Models during Decoding cites this paper.

Speculate, then Collaborate: Fusing Knowledge of Language Models during Decoding Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T11:11:44.653199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:11:44.653199Z digest=sha256:5babbb4482e314a3c956cc6d38edeaca5810796a372c2d61a3ff723549566a11

Observation c676fc58-5f5d-4722-a95e-3793b120a0b4 · inbound

Taming the Titans: A Survey of Efficient LLM Inference Serving cites this paper.

Taming the Titans: A Survey of Efficient LLM Inference Serving Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 123

Resolution
unresolved
no resolver link, observed 2026-08-16T05:49:10.203035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:49:10.203035Z digest=sha256:d29288ed6dea2582f36d3d0736e49c6e888c0314acbc289ec28e0d237ea2e9a6

Observation 8f392512-3f93-459d-94a8-510b05d9ea67 · inbound

Automatic Task Detection and Heterogeneous LLM Speculative Decoding cites this paper.

Automatic Task Detection and Heterogeneous LLM Speculative Decoding Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T21:57:56.200560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:57:56.200560Z digest=sha256:8bffa065b92b2276745fed4c2834972dc23631bbce8262d12b3702f6a1793e82

Observation 351ee2cc-f9ba-452b-b620-bb0869482f2a · inbound

VeriThinker: Learning to Verify Makes Reasoning Model Efficient cites this paper.

VeriThinker: Learning to Verify Makes Reasoning Model Efficient Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:14.802495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:14.802495Z digest=sha256:9e9ca8d35116b10d54c59b2d69a76372124770a329d6d32921941c1fddad929b

Observation 3aea8502-c3da-49e5-bd95-5429f5c2cbae · inbound

Think Before You Accept: Semantic Reflective Verification for Faster Speculative Decoding cites this paper.

Think Before You Accept: Semantic Reflective Verification for Faster Speculative Decoding Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:42.749845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:42.749845Z digest=sha256:519c0aa667e4d9cd8fe3485b2d108fb03406794106fb7cbf1964e920b9a88576

Observation f9be1957-6551-4f7e-b129-df13e605b179 · inbound

Consultant Decoding: Yet Another Synergistic Mechanism cites this paper.

Consultant Decoding: Yet Another Synergistic Mechanism Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:30:57.001513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:30:57.001513Z digest=sha256:fc4cae6ce589ec63a14d469f501d53675fa66ef02317e3b90b31fae3b396c69e

Observation bc2d3b77-700d-4849-9bce-2a0b5b5c1298 · inbound

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism cites this paper.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.384667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.384667Z digest=sha256:697bdab0fb9fac1f30fd2c29707a8e32789a7a5088e4112783b64f4e6cf9d29b

Observation 03a5e258-919c-4c63-9356-739963217701 · inbound

S$^4$C: Speculative Sampling with Syntactic and Semantic Coherence for Efficient Inference of Large Language Models cites this paper.

S$^4$C: Speculative Sampling with Syntactic and Semantic Coherence for Efficient Inference of Large Language Models Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:22:46.418338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:22:46.418338Z digest=sha256:a0ab5b712f71fc9cc27e854a695964b4be8b45ef68a89a62bc1185f84719d2dc

Observation 98267ba7-8eeb-4aeb-9e78-24ce97bd8e69 · inbound

DeltaLLM: A Training-Free Framework Exploiting Temporal Sparsity for Efficient Edge LLM Inference cites this paper.

DeltaLLM: A Training-Free Framework Exploiting Temporal Sparsity for Efficient Edge LLM Inference Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T14:17:44.088718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:17:44.088718Z digest=sha256:7df15a79138af8f961b649dd159401a7d17b5b775296e7a37bdd9b0f04bb534c

Observation 4e49a438-df9f-4712-870a-40c447917794 · inbound

When RL Meets Adaptive Speculative Training: A Unified Training-Serving System cites this paper.

When RL Meets Adaptive Speculative Training: A Unified Training-Serving System Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:37:28.884383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T06:33:41.860803Z digest=sha256:f6f569f160b5105b10fcc28aac4aef5632a0895f590794524e078657109de981

Observation 6291821a-e0de-4494-8be3-44ba6674a1fc · inbound

ECHO: Elastic Speculative Decoding with Sparse Gating for High-Concurrency Scenarios cites this paper.

ECHO: Elastic Speculative Decoding with Sparse Gating for High-Concurrency Scenarios Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:10:03.112922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T14:09:59.938676Z digest=sha256:b17f8127d268c4e49094e28f89dc65edc9eba0892017e51232d9e54f0b2dc2f3

Observation bbc09dcd-f2d6-4dd4-8436-93db38143891 · inbound

FASER: Fine-Grained Phase Management for Speculative Decoding in Dynamic LLM Serving cites this paper.

FASER: Fine-Grained Phase Management for Speculative Decoding in Dynamic LLM Serving Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-09T22:59:17.685404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-09T22:56:19.734262Z digest=sha256:6d38b81396403ac65b5cca30602fcc4ed6ebec7a11042d08695723d269333c52

Observation e609ecc1-d828-49f8-8e06-529c24ee3a5a · inbound

Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities cites this paper.

Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 169

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:16:10.385717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T09:45:57.201837Z digest=sha256:aff345f6ff97396efa4680267b5b60886932818c3e2ba30fa86307f01ba96484

Observation 80a3d0c8-c0ae-4104-a2ef-b64929f2a1a5 · inbound

When Hidden States Drift: Can KV Caches Rescue Long-Range Speculative Decoding? cites this paper.

When Hidden States Drift: Can KV Caches Rescue Long-Range Speculative Decoding? Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:21:26.630154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-07T11:08:25.132013Z digest=sha256:2a5fcbb391de8cd46d76043218fe2a54e9886d03cab70d0cacc0f4577433bb83

Observation 45b71c1b-b261-4f19-8aba-346a2c817270 · inbound

When Hidden States Drift: Can KV Caches Rescue Long-Range Speculative Decoding? cites this paper.

When Hidden States Drift: Can KV Caches Rescue Long-Range Speculative Decoding? Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:36:23.953655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T00:59:11.690911Z digest=sha256:661095584c7c2bf7a4bd70c647b263e3eaf556ba784e640bc6aa0f2c79a06e83

Observation c3f4aeeb-d47d-416f-8446-7e6dc91f13e3 · inbound

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration cites this paper.

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:51:30.281087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T03:51:52.375703Z digest=sha256:3b291508fc8650c5698dd2497d54106e064603235df5b0c157915d758b38263e

Observation 8313ebec-ae47-469f-8d9b-6f9299776a4f · inbound

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration cites this paper.

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:15:03.451418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T05:11:32.053440Z digest=sha256:72007086233f2ed1062464ce2d2963e5d4576e50a2c51ef1bd30d839af88d1aa

Observation 86141a99-b4a2-44e1-9db8-182f15452ebe · inbound

D-PACE: Dynamic Position-Aware Cross-Entropy for Parallel Speculative Drafting cites this paper.

D-PACE: Dynamic Position-Aware Cross-Entropy for Parallel Speculative Drafting Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:59:06.117035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T21:56:42.264380Z digest=sha256:cf1455cda5cd738b7cc174c8c24bf7ccf517952f4f88bbd1dc7221cc5d2c66e4

Observation 8b7018ec-e240-42a5-84ee-d0a9083ec63c · inbound

Draft Less, Retrieve More: Hybrid Tree Construction for Speculative Decoding cites this paper.

Draft Less, Retrieve More: Hybrid Tree Construction for Speculative Decoding Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:03:06.426899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T06:58:07.996335Z digest=sha256:1330bad9af9a7126f40915486c9802a83fb8308ffac135573dd156aa21c2e5ec

Observation def9a5c9-d4c8-4a89-af31-72992f9fc9d2 · inbound

RTP-LLM: High-Performance Alibaba LLM Inference Engine cites this paper.

RTP-LLM: High-Performance Alibaba LLM Inference Engine Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:52:49.154232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T23:52:40.763228Z digest=sha256:137b9b921d81185dab02b951100863787589fea8ca79bb6a6e82852bad196a01

Observation eddf3414-193d-46a8-ac71-c15c7f02f5c3 · inbound

SURF: Separation via Unsupervised Remixing Flow cites this paper.

SURF: Separation via Unsupervised Remixing Flow Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 185

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T10:16:52.609778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T05:07:10.235599Z digest=sha256:798d70dcf152cfd8ca242e863232c17e0302ef575e27b48a47293c2e2e25c7a4

Observation 03d59875-df8e-4692-a416-a10eb5d24238 · inbound

Speculation at a Distance: Where Edge-Cloud Speculative Decoding Actually Pays Off cites this paper.

Speculation at a Distance: Where Edge-Cloud Speculative Decoding Actually Pays Off Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-04T18:50:04.386306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-25T22:31:27.598322Z digest=sha256:b0e9f861c5f1072c08784c7c5317bacc806cf6fd2a52d0bf41f57acc06d56435

Observation 9a768270-878e-4961-9570-8d80e02cc279 · inbound

D-cut: Adaptive Verification Depth Pruning for Batched Speculative Decoding cites this paper.

D-cut: Adaptive Verification Depth Pruning for Batched Speculative Decoding Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:46.570419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:34:46.570419Z digest=sha256:5ac5c8bfbc5c9cdf102dced62e12901ab755eec857f46c37a6009808977f2570

Observation 3d44352d-c718-4c95-81c8-019008783cb1 · inbound

PyroDash: Cost-Efficient Token-Level Small-Large Language Model Collaborative Inference cites this paper.

PyroDash: Cost-Efficient Token-Level Small-Large Language Model Collaborative Inference Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T10:14:14.991489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:14:14.991489Z digest=sha256:33d05bbd8b77463f57174f0fc52281cf33dc2bb8fb7a2cd8eeb6add32a89eab5

Observation 9b1e2b73-e24a-4ae4-9db5-7ddffd24fded · inbound

AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding cites this paper.

AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T01:22:13.844047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:22:13.844047Z digest=sha256:a4ca809ba1c9b05692b97bbbc65c05578ec37da76e651b29364d17dedaf64f18

Observation 5389dc48-786e-454d-9406-0d90fd08b364 · inbound

Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes cites this paper.

Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T12:06:58.268790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:06:58.268790Z digest=sha256:9531d626a367b14b0678ee3655da6d93a7d8ee82d0f367f1abaf6ced7b8892ac

Observation 55687964-a304-4702-91dc-45e4facf98c1 · inbound

SPADE: Speculative Decoding for Precise and Low Cost Distributed Edge Cloud Inference cites this paper.

SPADE: Speculative Decoding for Precise and Low Cost Distributed Edge Cloud Inference Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T17:23:50.726665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:23:50.726665Z digest=sha256:ee7fe657727078f4a7c716413232d5cafc65ceb0e82db88bff7dfaa7cfb37fa6