Pith. sign in

Paper Citation Record · LEDGER

Cascade Speculative Drafting for Even Faster LLM Inference

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2312.11462.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.11462 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:44:46.574106Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T23:23:50.781491Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5c67dcdf-0e97-430d-b66c-3414d1333135 · inbound

EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty cites this paper.

EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty Cascade Speculative Drafting for Even Faster LLM Inference

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:15:49.372771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-15T00:15:49.303458Z digest=sha256:e7f9c1604042005e8418787781da97db411290d1835bd622dbc7c8c1234dcd47

Observation 2984eb47-39a1-4714-864b-ee6bb00366f2 · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models Cascade Speculative Drafting for Even Faster LLM Inference

Reference 245

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:39:33.371517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:53378afba1187fdbf79efa1a20a7ac1da3e2ee0f48790c039282af8ca176ba3e

Observation a6e5a1e9-63aa-4142-8490-bbb910c71888 · inbound

FastDraft: How to Train Your Draft cites this paper.

FastDraft: How to Train Your Draft Cascade Speculative Drafting for Even Faster LLM Inference

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T19:05:25.929523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:05:25.929523Z digest=sha256:c019c67ad325f2edb6ff2724fa0f21f30adff0487491a8a4cb2bab760b5e8d02

Observation dbb37119-5ca2-498f-9683-4fe5acd61444 · inbound

Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration cites this paper.

Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration Cascade Speculative Drafting for Even Faster LLM Inference

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T13:15:07.159423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:15:07.159423Z digest=sha256:62483130035b37f9d232f40476bc935b82e64e52d7beab45dd00f2435b63e598

Observation 85d3b22d-8c13-4fad-9bc1-0a8201207dca · inbound

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree cites this paper.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Cascade Speculative Drafting for Even Faster LLM Inference

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.580987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.580987Z digest=sha256:5556c16bd0e6c4f10927822f537f7af6076d79484cf5386be178cd75aeaf9ab4

Observation 0bc4f55b-fcdc-4c32-a090-e565eb9668dd · inbound

Reward-Guided Speculative Decoding for Efficient LLM Reasoning cites this paper.

Reward-Guided Speculative Decoding for Efficient LLM Reasoning Cascade Speculative Drafting for Even Faster LLM Inference

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T20:42:44.031679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:42:44.031679Z digest=sha256:292b83f01fbce014f62e1728d79459d7a34d820501260a94dd2f63fb8329f51a

Observation 53889ecc-1320-451a-982e-237acd9c79ca · inbound

Towards Harnessing the Collaborative Power of Large and Small Models for Domain Tasks cites this paper.

Towards Harnessing the Collaborative Power of Large and Small Models for Domain Tasks Cascade Speculative Drafting for Even Faster LLM Inference

Reference 132

Resolution
unresolved
no resolver link, observed 2026-08-16T10:44:46.574106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:44:46.574106Z digest=sha256:6fae9215b3f8abf315db6ac36f192ce6bd2086511c3349608c755b3b4cff3fd2

Observation ddcc36ca-767a-435b-9017-6b20126369aa · inbound

Accelerating Large Language Model Reasoning via Speculative Search cites this paper.

Accelerating Large Language Model Reasoning via Speculative Search Cascade Speculative Drafting for Even Faster LLM Inference

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T04:17:36.542450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:17:36.542450Z digest=sha256:2f32a43c82786a74c764e5be33ac08c668b0696573766edcdf7c6f3cce2ab971

Observation b78b656c-fd33-4ea8-b37c-61b10237bb90 · inbound

Automatic Task Detection and Heterogeneous LLM Speculative Decoding cites this paper.

Automatic Task Detection and Heterogeneous LLM Speculative Decoding Cascade Speculative Drafting for Even Faster LLM Inference

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T21:57:55.892172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:57:55.892172Z digest=sha256:2a0a4a50865b5001b36a0361a05fdfebb2d8087cb19e509341b3eff12de98337

Observation f2536f8f-5d7d-4163-bbe9-7a879f289637 · inbound

Reasoning Can Be Restored by Correcting a Few Decision Tokens cites this paper.

Reasoning Can Be Restored by Correcting a Few Decision Tokens Cascade Speculative Drafting for Even Faster LLM Inference

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:57:46.878572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-19T20:56:48.771058Z digest=sha256:fe99daba57d86af2bac3794b3ba4bd9fcdaeacfc6545ec38a1f0f56b200d6834

Observation 9ff69a76-141f-4466-9cd4-d80533d3c6df · inbound

UCCI: Calibrated Uncertainty for Cost-Optimal LLM Cascade Routing cites this paper.

UCCI: Calibrated Uncertainty for Cost-Optimal LLM Cascade Routing Cascade Speculative Drafting for Even Faster LLM Inference

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:23:50.785877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-20T23:23:45.561978Z digest=sha256:53d47adab8a54a3a35e7665074c08d29d4496e576346528239f0dfe141f3453d

Observation c91311fc-c99d-4751-b233-5b0f3821ab70 · inbound

Leaky Language Models: Stealing Architecture and Inference Optimizations via Per-Token Timing cites this paper.

Leaky Language Models: Stealing Architecture and Inference Optimizations via Per-Token Timing Cascade Speculative Drafting for Even Faster LLM Inference

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T09:35:53.518187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T09:35:53.518187Z digest=sha256:d8fa21678c9402ec10897c25ababdf8e0197e87b4ef67d5645b1f4bab1edee14