Pith. sign in

Paper Citation Record · LEDGER

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency

As of 19 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 0 inbound Pith citation observations for arXiv:2505.17074.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17074 v1

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:11:02.936052Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

34 of 34 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 42d5a527-ab7b-4ab9-8743-244263d7e28b · outbound

This paper cites Tam- ing throughput-latency tradeoff in llm inference with sarathi-serve.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Tam- ing throughput-latency tradeoff in llm inference with sarathi-serve

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.368899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:11:02.785380Z digest=sha256:02dfb3bfa2b7748c926e2bae2a9c148dc98ab81166a974bfea26725a9302dce3

Observation 9a25e2fb-e788-40b3-b8a5-063e15e7b745 · outbound

This paper cites Language models are few-shot learners.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Language models are few-shot learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:02.807182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:02.807182Z digest=sha256:15d532ac90c320f5fce1c59e0d78d3c873895828fb70d4ccf6a8ab8b723b8b22

Observation 8a0c9b30-9bfa-472d-9dfc-1f4b53200f03 · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Accelerating Large Language Model Decoding with Speculative Sampling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:02.817294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:02.817294Z digest=sha256:01b54668781d258cf4230708f1b0cf0ed834790a19cf5e090f10ddcfd0f2c5bc

Observation 31283eb5-b362-4c02-b03b-7dab19401254 · outbound

This paper cites Luan, Zhou Su, and Jing Deng.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Luan, Zhou Su, and Jing Deng

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.315834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:11:02.825252Z digest=sha256:d3e142e3756542ece2c66c4a05a790f45e8474c9f7edf83800761229b2c448cc

Observation 869aa006-063e-4a71-99e2-52cb0abc13e4 · outbound

This paper cites SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:02.834235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:02.834235Z digest=sha256:229cb96541531be6f99fd5554e2a1269a563c305160acf9507834c9032afe958

Observation 301a5537-351b-43a9-8537-4fcb47795eab · outbound

This paper cites Saving GPU Hours in LLM Inference System Development and Online Workloads with Simulation and DBMS-Inspired Cache Replacement Policies.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Saving GPU Hours in LLM Inference System Development and Online Workloads with Simulation and DBMS-Inspired Cache Replacement Policies

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:02.838412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:02.838412Z digest=sha256:f6dd42eab580e452b826b1fa43c688b9585f2827e668f933f7f7f2f7de97d795

Observation aded8af9-435f-40c4-b563-86738ff85e69 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Efficient memory management for large language model serving with pagedattention

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.294279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:11:02.844144Z digest=sha256:2b8cc3bac5650be6a9ba008361a785d431b59b3b02741b70346830ec1356fee3

Observation 7ef105d9-26f5-4bb8-83a3-fbbb544b1ef4 · outbound

This paper cites Incorporating spec- ulative execution into scheduling of control-flow-intensive designs.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Incorporating spec- ulative execution into scheduling of control-flow-intensive designs

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.281661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:11:02.848432Z digest=sha256:debffcebfda32e6a7ee2340953c7ab1c0266b8824ae87b8e85f863650c33e19e

Observation ddd5ed51-0972-44ff-8844-db2124302450 · outbound

This paper cites Al- paserve: Statistical multiplexing with model parallelism for deep learning serving.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Al- paserve: Statistical multiplexing with model parallelism for deep learning serving

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.259600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:11:02.855382Z digest=sha256:5079e099f8ce928e46b824c53ce7df7a3e37e0a77d8c31677f0b76e1d81fb3b2

Observation 9f98bf74-a110-4dce-994d-1b19faa538de · outbound

This paper cites Specpim: Accel- erating speculative inference on pim-enabled system via architecture-dataflow co-exploration.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Specpim: Accel- erating speculative inference on pim-enabled system via architecture-dataflow co-exploration

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.248622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:11:02.859162Z digest=sha256:f6e3a60962ab9b7c7c54189b586ae0b7197d8375497395212c8f551f0ab5ccbc

Observation d7ef7902-2662-4bc2-be5c-3e0e13f69e92 · outbound

This paper cites TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:02.863388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:02.863388Z digest=sha256:6b7ff473097542bb209bc48020284ccd75da9cc742abd96356215a2f150d5413

Observation b65a138e-aa24-41dc-9fd6-eb007cd29222 · outbound

This paper cites Specinfer: Accelerating large language model serv- ing with tree-based speculative inference and verification.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Specinfer: Accelerating large language model serv- ing with tree-based speculative inference and verification

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.236879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:11:02.867632Z digest=sha256:47f424dbdefe7ecd820d98c904a25bd56b22a2ff0024735094712fa669d6e41a

Observation 90743ccd-240a-4d0b-a89e-1cc33b704102 · outbound

This paper cites Exegpt: Constraint-aware resource scheduling for llm inference.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Exegpt: Constraint-aware resource scheduling for llm inference

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.225889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:11:02.871684Z digest=sha256:cfbbc2640c6a7a8aac9050090b574ebbc58f6705e3b3a9746c3bceedc5554a03

Observation 296aaf70-217a-475d-b61c-66badd4cbc96 · outbound

This paper cites Splitwise: Efficient generative llm infer- ence using phase splitting.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Splitwise: Efficient generative llm infer- ence using phase splitting

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.214905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:11:02.877104Z digest=sha256:c59f1412b5f1bf5bd4b69aec6435694a3d4899647d2bed93eb526c9a1553501c

Observation 4bb9ad2f-1c62-47ad-8acf-56a4bf3d2ade · outbound

This paper cites Efficient interactive llm serving with proxy model-based sequence length prediction.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Efficient interactive llm serving with proxy model-based sequence length prediction

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.203413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:11:02.881471Z digest=sha256:caf43401fe4894845cbef4524a65ba6c24a8b9b72a230f7f568b40f2728d0235

Observation 89138bb7-517d-4086-aeab-928cdd3abadd · outbound

This paper cites Analysis of las scheduling for job size distributions with high variance.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Analysis of las scheduling for job size distributions with high variance

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.190453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:11:02.885616Z digest=sha256:26556cabe5404a7d70e91026c3d953e39674a5e5f8f2472703e01a743ae23794

Observation 178b3282-06a9-4c07-be6f-fc38ab9859c7 · outbound

This paper cites Specexec: Massively parallel speculative de- coding for interactive LLM inference on consumer de- vices.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Specexec: Massively parallel speculative de- coding for interactive LLM inference on consumer de- vices

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.163232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:11:02.896303Z digest=sha256:019248b1ffb8f22524c5154f73e6121c224847b9921290e7a362a6f6803228a8

Observation 14604ae9-a14e-4f5d-8f30-adad5c27cf0b · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency LLaMA: Open and Efficient Foundation Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:02.899697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:02.899697Z digest=sha256:710751b23fd7c8774341998103da16653e3297173944895d38e7e78de931196e

Observation 61a5eb0a-65c0-45c4-8197-a90c3fe99502 · outbound

This paper cites Minions: Accelerating Large Language Model Inference with Aggregated Speculative Execution.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Minions: Accelerating Large Language Model Inference with Aggregated Speculative Execution

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:02.903168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:02.903168Z digest=sha256:504dc1e3c096f0dd25f0ac8351b8010c9c07f8c884f89ea32f2f845c487c4dd9

Observation 9b312937-4012-43cf-bcfa-512bafb608ca · outbound

This paper cites Fast Distributed Inference Serving for Large Language Models.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Fast Distributed Inference Serving for Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:02.907610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:02.907610Z digest=sha256:56a9fb2a13daaa0bc6e55f8997d99f5a8ed8b8baafc459f729fcd5c2c7529b31

Observation 0f6e6be7-143a-463a-8bd3-bb5065bb8bfa · outbound

This paper cites Mini- thinky dataset.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Mini- thinky dataset

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.149262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:11:02.912226Z digest=sha256:1e5d855434db14ba29baa397de84f1bbac87617cc94d40e29fdf0ffac8354df6

Observation c2f06694-8e64-48aa-9b9d-ee68be606a25 · outbound

This paper cites PerLLM: Personalized Inference Scheduling with Edge-Cloud Collaboration for Diverse LLM Services.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency PerLLM: Personalized Inference Scheduling with Edge-Cloud Collaboration for Diverse LLM Services

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:02.915966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:02.915966Z digest=sha256:1a88c9810ce96dd3564cb1a095f65ad269ed116d8dd98abbcd21f5fc1f96a016

Observation 8bc61a95-1e31-42f2-b78a-040c81c82a86 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Tree of thoughts: Deliberate problem solving with large language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.135718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:11:02.920303Z digest=sha256:de49748d1824c360741cbbecc937b842a918bafe67c6bf130ca20ac5e0e6b97e

Observation d45109dd-5775-4e1d-a6ca-cebc8f2d65d9 · outbound

This paper cites Adap- tive batch budget for llm inference.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Adap- tive batch budget for llm inference

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.122914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:11:02.923723Z digest=sha256:e474777f186e7d43aab06fea963a8b95888da5fa791660d7950c77a927762b7f

Observation 04a7b48d-01fa-4a22-aef5-f541461b2fbe · outbound

This paper cites Re- sponse length perception and sequence scheduling: An llm-empowered llm inference pipeline.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Re- sponse length perception and sequence scheduling: An llm-empowered llm inference pipeline

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.110445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:11:02.927453Z digest=sha256:d5885207af5fbee39d3ce078b0c20e9729969e7b61512d46e5c87380e3400359

Observation 45407d40-2d41-471f-b43a-9ed12450fee1 · outbound

This paper cites Response length perception and sequence scheduling: An llm-empowered llm inference pipeline.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Response length perception and sequence scheduling: An llm-empowered llm inference pipeline

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.099166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:11:02.931101Z digest=sha256:40d2ae56ae269565bf49ab124b535d6e1d28341b2b09388a288beeb13750ccd2

Observation fb78d568-5fe3-4870-8910-1e9137894864 · outbound

This paper cites Toolqa: A dataset for llm question answering with external tools.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Toolqa: A dataset for llm question answering with external tools

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.085461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:11:02.936052Z digest=sha256:11080028245f325d15ad9ea7179a937201b72efc8949efa7bcf092f291641757

Observation 76187568-019d-4aa1-8ebb-82dd39117b18 · outbound

This paper cites Fast inference from transformers via speculative decoding.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Fast inference from transformers via speculative decoding

Reference 2000

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.271405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:11:02.851841Z digest=sha256:5d3320d8bfdee080d86da767835ba1de66da8c0399522fb76fd1d53e652c1812

Observation 3cb0c213-b22a-45f9-affa-f5a5d523fd6e · outbound

This paper cites Accelerating LLM inference with staged speculative decoding.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Accelerating LLM inference with staged speculative decoding

Reference 2003

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.177782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:11:02.891871Z digest=sha256:32c764a4821f0ea3f70d907648ab8ca5971233ce11a61b57f0aa2e6cda483446

Observation d23c522d-39b5-4e2d-ba31-d39aaf79a327 · outbound

This paper cites Lee, Deming Chen, and Tri Dao.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Lee, Deming Chen, and Tri Dao

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.326987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:11:02.812256Z digest=sha256:244e836f4f5c4de55a1b711996c0b03f9b13967c8662b12363196a15e2467f93

Observation 57646099-8362-43cc-a5a8-75c58368988f · outbound

This paper cites Speculative streaming: Fast llm infer- ence without auxiliary models.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Speculative streaming: Fast llm infer- ence without auxiliary models

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.345380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:11:02.801375Z digest=sha256:61a7044b046dc8f8ed07d19cda86531ed0398cbcc7ed56c35e3982ed88a0bee1

Observation 57c104ff-bea9-483e-bdb8-b21201bdca93 · outbound

This paper cites Program Synthesis with Large Language Models.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Program Synthesis with Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:02.796594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:02.796594Z digest=sha256:78561822f74f2a52d2f025765ece6ebfad14e8eaf8d41b32f78ced5ccd2258b7

Observation e85c7fc4-dc85-4928-b59a-e47dcd8f6101 · outbound

This paper cites Chatbot instruc- tion prompts.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Chatbot instruc- tion prompts

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.356812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:11:02.792451Z digest=sha256:57c7cfde4152015a57ab15967d4419930e8921c9f301cf7168237f72d1888434

Observation b0aff8b5-208e-482c-8364-7401f37171f3 · outbound

This paper cites Glide with a cape: A low-hassle method to accelerate speculative decoding.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Glide with a cape: A low-hassle method to accelerate speculative decoding

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.305348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:11:02.830448Z digest=sha256:c30cd7883ee305ce22ce40fa586f9e57b582d7d44a353a5e670a4999a537acac

Pith citing papers

No inbound Pith citation observations are available.