Pith. sign in

Paper Citation Record · LEDGER

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency

As of 17 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 0 inbound Pith citation observations for arXiv:2505.17074.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17074 v1

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:11:02.936052Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

34 of 34 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 42d5a527-ab7b-4ab9-8743-244263d7e28b · outbound

This paper cites Tam- ing throughput-latency tradeoff in llm inference with sarathi-serve.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Tam- ing throughput-latency tradeoff in llm inference with sarathi-serve

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.368899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:11:02.785380Z digest=sha256:6b36b34ff2bd7e4772a5f9ec049a684aea36bebf62858e7a4533d8e4760ea844

Observation 9a25e2fb-e788-40b3-b8a5-063e15e7b745 · outbound

This paper cites Language models are few-shot learners.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Language models are few-shot learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:02.807182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:02.807182Z digest=sha256:555d2e6c73bba3a7e0bf217cd5041051d45b7131f5083ea3de542a584d666867

Observation 8a0c9b30-9bfa-472d-9dfc-1f4b53200f03 · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Accelerating Large Language Model Decoding with Speculative Sampling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:02.817294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:02.817294Z digest=sha256:2d154bd1bcdb5ac5b4c6741c8c6240dcf98e667cd6db88da7e9454d17eb3cdf1

Observation 31283eb5-b362-4c02-b03b-7dab19401254 · outbound

This paper cites Luan, Zhou Su, and Jing Deng.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Luan, Zhou Su, and Jing Deng

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.315834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:11:02.825252Z digest=sha256:8e38b0097615b6436abc99d5d64eb2a089008651a3a33552bff26ab9cc0041a1

Observation 869aa006-063e-4a71-99e2-52cb0abc13e4 · outbound

This paper cites SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:02.834235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:02.834235Z digest=sha256:a5765d4a3da48fda9b5293193515cd4e155168e2e79b0800f0609d70c5442008

Observation 301a5537-351b-43a9-8537-4fcb47795eab · outbound

This paper cites Saving GPU Hours in LLM Inference System Development and Online Workloads with Simulation and DBMS-Inspired Cache Replacement Policies.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Saving GPU Hours in LLM Inference System Development and Online Workloads with Simulation and DBMS-Inspired Cache Replacement Policies

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:02.838412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:02.838412Z digest=sha256:53122b1ede11437bf670018b247bc6148800fddc9962742219082c6057d97fac

Observation aded8af9-435f-40c4-b563-86738ff85e69 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Efficient memory management for large language model serving with pagedattention

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.294279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:11:02.844144Z digest=sha256:30cfe5c20fdb3bb347d2f1238dc2b31aa16d6fb7850b2055daf32842fee7010c

Observation 7ef105d9-26f5-4bb8-83a3-fbbb544b1ef4 · outbound

This paper cites Incorporating spec- ulative execution into scheduling of control-flow-intensive designs.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Incorporating spec- ulative execution into scheduling of control-flow-intensive designs

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.281661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:11:02.848432Z digest=sha256:4383c4a7241c0e6538fdd4a07d5f1b6366fe0e961eae3c047aea48245fee2c79

Observation ddd5ed51-0972-44ff-8844-db2124302450 · outbound

This paper cites Al- paserve: Statistical multiplexing with model parallelism for deep learning serving.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Al- paserve: Statistical multiplexing with model parallelism for deep learning serving

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.259600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:11:02.855382Z digest=sha256:e00f2ffa524f224d3f48b0f80b6389c58e2d7a2b0b2139ab0ee5eb8d9f2c26c4

Observation 9f98bf74-a110-4dce-994d-1b19faa538de · outbound

This paper cites Specpim: Accel- erating speculative inference on pim-enabled system via architecture-dataflow co-exploration.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Specpim: Accel- erating speculative inference on pim-enabled system via architecture-dataflow co-exploration

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.248622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:11:02.859162Z digest=sha256:e52ff339d4ede297ea324b67d5c97983e9ddde9d082045ac81a756ccde32ac40

Observation d7ef7902-2662-4bc2-be5c-3e0e13f69e92 · outbound

This paper cites TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:02.863388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:02.863388Z digest=sha256:9da9d704d2ed0a82d7f777235e2b0ee69a5e812938948a7ae184e6ca46608706

Observation b65a138e-aa24-41dc-9fd6-eb007cd29222 · outbound

This paper cites Specinfer: Accelerating large language model serv- ing with tree-based speculative inference and verification.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Specinfer: Accelerating large language model serv- ing with tree-based speculative inference and verification

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.236879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:11:02.867632Z digest=sha256:51f6b57597980adf466687bfec150c2e53344e605d10e2f244dfcba9e48a3658

Observation 90743ccd-240a-4d0b-a89e-1cc33b704102 · outbound

This paper cites Exegpt: Constraint-aware resource scheduling for llm inference.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Exegpt: Constraint-aware resource scheduling for llm inference

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.225889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:11:02.871684Z digest=sha256:30c4b69467eec611a169dc16e2fd182da75c485788dfc35ba1b0890a0e466fa9

Observation 296aaf70-217a-475d-b61c-66badd4cbc96 · outbound

This paper cites Splitwise: Efficient generative llm infer- ence using phase splitting.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Splitwise: Efficient generative llm infer- ence using phase splitting

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.214905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:11:02.877104Z digest=sha256:7b174587562681424bb0805e9471379b792564ad47821c6ef00a6182ed8be391

Observation 4bb9ad2f-1c62-47ad-8acf-56a4bf3d2ade · outbound

This paper cites Efficient interactive llm serving with proxy model-based sequence length prediction.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Efficient interactive llm serving with proxy model-based sequence length prediction

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.203413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:11:02.881471Z digest=sha256:9bad03d2c1257ecdfdd01935a7397243a3d1380229a38d72e0232df5faa25c09

Observation 89138bb7-517d-4086-aeab-928cdd3abadd · outbound

This paper cites Analysis of las scheduling for job size distributions with high variance.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Analysis of las scheduling for job size distributions with high variance

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.190453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:11:02.885616Z digest=sha256:e7bd9e620a92b3612e52d94947a6942e2b38e4d97715541c76a1af69b5afb301

Observation 178b3282-06a9-4c07-be6f-fc38ab9859c7 · outbound

This paper cites Specexec: Massively parallel speculative de- coding for interactive LLM inference on consumer de- vices.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Specexec: Massively parallel speculative de- coding for interactive LLM inference on consumer de- vices

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.163232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:11:02.896303Z digest=sha256:26985d96e7e024083d284c93dd02abe8103ce2443eb9da7602685b22ce67f07a

Observation 14604ae9-a14e-4f5d-8f30-adad5c27cf0b · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency LLaMA: Open and Efficient Foundation Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:02.899697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:02.899697Z digest=sha256:9f38f04fed19b34e64256fca9f410ecf1051934a02e8f22eba2115741d637924

Observation 61a5eb0a-65c0-45c4-8197-a90c3fe99502 · outbound

This paper cites Minions: Accelerating Large Language Model Inference with Aggregated Speculative Execution.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Minions: Accelerating Large Language Model Inference with Aggregated Speculative Execution

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:02.903168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:02.903168Z digest=sha256:6e30ae279944ae6a53eee5a461db73cdeff6879a63f2e8a47f27a7d2adf9e8fe

Observation 9b312937-4012-43cf-bcfa-512bafb608ca · outbound

This paper cites Fast Distributed Inference Serving for Large Language Models.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Fast Distributed Inference Serving for Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:02.907610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:02.907610Z digest=sha256:9cd74a524c0b98b4172ba9e98154706263992f0c05776c23e7860cd9bd88d444

Observation 0f6e6be7-143a-463a-8bd3-bb5065bb8bfa · outbound

This paper cites Mini- thinky dataset.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Mini- thinky dataset

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.149262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:11:02.912226Z digest=sha256:9d9b60f7fdb890ab8e3821b039bad1de7eb75c503a0773ea5d3c14ce70700286

Observation c2f06694-8e64-48aa-9b9d-ee68be606a25 · outbound

This paper cites PerLLM: Personalized Inference Scheduling with Edge-Cloud Collaboration for Diverse LLM Services.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency PerLLM: Personalized Inference Scheduling with Edge-Cloud Collaboration for Diverse LLM Services

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:02.915966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:02.915966Z digest=sha256:9c5171706bef48d491877ae29dcfeb753057800e2b1291f45f54a55c94818132

Observation 8bc61a95-1e31-42f2-b78a-040c81c82a86 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Tree of thoughts: Deliberate problem solving with large language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.135718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:11:02.920303Z digest=sha256:cddd81e34ba01f6081fc81305057489e8e60cc4d6876e50f7632fd10bb07b13e

Observation d45109dd-5775-4e1d-a6ca-cebc8f2d65d9 · outbound

This paper cites Adap- tive batch budget for llm inference.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Adap- tive batch budget for llm inference

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.122914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:11:02.923723Z digest=sha256:682e6ed3fa11874c315acd99fd1bfbde26f69ca2b28d59b386b91ae0d9ad92ac

Observation 04a7b48d-01fa-4a22-aef5-f541461b2fbe · outbound

This paper cites Re- sponse length perception and sequence scheduling: An llm-empowered llm inference pipeline.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Re- sponse length perception and sequence scheduling: An llm-empowered llm inference pipeline

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.110445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:11:02.927453Z digest=sha256:0104ffb7da126651b68c5728b7b900bbef470edbf18c54c1fa9c17b881448dcb

Observation 45407d40-2d41-471f-b43a-9ed12450fee1 · outbound

This paper cites Response length perception and sequence scheduling: An llm-empowered llm inference pipeline.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Response length perception and sequence scheduling: An llm-empowered llm inference pipeline

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.099166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:11:02.931101Z digest=sha256:d8e72728526ad5ce2da301b951d9b4d3dfcea438c5398d99702f3cc916deba34

Observation fb78d568-5fe3-4870-8910-1e9137894864 · outbound

This paper cites Toolqa: A dataset for llm question answering with external tools.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Toolqa: A dataset for llm question answering with external tools

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.085461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:11:02.936052Z digest=sha256:cc7e1b6cdaa021962fa3464377f8a9d6ae94340b66278cce14982d6683b8ea76

Observation 76187568-019d-4aa1-8ebb-82dd39117b18 · outbound

This paper cites Fast inference from transformers via speculative decoding.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Fast inference from transformers via speculative decoding

Reference 2000

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.271405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:11:02.851841Z digest=sha256:ee5f1e66222f2b8c62d7d883dcdeb3a3d55949a03b40668dc6162ce2428489ca

Observation 3cb0c213-b22a-45f9-affa-f5a5d523fd6e · outbound

This paper cites Accelerating LLM inference with staged speculative decoding.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Accelerating LLM inference with staged speculative decoding

Reference 2003

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.177782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:11:02.891871Z digest=sha256:af9adb73ebcc81071b789d0013a61eb05dc4f45681f453ef6de855484eb898d0

Observation d23c522d-39b5-4e2d-ba31-d39aaf79a327 · outbound

This paper cites Lee, Deming Chen, and Tri Dao.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Lee, Deming Chen, and Tri Dao

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.326987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:11:02.812256Z digest=sha256:77f091981d376e0ff7fb2e28c56f9b0be94b3f049331148ae1a67bda76fe59bf

Observation 57646099-8362-43cc-a5a8-75c58368988f · outbound

This paper cites Speculative streaming: Fast llm infer- ence without auxiliary models.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Speculative streaming: Fast llm infer- ence without auxiliary models

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.345380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:11:02.801375Z digest=sha256:a86bdff64ae161d4dd98a1a577667bd7f6c5299c8786162b36e0342e0de455d9

Observation 57c104ff-bea9-483e-bdb8-b21201bdca93 · outbound

This paper cites Program Synthesis with Large Language Models.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Program Synthesis with Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T20:11:02.796594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:11:02.796594Z digest=sha256:98fe05c98f280666f7d6777b50c5e62f4f88850581672ec9d4a412c338c9c23d

Observation e85c7fc4-dc85-4928-b59a-e47dcd8f6101 · outbound

This paper cites Chatbot instruc- tion prompts.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Chatbot instruc- tion prompts

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.356812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:11:02.792451Z digest=sha256:3af9c148881c6eb2fcbca04b92b859f9c41d37dd6af499f2ecc98a323ed223f5

Observation b0aff8b5-208e-482c-8364-7401f37171f3 · outbound

This paper cites Glide with a cape: A low-hassle method to accelerate speculative decoding.

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Glide with a cape: A low-hassle method to accelerate speculative decoding

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:11:03.305348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:11:02.830448Z digest=sha256:80017f8eea6fef030cd2c632c3e4c6a687a79db46805225d91696a00568ae16e

Pith citing papers

No inbound Pith citation observations are available.