Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:11:02.936052Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 0 inbound Pith citation observations for arXiv:2505.17074.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:11:02.936052Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
34 of 34 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 42d5a527-ab7b-4ab9-8743-244263d7e28b · outbound
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Tam- ing throughput-latency tradeoff in llm inference with sarathi-serve
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9a25e2fb-e788-40b3-b8a5-063e15e7b745 · outbound
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Language models are few-shot learners
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a0c9b30-9bfa-472d-9dfc-1f4b53200f03 · outbound
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Accelerating Large Language Model Decoding with Speculative Sampling
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31283eb5-b362-4c02-b03b-7dab19401254 · outbound
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Luan, Zhou Su, and Jing Deng
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 869aa006-063e-4a71-99e2-52cb0abc13e4 · outbound
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 301a5537-351b-43a9-8537-4fcb47795eab · outbound
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Saving GPU Hours in LLM Inference System Development and Online Workloads with Simulation and DBMS-Inspired Cache Replacement Policies
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aded8af9-435f-40c4-b563-86738ff85e69 · outbound
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Efficient memory management for large language model serving with pagedattention
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7ef105d9-26f5-4bb8-83a3-fbbb544b1ef4 · outbound
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Incorporating spec- ulative execution into scheduling of control-flow-intensive designs
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ddd5ed51-0972-44ff-8844-db2124302450 · outbound
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Al- paserve: Statistical multiplexing with model parallelism for deep learning serving
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9f98bf74-a110-4dce-994d-1b19faa538de · outbound
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Specpim: Accel- erating speculative inference on pim-enabled system via architecture-dataflow co-exploration
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d7ef7902-2662-4bc2-be5c-3e0e13f69e92 · outbound
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b65a138e-aa24-41dc-9fd6-eb007cd29222 · outbound
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Specinfer: Accelerating large language model serv- ing with tree-based speculative inference and verification
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 90743ccd-240a-4d0b-a89e-1cc33b704102 · outbound
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Exegpt: Constraint-aware resource scheduling for llm inference
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 296aaf70-217a-475d-b61c-66badd4cbc96 · outbound
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Splitwise: Efficient generative llm infer- ence using phase splitting
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4bb9ad2f-1c62-47ad-8acf-56a4bf3d2ade · outbound
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Efficient interactive llm serving with proxy model-based sequence length prediction
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 89138bb7-517d-4086-aeab-928cdd3abadd · outbound
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Analysis of las scheduling for job size distributions with high variance
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 178b3282-06a9-4c07-be6f-fc38ab9859c7 · outbound
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Specexec: Massively parallel speculative de- coding for interactive LLM inference on consumer de- vices
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 14604ae9-a14e-4f5d-8f30-adad5c27cf0b · outbound
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency LLaMA: Open and Efficient Foundation Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61a5eb0a-65c0-45c4-8197-a90c3fe99502 · outbound
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Minions: Accelerating Large Language Model Inference with Aggregated Speculative Execution
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b312937-4012-43cf-bcfa-512bafb608ca · outbound
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Fast Distributed Inference Serving for Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f6e6be7-143a-463a-8bd3-bb5065bb8bfa · outbound
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Mini- thinky dataset
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c2f06694-8e64-48aa-9b9d-ee68be606a25 · outbound
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency PerLLM: Personalized Inference Scheduling with Edge-Cloud Collaboration for Diverse LLM Services
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bc61a95-1e31-42f2-b78a-040c81c82a86 · outbound
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Tree of thoughts: Deliberate problem solving with large language models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d45109dd-5775-4e1d-a6ca-cebc8f2d65d9 · outbound
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Adap- tive batch budget for llm inference
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 04a7b48d-01fa-4a22-aef5-f541461b2fbe · outbound
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Re- sponse length perception and sequence scheduling: An llm-empowered llm inference pipeline
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 45407d40-2d41-471f-b43a-9ed12450fee1 · outbound
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Response length perception and sequence scheduling: An llm-empowered llm inference pipeline
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation fb78d568-5fe3-4870-8910-1e9137894864 · outbound
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Toolqa: A dataset for llm question answering with external tools
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 76187568-019d-4aa1-8ebb-82dd39117b18 · outbound
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Fast inference from transformers via speculative decoding
Reference 2000
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3cb0c213-b22a-45f9-affa-f5a5d523fd6e · outbound
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Accelerating LLM inference with staged speculative decoding
Reference 2003
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d23c522d-39b5-4e2d-ba31-d39aaf79a327 · outbound
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Lee, Deming Chen, and Tri Dao
Reference 2020
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 57646099-8362-43cc-a5a8-75c58368988f · outbound
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Speculative streaming: Fast llm infer- ence without auxiliary models
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 57c104ff-bea9-483e-bdb8-b21201bdca93 · outbound
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Program Synthesis with Large Language Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e85c7fc4-dc85-4928-b59a-e47dcd8f6101 · outbound
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Chatbot instruc- tion prompts
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b0aff8b5-208e-482c-8364-7401f37171f3 · outbound
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency Glide with a cape: A low-hassle method to accelerate speculative decoding
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
No inbound Pith citation observations are available.