Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:15:45.012578Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 2 inbound Pith citation observations for arXiv:2506.11309.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:15:45.012578Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:45:04.545787Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T18:45:05.618170Z
51 of 51 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3fb5538b-276a-41a3-b1cb-10b4667fce22 · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Tam- ing Throughput-Latency tradeoff in LLM inference with Sarathi-Serve
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8d7ff1db-4df9-4718-89e9-f354a6150ac2 · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding PipeInfer: Accelerating LLM Inference using Asynchronous Pipelined Speculation
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4592b5d3-f84d-4c94-824e-057e5bf26bf0 · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9963ce5-e8b9-4e6b-8d8b-754c5d1f4f90 · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Large Language Models vs. Search Engines: Evaluating User Preferences Across Varied Information Retrieval Scenarios
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c754f5f-5702-4004-a0aa-5fb1673780cf · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Accelerating Large Language Model Decoding with Speculative Sampling
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdd14e65-f8a2-4532-b295-7d5ed4d5594e · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Evaluating Large Language Models Trained on Code
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96447f0a-0d6b-47a8-b1b3-81e18e68e95e · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Training Verifiers to Solve Math Word Problems
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf6635c0-7cf7-4eaf-a680-5b53ed6b59d1 · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a82ae83-f87f-4e58-95e9-f708bc578699 · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding DAPPLE: A Pipelined Data Parallel Approach for Training Large Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 49df5af8-e974-4540-b019-993b34b03e69 · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding LongCoder: A Long-Range Pre-trained Language Model for Code Completion
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bb237d4-b7dc-4d72-8cdf-2f00cd1d398e · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ef14264-c6a6-49ae-9bcb-ad5d004cddff · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Unlocking the Potential of ChatGPT: A Comprehensive Exploration of its Applications, Advantages, Limitations, and Future Directions in Natural Language Processing
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ab8f421f-7074-471e-a655-4a27f89380e3 · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Le, Yonghui Wu, and Zhifeng Chen
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 87637fcd-8ea2-4e70-8d37-31eb934ce7ee · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Language Models for Code Completion: A Practical Evaluation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46f4ee4b-5f47-4d96-9197-7ef6c6fd25af · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation aceba6d5-f5de-4472-912c-3836a932dde2 · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87745101-6bcf-4d75-8363-4a2416d0daea · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Fast inference from transformers via speculative de- coding, 2023
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3e99ce16-7dd7-4564-bf61-74a784243bc8 · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f1a7488-ce17-4678-a6ba-c130cc8374ba · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c5ef95b-fca0-489f-ba9f-d01467c1509d · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding PEARL: Parallel Speculative Decoding with Adaptive Draft Length
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff29de01-5fa9-470f-b85a-f1ce27a5a5ab · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d04da29a-e91d-4121-9b85-9c9c9a1bacee · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4e72ecd-1580-42b7-a2b4-9e8959a0f456 · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding AMUSD: Asynchronous Multi-Device Speculative Decoding for LLM Acceleration
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a796db2-febf-4c3c-9bad-4b3df007c4fe · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Specinfer: Accelerating large language model serving with tree-based speculative infer- ence and verification
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27849d61-0697-409b-9c04-2c4260d8ae7d · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Abstractive Text Summarization Using Sequence-to-Sequence RNNs and Beyond
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0ab6bb1-7e66-4b45-a031-01bd48dfa83d · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Nvidia/tensorrt-llm: A tensorrt toolbox for optimized large language model inference
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2e8f0325-e1b0-49e7-9a34-bdf2c8fd3340 · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding GPT-4 Technical Report
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 413a2d54-8dab-4c49-b80e-d2ca69d2e918 · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Splitwise: Efficient generative LLM inference using phase splitting
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45eb26d6-6f58-4d29-a6dd-14bc43a42ea6 · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eaa4fb03-50e6-4477-8c55-81b5318b63b5 · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Tatsu-lab/stanford-alpaca: Code and documentation to train stanford’s alpaca mod- els, and generate the data
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b772985c-390f-4f4f-8d2f-06ffb9e55c33 · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding CUTLASS, January
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a5f09fc7-f62e-4b0a-b8fb-43ac80d9e983 · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Llama: Open and efficient foundation language models, 2023
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b4221c45-d13b-498c-9c04-b597e23fd764 · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Gomez, Lukasz Kaiser, and Illia Polosukhin
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79725ce1-bde8-4b1a-b38b-9657020b78eb · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcd83ab8-57ba-4e93-983b-2eeb8c57a941 · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9686e40c-03cc-4790-8961-f2f059f63c0b · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding When Search Engine Services meet Large Language Models: Visions and Challenges
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe5bd438-7080-4835-87e7-d1c32d86c99e · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Qwen2 Technical Report
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e37bbea3-b624-4247-badc-0985b0956d72 · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding PALR: Personalization Aware LLMs for Recommendation
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b6c1c76-ca47-4f89-a99e-90a37ce29d59 · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Orca: A distributed serving system for Transformer-Based generative models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e17bc3b8-1bfe-46e7-9b63-be6229c899cf · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Fltrnn: Faithful long-horizon taskplanningforroboticswithlargelanguagemodels
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5390ded-a5b2-4eb2-9f0b-14ca9aa49a33 · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Prepacking: A simple method for fast prefilling and increased throughput in large language models, 2024
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 64966e2b-fdec-46df-9dff-219b634fbae3 · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5448a914-ed7d-4f69-9198-e5bea3bd5bf5 · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Gonzalez, Clark Barrett, and Ying Sheng
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 45bfc0d8-6903-4663-91a2-b67ab244d231 · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 78327a17-ce95-4d19-8262-07a9811bd1bb · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44f94357-1a95-4a21-8b21-fc9328f7be3d · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding SGLang: Efficient Execution of Structured Language Model Programs
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e7eda31-2e22-4c88-94dc-78f0a44fab71 · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding ISBN 978-1-939133- 40-3
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8b24c817-30aa-4d89-83b7-fb108d0329d0 · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcf9a324-ef29-4afa-a642-0e47de74a6b9 · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding ISBN 978-1-939133- 28-1
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 82fe9fe4-a7e4-402b-8803-a3c012037936 · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Unresolved cited work
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1b9cf4b0-5cc1-448d-b5d8-fdd725b6a0b1 · outbound
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding ISBN 978-1-939133- 40-3
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 73e76fe6-46b5-44f6-a9f2-eec8c7d42f7f · inbound
ODIA: Oriented Distillation for Inline Acceleration of LLM-based Function Calling SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 44023fab-b54f-4e68-933e-80447626be8f · inbound
Unlocking Parallelism in Autoregressive Language Models via Speculative Decoding with Progressive Tree Drafting SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.