Pith. sign in

Paper Citation Record · LEDGER

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding

As of 23 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 2 inbound Pith citation observations for arXiv:2506.11309.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11309 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:15:45.012578Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:45:04.545787Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T18:45:05.618170Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact2
  • verified fuzzy15
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3fb5538b-276a-41a3-b1cb-10b4667fce22 · outbound

This paper cites Tam- ing Throughput-Latency tradeoff in LLM inference with Sarathi-Serve.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Tam- ing Throughput-Latency tradeoff in LLM inference with Sarathi-Serve

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:15:45.562565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T04:15:44.865054Z digest=sha256:e4c894dbd1b2d7dcbaba3fe91d604900fd9d84cc78130195e648dbc4b881c2db

Observation 8d7ff1db-4df9-4718-89e9-f354a6150ac2 · outbound

This paper cites PipeInfer: Accelerating LLM Inference using Asynchronous Pipelined Speculation.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding PipeInfer: Accelerating LLM Inference using Asynchronous Pipelined Speculation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:44.871917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:15:44.871917Z digest=sha256:3c9ea68f50d269f4f8264fdb3ef838598224ce92edc4eb886d84c54cc90ee8fd

Observation 4592b5d3-f84d-4c94-824e-057e5bf26bf0 · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:44.875561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:15:44.875561Z digest=sha256:5681bae2f721df9b7f0c4223e335718123f6ed5577d25814d60d61d143cad4a7

Observation e9963ce5-e8b9-4e6b-8d8b-754c5d1f4f90 · outbound

This paper cites Large Language Models vs. Search Engines: Evaluating User Preferences Across Varied Information Retrieval Scenarios.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Large Language Models vs. Search Engines: Evaluating User Preferences Across Varied Information Retrieval Scenarios

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:44.879263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:15:44.879263Z digest=sha256:5608cd5e5a2506eadf4d1543563a4c47c6b32d259355a53cb847ba3d74853b3e

Observation 7c754f5f-5702-4004-a0aa-5fb1673780cf · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Accelerating Large Language Model Decoding with Speculative Sampling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:44.882400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:15:44.882400Z digest=sha256:59ffcfdb4fc6b0548a73c1bdcd8d4994c6ee11240fdebb7a4ba0452f5295de86

Observation bdd14e65-f8a2-4532-b295-7d5ed4d5594e · outbound

This paper cites Evaluating Large Language Models Trained on Code.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Evaluating Large Language Models Trained on Code

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:44.888884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:15:44.888884Z digest=sha256:751a5cea87bf4a1ea4643e19bbf2a6c4bc78612492fd12477e511e86cc30dccb

Observation 96447f0a-0d6b-47a8-b1b3-81e18e68e95e · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:44.892665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:15:44.892665Z digest=sha256:371946d4df224b56f42caf0149f821c6fe901df99c6f2ab3a5416353ff30a3fa

Observation bf6635c0-7cf7-4eaf-a680-5b53ed6b59d1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:44.895818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:15:44.895818Z digest=sha256:8742f6398314f9b95f50f0d5dfeacec3d59dbe2d0dc7d50a07b39445a612d5c8

Observation 6a82ae83-f87f-4e58-95e9-f708bc578699 · outbound

This paper cites DAPPLE: A Pipelined Data Parallel Approach for Training Large Models.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding DAPPLE: A Pipelined Data Parallel Approach for Training Large Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:15:45.367229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T04:15:44.900179Z digest=sha256:24adb3bd281f1434588239b51fa4d43dec814ad8ec848e452db96006622f139e

Observation 49df5af8-e974-4540-b019-993b34b03e69 · outbound

This paper cites LongCoder: A Long-Range Pre-trained Language Model for Code Completion.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding LongCoder: A Long-Range Pre-trained Language Model for Code Completion

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:44.902945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:15:44.902945Z digest=sha256:31db304df99c0e0ceb48d42fa68a2d36a7caaf340420ef26f85ecf0a0724860f

Observation 5bb237d4-b7dc-4d72-8cdf-2f00cd1d398e · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:44.905754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:15:44.905754Z digest=sha256:a8058727b4b0a6f8491b977c86affde3a712b0054e2aa23506ac83a990f94983

Observation 0ef14264-c6a6-49ae-9bcb-ad5d004cddff · outbound

This paper cites Unlocking the Potential of ChatGPT: A Comprehensive Exploration of its Applications, Advantages, Limitations, and Future Directions in Natural Language Processing.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Unlocking the Potential of ChatGPT: A Comprehensive Exploration of its Applications, Advantages, Limitations, and Future Directions in Natural Language Processing

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:15:45.338849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T04:15:44.908489Z digest=sha256:0bb6624f5e87232b765dad2191ed3578989fde9e9503300e97829194eb211bf3

Observation ab8f421f-7074-471e-a655-4a27f89380e3 · outbound

This paper cites Le, Yonghui Wu, and Zhifeng Chen.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Le, Yonghui Wu, and Zhifeng Chen

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:15:45.546350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T04:15:44.911012Z digest=sha256:d7381ec0e11c745f47621152eaf05a3163bb4bd26f1cace3352babaa0cd5fc60

Observation 87637fcd-8ea2-4e70-8d37-31eb934ce7ee · outbound

This paper cites Language Models for Code Completion: A Practical Evaluation.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Language Models for Code Completion: A Practical Evaluation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:44.916662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:15:44.916662Z digest=sha256:2ef1b3d18ca25097883cd429e40b5eab6efb92e071c4348e69ce66d237706837

Observation 46f4ee4b-5f47-4d96-9197-7ef6c6fd25af · outbound

This paper cites Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:15:45.537938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T04:15:44.919515Z digest=sha256:5e383686e77ccfbd8312b8e2b2e9ccdcf67d66dc65ce694403ab63864ca4e4ae

Observation aceba6d5-f5de-4472-912c-3836a932dde2 · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:44.922026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:15:44.922026Z digest=sha256:14301985949998a9e8685b8b3d3e1e9069bab3634e809d1b205f99bbfd29ac0a

Observation 87745101-6bcf-4d75-8363-4a2416d0daea · outbound

This paper cites Fast inference from transformers via speculative de- coding, 2023.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Fast inference from transformers via speculative de- coding, 2023

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:15:45.529451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T04:15:44.924727Z digest=sha256:95caf2f68b621df923a1207534869975e7eccd12c3724bf28b173bfbbd23239f

Observation 3e99ce16-7dd7-4564-bf61-74a784243bc8 · outbound

This paper cites EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:44.927085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:15:44.927085Z digest=sha256:ebf3799c572359fa96d5e931afb9d978345a8413f8fe6945d973f7136abad2b4

Observation 2f1a7488-ce17-4678-a6ba-c130cc8374ba · outbound

This paper cites AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:44.929623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:15:44.929623Z digest=sha256:595f7a820f865686c0d9543549cd0b3bd77085e94ceac5916612d78d120bbfb1

Observation 6c5ef95b-fca0-489f-ba9f-d01467c1509d · outbound

This paper cites PEARL: Parallel Speculative Decoding with Adaptive Draft Length.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding PEARL: Parallel Speculative Decoding with Adaptive Draft Length

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:44.932462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:15:44.932462Z digest=sha256:866fde85315a7dca67a378ded9127b4e8869b817f5d8480ba797bf31b722b070

Observation ff29de01-5fa9-470f-b85a-f1ce27a5a5ab · outbound

This paper cites TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:44.934977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:15:44.934977Z digest=sha256:4c5eed1660fdc835cee187f7305d4e8ffc9e788fed1a0dedeac9a630ec71868c

Observation d04da29a-e91d-4121-9b85-9c9c9a1bacee · outbound

This paper cites CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:44.937596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:15:44.937596Z digest=sha256:50d551e0e07c2cd4bb610ceb797ec904af22b62db9a501fd5712f279825e6e09

Observation c4e72ecd-1580-42b7-a2b4-9e8959a0f456 · outbound

This paper cites AMUSD: Asynchronous Multi-Device Speculative Decoding for LLM Acceleration.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding AMUSD: Asynchronous Multi-Device Speculative Decoding for LLM Acceleration

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:44.940584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:15:44.940584Z digest=sha256:a4a4748373434bb6aa0e7c0a87c4d1aec93b5dc57af073e822ff5082887e0411

Observation 5a796db2-febf-4c3c-9bad-4b3df007c4fe · outbound

This paper cites Specinfer: Accelerating large language model serving with tree-based speculative infer- ence and verification.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Specinfer: Accelerating large language model serving with tree-based speculative infer- ence and verification

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:44.943253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:15:44.943253Z digest=sha256:e29cda63bf5faeeaaa879ef136a928d286a6dce343cc2dbefb6a8fb41cdecdf5

Observation 27849d61-0697-409b-9c04-2c4260d8ae7d · outbound

This paper cites Abstractive Text Summarization Using Sequence-to-Sequence RNNs and Beyond.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Abstractive Text Summarization Using Sequence-to-Sequence RNNs and Beyond

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:44.945536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:15:44.945536Z digest=sha256:ec5440d2c9be2bdeb443e97fe01f0adeb0bade07ed851553b00891ecbc19019b

Observation d0ab6bb1-7e66-4b45-a031-01bd48dfa83d · outbound

This paper cites Nvidia/tensorrt-llm: A tensorrt toolbox for optimized large language model inference.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Nvidia/tensorrt-llm: A tensorrt toolbox for optimized large language model inference

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:15:45.520123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T04:15:44.948239Z digest=sha256:a4658b0d0445ad1768af9d2d40558d353fc50d4cb0649223033f08d4a968214c

Observation 2e8f0325-e1b0-49e7-9a34-bdf2c8fd3340 · outbound

This paper cites GPT-4 Technical Report.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding GPT-4 Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:44.950907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:15:44.950907Z digest=sha256:3812a36c70b6bb43e365a1ab733d54a481c370d5729ef8c4ac84ad02f7ad896f

Observation 413a2d54-8dab-4c49-b80e-d2ca69d2e918 · outbound

This paper cites Splitwise: Efficient generative LLM inference using phase splitting.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Splitwise: Efficient generative LLM inference using phase splitting

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:44.953411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:15:44.953411Z digest=sha256:1f9799defc35a2a8f45d5a5e85dcb878d87f37e53c09c359a104fa105fdcc16a

Observation 45eb26d6-6f58-4d29-a6dd-14bc43a42ea6 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:44.956041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:15:44.956041Z digest=sha256:f0809d34e78437f6f849f23eb75d503fab075103875accfb42896c153ab1be51

Observation eaa4fb03-50e6-4477-8c55-81b5318b63b5 · outbound

This paper cites Tatsu-lab/stanford-alpaca: Code and documentation to train stanford’s alpaca mod- els, and generate the data.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Tatsu-lab/stanford-alpaca: Code and documentation to train stanford’s alpaca mod- els, and generate the data

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:15:45.510400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T04:15:44.958715Z digest=sha256:d09a377e5d8b676710a227b1510361db59fb5766e2f5f694af113f6c5c204e9a

Observation b772985c-390f-4f4f-8d2f-06ffb9e55c33 · outbound

This paper cites CUTLASS, January.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding CUTLASS, January

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:15:45.500994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T04:15:44.961662Z digest=sha256:90ba9eaedadabadddaf8a709d0398a8f7b2c04e11ab54730c3cea0317336147a

Observation a5f09fc7-f62e-4b0a-b8fb-43ac80d9e983 · outbound

This paper cites Llama: Open and efficient foundation language models, 2023.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Llama: Open and efficient foundation language models, 2023

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:15:45.484956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T04:15:44.966734Z digest=sha256:b8da49a6b236e3f0fe9abd8ef7da7943f79877b777e1dbdadaf1ec8f90457c7f

Observation b4221c45-d13b-498c-9c04-b597e23fd764 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:44.969070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:15:44.969070Z digest=sha256:54673e7874a84e6065e1266c0af0db7f93cf4525ccf6688d40a0a81989bcb38c

Observation 79725ce1-bde8-4b1a-b38b-9657020b78eb · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:44.971675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:15:44.971675Z digest=sha256:d4d37efcc59d71750728d933bcfe8039d430dd0013e9ad04519845e1546007fa

Observation fcd83ab8-57ba-4e93-983b-2eeb8c57a941 · outbound

This paper cites SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:44.974746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:15:44.974746Z digest=sha256:7ada6355b86cf874ca35cda6b31fd416856b8fdd1ebcbf5862924d78d28dee2a

Observation 9686e40c-03cc-4790-8961-f2f059f63c0b · outbound

This paper cites When Search Engine Services meet Large Language Models: Visions and Challenges.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding When Search Engine Services meet Large Language Models: Visions and Challenges

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:44.977729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:15:44.977729Z digest=sha256:299b39da523c008f7e1885f0944e39c565d8140057c5adcece8d63bedcb581fd

Observation fe5bd438-7080-4835-87e7-d1c32d86c99e · outbound

This paper cites Qwen2 Technical Report.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Qwen2 Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:44.980434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:15:44.980434Z digest=sha256:34b57f0e0c4e16f593c866973287b43eac07b2dbb90710c4a951844288b9f0b6

Observation e37bbea3-b624-4247-badc-0985b0956d72 · outbound

This paper cites PALR: Personalization Aware LLMs for Recommendation.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding PALR: Personalization Aware LLMs for Recommendation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:44.983354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:15:44.983354Z digest=sha256:73d8d0cb7e08bac9e3d4de713657bd37ef5fae846d8757ef6fadf89d30f54e6f

Observation 2b6c1c76-ca47-4f89-a99e-90a37ce29d59 · outbound

This paper cites Orca: A distributed serving system for Transformer-Based generative models.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Orca: A distributed serving system for Transformer-Based generative models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:15:45.471867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T04:15:44.986524Z digest=sha256:62608f925a26b17ad6b6f46083f3b13f1123901bf6e37ce1e03d64d811be2f54

Observation e17bc3b8-1bfe-46e7-9b63-be6229c899cf · outbound

This paper cites Fltrnn: Faithful long-horizon taskplanningforroboticswithlargelanguagemodels.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Fltrnn: Faithful long-horizon taskplanningforroboticswithlargelanguagemodels

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:44.991410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:15:44.991410Z digest=sha256:1d0ac2431596bf8b8ce196cd98cfecf1755c72f357001a01562ff27f22f66508

Observation d5390ded-a5b2-4eb2-9f0b-14ca9aa49a33 · outbound

This paper cites Prepacking: A simple method for fast prefilling and increased throughput in large language models, 2024.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Prepacking: A simple method for fast prefilling and increased throughput in large language models, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:15:45.454968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T04:15:44.993835Z digest=sha256:9b5417ec817fdc251c3f3fd66918b1024465901503acbeab12fc991ddcd12a69

Observation 64966e2b-fdec-46df-9dff-219b634fbae3 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:44.996326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:15:44.996326Z digest=sha256:a9c016c6081cd240c145212cf2919fd43d5d5078f415d32a0d7fbc8432c446e0

Observation 5448a914-ed7d-4f69-9198-e5bea3bd5bf5 · outbound

This paper cites Gonzalez, Clark Barrett, and Ying Sheng.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Gonzalez, Clark Barrett, and Ying Sheng

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:15:45.446325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T04:15:44.999082Z digest=sha256:66bd557cab0bf9863c8d60e692501e49f49c20600b071d21abd20a808a63cf8d

Observation 45bfc0d8-6903-4663-91a2-b67ab244d231 · outbound

This paper cites DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:15:45.438465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T04:15:45.007350Z digest=sha256:410f5968bc1d50a8a106dfc2da77b3129906c51cfb998ffede3723eaaac28096

Observation 78327a17-ce95-4d19-8262-07a9811bd1bb · outbound

This paper cites Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:45.012578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:15:45.012578Z digest=sha256:43086f73a76fb4f46529dacaf05793bee839cc4ec8daba69feec6750e3d00630

Observation 44f94357-1a95-4a21-8b21-fc9328f7be3d · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding SGLang: Efficient Execution of Structured Language Model Programs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:45.001651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:15:45.001651Z digest=sha256:14f9d9262d529ad523e525a97df6543e6d134035607c3f70be8edd1f2adfe2f6

Observation 8e7eda31-2e22-4c88-94dc-78f0a44fab71 · outbound

This paper cites ISBN 978-1-939133- 40-3.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding ISBN 978-1-939133- 40-3

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:15:45.430088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T04:15:45.010054Z digest=sha256:3c62188e2228dd4e13e852013bcc567488517635c92d701c47c18a57c61e1228

Observation 8b24c817-30aa-4d89-83b7-fb108d0329d0 · outbound

This paper cites GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:44.913577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:15:44.913577Z digest=sha256:5e6ef9a7a91f034fe6eb88e5fbfd09461ff1a839c0acb1fa5da401e8525104bd

Observation dcf9a324-ef29-4afa-a642-0e47de74a6b9 · outbound

This paper cites ISBN 978-1-939133- 28-1.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding ISBN 978-1-939133- 28-1

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:15:45.463041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T04:15:44.989023Z digest=sha256:c60e03c7ebedb854c1ca88d830cbc0e01118a06c4c0f0bf32156212f423821b6

Observation 82fe9fe4-a7e4-402b-8803-a3c012037936 · outbound

This paper cites an unresolved cited work.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding Unresolved cited work

Reference 2023

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:15:45.492822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T04:15:44.964200Z digest=sha256:5e350b6da728bf57cd779431c81d809eead3d22a6810c36549b61681ba27d7ae

Observation 1b9cf4b0-5cc1-448d-b5d8-fdd725b6a0b1 · outbound

This paper cites ISBN 978-1-939133- 40-3.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding ISBN 978-1-939133- 40-3

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:15:45.554474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T04:15:44.868963Z digest=sha256:ad6695d3892bfe3cd0c38711ae737d7e8412cfc8042b2c4baf315699cca02c35

Pith citing papers

Observation 73e76fe6-46b5-44f6-a9f2-eec8c7d42f7f · inbound

ODIA: Oriented Distillation for Inline Acceleration of LLM-based Function Calling cites this paper.

ODIA: Oriented Distillation for Inline Acceleration of LLM-based Function Calling SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:45:05.693179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T18:45:04.545787Z digest=sha256:3b09eccce24e8ca4fa30a6ad204afb282b4be1ff587c160bd574b282c55b4eb2

Observation 44023fab-b54f-4e68-933e-80447626be8f · inbound

Unlocking Parallelism in Autoregressive Language Models via Speculative Decoding with Progressive Tree Drafting cites this paper.

Unlocking Parallelism in Autoregressive Language Models via Speculative Decoding with Progressive Tree Drafting SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T10:09:31.872801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:09:31.872801Z digest=sha256:3726e5ca94f17c6e31da3136ff5021cbdd5f95bea7e355fa18cb245a1ca42010