Pith. sign in

Paper Citation Record · LEDGER

FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2403.11421.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.11421 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T20:31:35.676302Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T15:30:17.948710Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1aabc17c-9418-4901-9a3f-caef676b27bc · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 252

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:39:33.397582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:836b64a725a189caae82dfa669984405015582e3862a2f4c81b10dc4dac86439

Observation 610edb82-c357-4d9a-9513-9c86109dd002 · inbound

EcoServe: Designing Carbon-Aware AI Inference Systems cites this paper.

EcoServe: Designing Carbon-Aware AI Inference Systems FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T20:31:35.676302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:31:35.676302Z digest=sha256:bdab6158738dae2196d2f74b4a217e7bfa2d61fff22f65e1767e6b80435bab2c

Observation f576932a-1cf8-4630-82db-d0af173eecb0 · inbound

Hardware-Efficient Attention for Fast Decoding cites this paper.

Hardware-Efficient Attention for Fast Decoding FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:30.932279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:30.932279Z digest=sha256:9f2c3b9bc9ed285295bf75d5733a28be0a8080ac077f35106edf03691bb54203

Observation 25305e44-e40b-428b-8584-0e4fcba07bfe · inbound

Kinetics: Rethinking Test-Time Scaling Laws cites this paper.

Kinetics: Rethinking Test-Time Scaling Laws FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:30:34.224262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:30:34.224262Z digest=sha256:936103d7ae760408ff478154bc44eca194e48d512979cf34bde3bb031f32f7d8

Observation 1db91454-2a3d-4e31-908e-ff723db9d091 · inbound

Beyond the Buzz: A Pragmatic Take on Inference Disaggregation cites this paper.

Beyond the Buzz: A Pragmatic Take on Inference Disaggregation FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:26:32.741822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:26:32.741822Z digest=sha256:697c183670da3b99bbcd93859c11eb4d8e427c854849ed4d2ca22bbe4824a065

Observation 1b096e79-904e-4e7a-9c35-dab5e3ecdf9f · inbound

MoQAE: Mixed-Precision Quantization for Long-Context LLM Inference via Mixture of Quantization-Aware Experts cites this paper.

MoQAE: Mixed-Precision Quantization for Long-Context LLM Inference via Mixture of Quantization-Aware Experts FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:52.754758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:37:52.754758Z digest=sha256:e7aebb661c2f9376ff5f4efbc0d28c9ec8b4e53372f7fa797c60a44fdfa1e3dc

Observation a9c97e9d-4fe5-4fba-ba4d-0a63006dc951 · inbound

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding cites this paper.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:37.472671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:37.472671Z digest=sha256:81b51bce89d9579f30d70740873ce5aa9b41ebc9b85a5c171138fa96ccbd13a0

Observation 5737006c-aef0-4827-b38a-d548350d40ee · inbound

Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead cites this paper.

Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 123

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:32.334492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:32.334492Z digest=sha256:dbbe7b22a7cec4fd1a000fcd75223a76cc351d4aa1caec7b8f46254e04a9f4ce

Observation 3104cf80-a1c0-4618-84aa-fcf5c0feb196 · inbound

KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding cites this paper.

KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T17:20:26.333389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:20:26.333389Z digest=sha256:8d102612750250ee52853d48f020f1ef58c47473249f446203aa04478f9e747e

Observation 950360d6-18a2-4927-9922-c48f3179ece1 · inbound

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs cites this paper.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:37.907873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:37.907873Z digest=sha256:e8bc290a096b7ec55743fe7226e18001428f4f0e98b41c6923b3ccee72f2c434

Observation 3e8225a2-deb0-4c9f-b07a-c421677e2b82 · inbound

Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism cites this paper.

Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T20:53:31.847951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:53:31.847951Z digest=sha256:19b2fa9c7a7af4e794ee3f8e71c4448e0de210da7a8be24389e47c21952f73e2

Observation 2f6b9d54-a29c-4d28-973a-c09b00bd5bb0 · inbound

SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips cites this paper.

SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-21T15:30:17.950957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T15:26:01.283448Z digest=sha256:633959ba2436a9fc3e97bd075f6b3e54fbbfd095eee0efb14a4652bbfb0bffe9

Observation 90749329-89f8-4a3d-ba2b-73467a013d77 · inbound

Understanding Rate-Distortion Performance in Distributed Transformer Inference cites this paper.

Understanding Rate-Distortion Performance in Distributed Transformer Inference FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T10:02:42.597700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T10:01:53.390831Z digest=sha256:ed44e1fb66d8f86cb760d5023dc407f45b39dcc4128e651a41f034a40edb0e15

Observation 8a7c8d74-18bd-448f-848f-31ac9a830ebc · inbound

Understanding Rate-Distortion Performance in Distributed Transformer Inference cites this paper.

Understanding Rate-Distortion Performance in Distributed Transformer Inference FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T06:51:13.983136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:51:13.983136Z digest=sha256:9eb5767656cd137cf70fc15d2337fb86b9d076548f431a30f399103fba9985cf

Observation 0a05612f-965f-40c6-acb2-8d82e7d567af · inbound

Understanding Rate-Distortion Performance in Distributed Transformer Inference cites this paper.

Understanding Rate-Distortion Performance in Distributed Transformer Inference FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T06:23:13.506645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:23:13.506645Z digest=sha256:77eb236ca56835325243a731cf187ebbe4b5d1d0fca4a355e6166f540462b73a

Observation 63dae347-e822-4f8a-80ad-d723fa99cf1c · inbound

DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch cites this paper.

DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T15:06:54.210093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:06:54.210093Z digest=sha256:08ea06379684725cd4757ae7801deb3312ca66ccfd6c983b5e2cd221cb257c6e

Observation e3cec5c8-1849-4e2e-89b7-01434f9dcc5b · inbound

Training-Free Hashing-Based Attention via Binary Principal Components cites this paper.

Training-Free Hashing-Based Attention via Binary Principal Components FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:48.467163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:45:48.467163Z digest=sha256:640f429573ceba09ee8d9572e43ebd3221877d1e4088d22ca8e9636c0838c1d9