Pith. sign in

Paper Citation Record · LEDGER

DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 46 inbound Pith citation observations for arXiv:2401.08671.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.08671 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 46 of 46 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:27:17.268072Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T11:59:50.582581Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9ea946e8-56ae-4cfe-9859-67f8ea75d258 · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 279

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:39:33.490879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:7957686068f2bce0e13dd82e279bccd3f0bdb7ffda54e0457c4e450af1529791

Observation 852d3521-7c90-462c-832a-bcc96f884e8b · inbound

HybridFlow: A Flexible and Efficient RLHF Framework cites this paper.

HybridFlow: A Flexible and Efficient RLHF Framework DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:53:38.983721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T07:53:38.715353Z digest=sha256:ecd9205e8dc3b959fe009f04caa0ea1531ff902adf170d6fbe021be999afd54f

Observation 816fb079-be1e-4d89-b5fb-5e5ad5510036 · inbound

Ensuring Fair LLM Serving Amid Diverse Applications cites this paper.

Ensuring Fair LLM Serving Amid Diverse Applications DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T13:44:18.524777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:44:18.524777Z digest=sha256:647ea8d1b4b25117cec03232cf672f2a95cba9c4d90492dff81c126b99c086ff

Observation 91aa30dc-2c98-4911-8bd1-504ade85231f · inbound

BlendServe: Optimizing Offline Inference for Auto-regressive Large Models with Resource-aware Batching cites this paper.

BlendServe: Optimizing Offline Inference for Auto-regressive Large Models with Resource-aware Batching DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T13:36:43.930205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:36:43.930205Z digest=sha256:cae0713881e83fad0370ba95ae041f618e5ddf711cd2b1e76fd2fbbfcb17e14a

Observation a129fb99-037a-4081-a384-96d4579e7a79 · inbound

BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching cites this paper.

BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-23T16:58:11.937207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T16:57:46.645061Z digest=sha256:f42bd5ba2a2b8b87c8080bd941ba9b3a9a18ba3a6c619d47a14915d29158f50d

Observation 0c8721d8-0d08-4ff2-ad35-51587c21c02f · inbound

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design cites this paper.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:57:40.262382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:75eb04107ca4e6d01d50e8b0ecf9f7da0cae75fd9b49376089379042a703e1d6

Observation 5c526a61-35fc-4073-9207-c18d2d746ee0 · inbound

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels cites this paper.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.498348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.498348Z digest=sha256:032d303eed5d499c637d5bcf9c1d6cc32d3d1284698c60a8f65d35dd36f94ffc

Observation ba259c45-6fa6-4ff5-9025-48c8fcdfd722 · inbound

Efficiently Scaling LLM Reasoning with Certaindex cites this paper.

Efficiently Scaling LLM Reasoning with Certaindex DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T23:09:44.472681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:09:44.472681Z digest=sha256:bd1ca9fd6b8994acebf2af1433c294cf1344f6c0379e793b42dadbd57450250f

Observation 4346f9de-5e52-4366-852c-485d078ea07e · inbound

AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding cites this paper.

AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T17:39:10.292188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:39:10.292188Z digest=sha256:8b5f6e8888a3bba7b1c722dfd05e974e3e158905fbb3539fc4d7664df4a1fb19

Observation 9565420e-dca0-4ebd-8b5b-4b53a0e33883 · inbound

iServe: An Intent-based Serving System for LLMs cites this paper.

iServe: An Intent-based Serving System for LLMs DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T21:37:14.252651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:37:14.252651Z digest=sha256:0d6b78c0d32cde25bcf3c30e0c375a9d547258ad6f24e09aa4a3d80b53c8f78a

Observation e0a45e8e-b80a-420d-b3ee-031f62db8e6d · inbound

KVDirect: Distributed Disaggregated LLM Inference cites this paper.

KVDirect: Distributed Disaggregated LLM Inference DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:14.577889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:14.577889Z digest=sha256:631f7eef5e51e0f454877bf06077d7c7af20852319cb430105d76f993792f11b

Observation cca8c0a0-3680-428a-9ec0-d540affbe1ac · inbound

Klotski: Efficient Mixture-of-Expert Inference via Expert-Aware Multi-Batch Pipeline cites this paper.

Klotski: Efficient Mixture-of-Expert Inference via Expert-Aware Multi-Batch Pipeline DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T17:55:54.042509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:55:54.042509Z digest=sha256:471564ab05842c8f79ca1c7d0e4202bc6e747b9dce5d1a1052171d8db508b1c7

Observation 6bedb766-c3d8-4009-a6f4-8cabf5d05c7c · inbound

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees cites this paper.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.872509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.872509Z digest=sha256:a69fcc92a7b5640e6bc54c8dcbbc2afe7a98dc55935500215a189458f84a2ce3

Observation ebbd282c-edd5-43e1-ae96-9cf2bf0587d4 · inbound

MIST: A Co-Design Framework for Heterogeneous, Multi-Stage LLM Inference cites this paper.

MIST: A Co-Design Framework for Heterogeneous, Multi-Stage LLM Inference DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:17:08.081299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T21:16:31.655330Z digest=sha256:81c15c172702c9fd390b4e0a610c15edceca2eec3499472b4048cd7e4cb8d88b

Observation d2c85ef4-c9e7-4393-b965-2be97d35296d · inbound

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference cites this paper.

SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:17.268072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:17.268072Z digest=sha256:b01d3aaf2699db97721fe83bafa8a87ddf9211e706315345c4434f6fdc398f8a

Observation 096efbf6-39e3-4d39-a19d-496313f4af06 · inbound

Taming the Titans: A Survey of Efficient LLM Inference Serving cites this paper.

Taming the Titans: A Survey of Efficient LLM Inference Serving DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T05:49:09.813521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:49:09.813521Z digest=sha256:b741ee4be0bdbec441478881aab3bbd2ff253f3e4a5877a5ab7f9ab251d1a1d7

Observation 2f0a985d-a116-40b9-8d50-f3648ef4cc2f · inbound

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage cites this paper.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.325191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.325191Z digest=sha256:15f47c0e107988fe9a93e254de90e2440710200e723c2fba1cd7678edd65d18f

Observation 72c79d06-ef02-4f26-9420-a0d9059c4842 · inbound

Ascendra: Dynamic Request Prioritization for Efficient LLM Serving cites this paper.

Ascendra: Dynamic Request Prioritization for Efficient LLM Serving DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T05:23:26.935869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:23:26.935869Z digest=sha256:37291eb21790037a8bceaa4b57a640667907b508152e22b0eaba5dc54055d150

Observation b06f25c7-381a-4421-a4e4-1855b33a2e6f · inbound

ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production cites this paper.

ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-22T15:44:58.084998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T15:42:05.266854Z digest=sha256:b7f3936d5e133ab60ca28925e1be28dd6c56770fcb9b17cb390de48c59b9f591

Observation 313443c8-0e08-4e59-b929-60f45ebddbc7 · inbound

AnchorAttention: Difference-Aware Sparse Attention with Stripe Granularity cites this paper.

AnchorAttention: Difference-Aware Sparse Attention with Stripe Granularity DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:52.961383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:52.961383Z digest=sha256:59e26f537622bec0ffa81f32fa2d6cbeb71c83bc03fcff56e0d7281088af469c

Observation 17c4b2d3-7d80-4f7e-a321-478c29739989 · inbound

Rectified Sparse Attention cites this paper.

Rectified Sparse Attention DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.938385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.938385Z digest=sha256:df160c1f77f64f9526365c703707b32777469633a7fa33a5e6e1e777d64fa7df

Observation 3f75106c-2cda-4921-a977-c6699722549e · inbound

Chat-Ghosting: A Comparative Study of Methods for Auto-Completion in Dialog Systems cites this paper.

Chat-Ghosting: A Comparative Study of Methods for Auto-Completion in Dialog Systems DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:19:10.069938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:19:10.069938Z digest=sha256:9392c48d5c992040f1f1a31315bd1b6bffb5d55ce46fe53a8bb558b5f4c16105

Observation f846ec65-36d7-4250-8207-3cbfadd1c795 · inbound

Rethinking Caching for LLM Serving Systems: Beyond Traditional Heuristics cites this paper.

Rethinking Caching for LLM Serving Systems: Beyond Traditional Heuristics DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T16:21:09.103693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:21:09.103693Z digest=sha256:4ad3961f1211ba7547801d275e81d945a5387f73b2b30bd5894aaac30751f7de

Observation 804a2c0c-1dc3-47dc-a9f9-82497e7d3e21 · inbound

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference cites this paper.

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T15:53:29.009532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:53:29.009532Z digest=sha256:4fa9cbb7022e650b4231a449798f21dc7b2399cfda1a46004116bcba51ce2fce

Observation c81406e8-a4fc-4464-adde-fb4e1b76914f · inbound

WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving cites this paper.

WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-22T12:24:51.224216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T12:22:56.972437Z digest=sha256:585f4dc714de74815b4d2be19a539f6dcef76edc7caa4b17ff747e2e44a21a34

Observation 268c5e27-d697-4056-a7a9-becc9c17b726 · inbound

PrefixWall: Mitigating Prefix Caching Side Channels in Shared LLM Systems cites this paper.

PrefixWall: Mitigating Prefix Caching Side Channels in Shared LLM Systems DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-21T12:15:06.799217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T12:14:09.509302Z digest=sha256:1cab771dc4395d7faf7453177c9e23e048e5302eef9a19f8a270a8116fec37b8

Observation ffa085a0-9695-41a7-8065-f41d1ffbd4c5 · inbound

FASTER: Rethinking Real-Time Flow VLAs cites this paper.

FASTER: Rethinking Real-Time Flow VLAs DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T08:05:15.319681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T08:02:13.188363Z digest=sha256:e79a0210cfb9009a70fc302d88bc97575670c24d511f7cf8fe81d39a7ea5c11a

Observation aaf05164-07a0-4f3b-a171-ee9197516154 · inbound

FASTER: Rethinking Real-Time Flow VLAs cites this paper.

FASTER: Rethinking Real-Time Flow VLAs DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T10:50:01.062812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T10:48:56.280105Z digest=sha256:d77ee1e1647375c667241845ce24720e587a2c27c60e864d8fff1727020507e1

Observation 9c00d7a0-282f-44cf-acd1-f1192872aa99 · inbound

Towards Efficient Large Vision-Language Models: A Comprehensive Survey on Inference Strategies cites this paper.

Towards Efficient Large Vision-Language Models: A Comprehensive Survey on Inference Strategies DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:43:00.670264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T21:40:55.343169Z digest=sha256:ceea0a3eadf709847a7ccf4c8dcca48dc0eb2894f913642db7cba28b497ea173

Observation 876815a3-5aa7-486c-8378-6ce3c7e3a195 · inbound

Characterizing Performance-Energy Trade-offs of Large Language Models in Multi-Request Workflows cites this paper.

Characterizing Performance-Energy Trade-offs of Large Language Models in Multi-Request Workflows DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:30:00.274373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T12:28:38.809950Z digest=sha256:cbbb436d894645c180a0f6b1c9c48bbc50ab3a6d91dd5283be0f9b711fdafd47

Observation d57681e3-f3fb-42fc-8982-ad26d7846e04 · inbound

Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines cites this paper.

Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:59:03.464710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T09:52:40.057014Z digest=sha256:1a91fcb70165ff5f6cbbaa2f61a6d5361eb4e4501c89cbe900c327fe7ec60208

Observation 61b831f4-fcd0-47ff-adc3-18b849d53a70 · inbound

Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities cites this paper.

Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:16:10.210742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T09:45:57.201837Z digest=sha256:30fcf3a331682e4fcf8ed9cc2c53053ffcf9fa60d002b65bd365b24aad097e16

Observation c0ab0b09-ffc3-4798-9256-856c3064c131 · inbound

MARS: Efficient, Adaptive Co-Scheduling for Heterogeneous Agentic Systems cites this paper.

MARS: Efficient, Adaptive Co-Scheduling for Heterogeneous Agentic Systems DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:35:35.686543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T14:30:56.899306Z digest=sha256:ec5ee5daf38711a43ff4558157a24dc002f19f53c5518b616972effd8cbcee85

Observation bdb1d749-6c65-4651-bbda-4dec91c8d58d · inbound

Federation of Experts: Communication Efficient Distributed Inference for Large Language Models cites this paper.

Federation of Experts: Communication Efficient Distributed Inference for Large Language Models DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:51:07.551006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T13:40:50.411198Z digest=sha256:0bdf22676580b3cc9d555c2b7da463300bf08fad1c0e51a0d7a102cb0b437410

Observation 4baef794-3b5c-4841-b783-3006804077fb · inbound

Focused Forcing: Content-Aware Per-Frame KV Selection for Efficient Autoregressive Video Diffusion cites this paper.

Focused Forcing: Content-Aware Per-Frame KV Selection for Efficient Autoregressive Video Diffusion DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:23:14.150643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T11:19:21.721494Z digest=sha256:eb9c7b4b6d00d893f293894c65e026165568f926b410d34d6940dd978d9c050b

Observation 13fd3d8f-a36d-4899-832c-359a99beb400 · inbound

AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System cites this paper.

AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:15:17.294037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-25T03:12:49.028342Z digest=sha256:6de7136a45f5aa8ef96e297e837e76f89fd3da308bad51694a5a68168fa471dc

Observation 83d9298f-7b54-40a6-9b32-f483749567fe · inbound

Observation, Not Prediction: Conversation-Level Disaggregated Scheduling for Agentic Serving cites this paper.

Observation, Not Prediction: Conversation-Level Disaggregated Scheduling for Agentic Serving DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:56:25.651067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T12:56:16.768455Z digest=sha256:09905ec89088e8ae491fcc93c385e5d0e98b577da9f05f5889a91737171b6128

Observation 6e215635-2261-408f-9102-c420675d9305 · inbound

Multi-Segment Attention: Enabling Efficient KV-Cache Management for Faster Large Language Model Serving cites this paper.

Multi-Segment Attention: Enabling Efficient KV-Cache Management for Faster Large Language Model Serving DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:36:26.302619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T11:38:05.435505Z digest=sha256:dc44007528bcf60f859148966c3886acd028583fd0b29d44117ca2cf5fd3be81

Observation f690c26d-f209-43f4-b662-660346d06641 · inbound

Beyond Greedy Chunking: SLO-Aware Sliding-Window Scheduling for LLM Inference cites this paper.

Beyond Greedy Chunking: SLO-Aware Sliding-Window Scheduling for LLM Inference DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:27:06.039657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T23:49:28.318260Z digest=sha256:b6596d146a9cedfb242fb3b85b0fad1bcbb90f0fe33bdfe97ebb6658796adb39

Observation e407d765-5959-43a4-b272-097e3e6cc9fa · inbound

Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents cites this paper.

Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:36:59.463922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T01:07:14.691347Z digest=sha256:a61153c1cce8d39a46b8eef973caafc93d21b9b25e5a7ab264885e54f2ef3144

Observation 338310bb-7333-4cb7-9032-c6507b74679f · inbound

ReMP: Low-Downtime Runtime Model-Parallelism Reconfiguration for LLM Serving cites this paper.

ReMP: Low-Downtime Runtime Model-Parallelism Reconfiguration for LLM Serving DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:39:24.932913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T19:32:31.882977Z digest=sha256:cf593037d1cd9f35ecbf5f510890de4a9e538753729d06ef7f4a628d8966cae1

Observation a6e62147-93fa-408e-88c6-1da34df551c8 · inbound

LiveServe: Interaction-Aware Serving for Real-Time Omni-Modal LLMs cites this paper.

LiveServe: Interaction-Aware Serving for Real-Time Omni-Modal LLMs DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:59:50.585382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T07:26:07.356352Z digest=sha256:c839a38b9ee0ee35a43b3feae5ee96e178079233469e6bdf744eaceb6f429745

Observation 286d72a8-abb5-474e-8411-d1aa48d0e656 · inbound

SmoothAgent: Efficient Long-Horizon LLM-Based Agent Serving with Lookahead Context Engineering cites this paper.

SmoothAgent: Efficient Long-Horizon LLM-Based Agent Serving with Lookahead Context Engineering DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:14.671196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-02T17:19:53.560039Z digest=sha256:dfbea00296acf95fbe5438aab1e9c183b06ca5bae1a98ab8068c036e39698f61

Observation 208c00e4-9b51-4ca6-a064-f638bf559b5f · inbound

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving cites this paper.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:03be1e6abbb029274b7c439f8394afea801c93d21449e774c2ca59a073ed33fc

Observation b03ec011-2f9d-4d74-9362-1893d5dec339 · inbound

BlockServe: Block-Grained Continuous Batching for High-Throughput Diffusion LLM Serving cites this paper.

BlockServe: Block-Grained Continuous Batching for High-Throughput Diffusion LLM Serving DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T05:43:12.690359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:43:12.690359Z digest=sha256:9cef4ee2fd5889c2912c515d331187e6eec91b2ca3f0b2695d30a89ed181bcf4

Observation a0422bac-16ce-46b9-823a-0e3a1899cb82 · inbound

Beyond Storage: State as a Runtime Control Problem in Parallel and Distributed Systems cites this paper.

Beyond Storage: State as a Runtime Control Problem in Parallel and Distributed Systems DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-01T19:51:06.098713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:51:06.098713Z digest=sha256:6be528399a8ff7397c40583c3d3df022638026a50fbaa9055725b0a25a3eb24e