Pith. sign in

Paper Citation Record · LEDGER

DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 35 inbound Pith citation observations for arXiv:2401.08671.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.08671 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:55:54.042509Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T11:59:50.582581Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9ea946e8-56ae-4cfe-9859-67f8ea75d258 · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 279

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:39:33.490879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:419eca2543b86ed346a3ca3cde919dee6e0e3e57af5e1b23d5f2f7efdaf4b98c

Observation 852d3521-7c90-462c-832a-bcc96f884e8b · inbound

HybridFlow: A Flexible and Efficient RLHF Framework cites this paper.

HybridFlow: A Flexible and Efficient RLHF Framework DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:53:38.983721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T07:53:38.715353Z digest=sha256:ad7256b7d7324a52ace2b55f1bd8fb5ad64a666bad8dcd669da7b3566a1f529a

Observation a129fb99-037a-4081-a384-96d4579e7a79 · inbound

BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching cites this paper.

BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-23T16:58:11.937207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T16:57:46.645061Z digest=sha256:43a5f23f6bee3e0c5138ebd6f80d70ebddd56efe053b163bade62510839301a7

Observation 0c8721d8-0d08-4ff2-ad35-51587c21c02f · inbound

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design cites this paper.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:57:40.262382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:f88548c96564fcbe2d222fdbccad0a4dbb2e3fcb02e7fb40f28b25798611225e

Observation cca8c0a0-3680-428a-9ec0-d540affbe1ac · inbound

Klotski: Efficient Mixture-of-Expert Inference via Expert-Aware Multi-Batch Pipeline cites this paper.

Klotski: Efficient Mixture-of-Expert Inference via Expert-Aware Multi-Batch Pipeline DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T17:55:54.042509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:55:54.042509Z digest=sha256:3a4da66639ae700236338353c37f22c305333f19b0d5a98e48c0b7ff4ea1078e

Observation 6bedb766-c3d8-4009-a6f4-8cabf5d05c7c · inbound

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees cites this paper.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.872509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.872509Z digest=sha256:1fae67ba280544561033f8dfb36d96d97305d74a5c088dc71ba588c72b2d2267

Observation ebbd282c-edd5-43e1-ae96-9cf2bf0587d4 · inbound

MIST: A Co-Design Framework for Heterogeneous, Multi-Stage LLM Inference cites this paper.

MIST: A Co-Design Framework for Heterogeneous, Multi-Stage LLM Inference DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:17:08.081299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T21:16:31.655330Z digest=sha256:68aa346e53cdb612f648ce534bc45bc0a64e5f11d7ec55a6f75ddb5b619c52df

Observation b06f25c7-381a-4421-a4e4-1855b33a2e6f · inbound

ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production cites this paper.

ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-22T15:44:58.084998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T15:42:05.266854Z digest=sha256:5168771d94cf96f378560472b1af5f6de0cc7b7d452962085a730686212b637f

Observation 313443c8-0e08-4e59-b929-60f45ebddbc7 · inbound

AnchorAttention: Difference-Aware Sparse Attention with Stripe Granularity cites this paper.

AnchorAttention: Difference-Aware Sparse Attention with Stripe Granularity DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:52.961383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:52.961383Z digest=sha256:44a65daa82f6b0a4a81c365a381d8c31309bb763bc514053d346ac17c2240829

Observation 17c4b2d3-7d80-4f7e-a321-478c29739989 · inbound

Rectified Sparse Attention cites this paper.

Rectified Sparse Attention DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.938385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.938385Z digest=sha256:b8bdafa6450187f61779b5480531061c2153c719ef283a72dc2e63f3c68d7da7

Observation 3f75106c-2cda-4921-a977-c6699722549e · inbound

Chat-Ghosting: A Comparative Study of Methods for Auto-Completion in Dialog Systems cites this paper.

Chat-Ghosting: A Comparative Study of Methods for Auto-Completion in Dialog Systems DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:19:10.069938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:19:10.069938Z digest=sha256:33b8a319e0d78184809bc4f66c641980dd2545c2594bc874f6cb25fccbcfd255

Observation f846ec65-36d7-4250-8207-3cbfadd1c795 · inbound

Rethinking Caching for LLM Serving Systems: Beyond Traditional Heuristics cites this paper.

Rethinking Caching for LLM Serving Systems: Beyond Traditional Heuristics DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T16:21:09.103693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:21:09.103693Z digest=sha256:5acaac1f8c7dfe7fab8c5743b9525ba8ac4b6294dabe45dca33963edc2eb3fab

Observation 804a2c0c-1dc3-47dc-a9f9-82497e7d3e21 · inbound

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference cites this paper.

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T15:53:29.009532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:53:29.009532Z digest=sha256:50bde247e7d272c7f0491eddd02c5d87a087997ba5aa6fb33327d448cd85c2b0

Observation c81406e8-a4fc-4464-adde-fb4e1b76914f · inbound

WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving cites this paper.

WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-22T12:24:51.224216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T12:22:56.972437Z digest=sha256:2822718348efaf2fedb2348c1519516902f60b0aed7ec92250ec2701e71f4601

Observation 268c5e27-d697-4056-a7a9-becc9c17b726 · inbound

PrefixWall: Mitigating Prefix Caching Side Channels in Shared LLM Systems cites this paper.

PrefixWall: Mitigating Prefix Caching Side Channels in Shared LLM Systems DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-21T12:15:06.799217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T12:14:09.509302Z digest=sha256:2d20c63bf09894ca3570cfaa1c41cee36600e4ef09b3f35d312180c67fc50221

Observation ffa085a0-9695-41a7-8065-f41d1ffbd4c5 · inbound

FASTER: Rethinking Real-Time Flow VLAs cites this paper.

FASTER: Rethinking Real-Time Flow VLAs DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T08:05:15.319681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T08:02:13.188363Z digest=sha256:98d34d65ad67b03bed3a9bceec8bc927795d23cb5ca934f1aa1dfb622a6b002a

Observation aaf05164-07a0-4f3b-a171-ee9197516154 · inbound

FASTER: Rethinking Real-Time Flow VLAs cites this paper.

FASTER: Rethinking Real-Time Flow VLAs DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T10:50:01.062812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T10:48:56.280105Z digest=sha256:458561c200502f0a3b15d4afa3d8f107245868e77212a86f23866e5cabc38be9

Observation 9c00d7a0-282f-44cf-acd1-f1192872aa99 · inbound

Towards Efficient Large Vision-Language Models: A Comprehensive Survey on Inference Strategies cites this paper.

Towards Efficient Large Vision-Language Models: A Comprehensive Survey on Inference Strategies DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:43:00.670264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T21:40:55.343169Z digest=sha256:8ffddb22524f7a6ad669af40edddba4db1b1e44440bd9d70aebb44fc57f33f50

Observation 876815a3-5aa7-486c-8378-6ce3c7e3a195 · inbound

Characterizing Performance-Energy Trade-offs of Large Language Models in Multi-Request Workflows cites this paper.

Characterizing Performance-Energy Trade-offs of Large Language Models in Multi-Request Workflows DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:30:00.274373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T12:28:38.809950Z digest=sha256:f1d989343956bd92005876d09f8c0e597546303d9901eacfab2ded7cf298899d

Observation d57681e3-f3fb-42fc-8982-ad26d7846e04 · inbound

Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines cites this paper.

Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:59:03.464710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T09:52:40.057014Z digest=sha256:2f9ca1574aaa46b15bc5b4c43ccc44be237b6aa3a1cc86def1ef6b82a8e1c011

Observation 61b831f4-fcd0-47ff-adc3-18b849d53a70 · inbound

Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities cites this paper.

Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:16:10.210742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T09:45:57.201837Z digest=sha256:2778e424a20cda23be72a9626b17bb5a3d4683e454793effcfd92e8bd5228700

Observation c0ab0b09-ffc3-4798-9256-856c3064c131 · inbound

MARS: Efficient, Adaptive Co-Scheduling for Heterogeneous Agentic Systems cites this paper.

MARS: Efficient, Adaptive Co-Scheduling for Heterogeneous Agentic Systems DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:35:35.686543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T14:30:56.899306Z digest=sha256:0acb65fb3dec40fa7a2f8a7b5e167764a094ad834e9993bc2689f2ce694f01c1

Observation bdb1d749-6c65-4651-bbda-4dec91c8d58d · inbound

Federation of Experts: Communication Efficient Distributed Inference for Large Language Models cites this paper.

Federation of Experts: Communication Efficient Distributed Inference for Large Language Models DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:51:07.551006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T13:40:50.411198Z digest=sha256:a73d760cf836891df7cc46298db0b0a03a8360edb4f162a6e643c9b167e56ac6

Observation 4baef794-3b5c-4841-b783-3006804077fb · inbound

Focused Forcing: Content-Aware Per-Frame KV Selection for Efficient Autoregressive Video Diffusion cites this paper.

Focused Forcing: Content-Aware Per-Frame KV Selection for Efficient Autoregressive Video Diffusion DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:23:14.150643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T11:19:21.721494Z digest=sha256:ce61f0f136bcbcbf069d127c931e767cbf62aa7da29403476cc787c443da76d5

Observation 13fd3d8f-a36d-4899-832c-359a99beb400 · inbound

AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System cites this paper.

AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:15:17.294037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T03:12:49.028342Z digest=sha256:b388e4b985930194457365cb5f70f991bb9db7a91da22d2bd086cb596bc3ab12

Observation 83d9298f-7b54-40a6-9b32-f483749567fe · inbound

Observation, Not Prediction: Conversation-Level Disaggregated Scheduling for Agentic Serving cites this paper.

Observation, Not Prediction: Conversation-Level Disaggregated Scheduling for Agentic Serving DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:56:25.651067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T12:56:16.768455Z digest=sha256:5d4fdb68b046be73f325246529a97e50ee4b51f99f7d30bcc7a0177277aa6c29

Observation 6e215635-2261-408f-9102-c420675d9305 · inbound

Multi-Segment Attention: Enabling Efficient KV-Cache Management for Faster Large Language Model Serving cites this paper.

Multi-Segment Attention: Enabling Efficient KV-Cache Management for Faster Large Language Model Serving DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:36:26.302619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T11:38:05.435505Z digest=sha256:c5518a93d90175c8cd90b9e754137a5ae59eb0bfb0dd54fca89f0478aa7c89ff

Observation f690c26d-f209-43f4-b662-660346d06641 · inbound

Beyond Greedy Chunking: SLO-Aware Sliding-Window Scheduling for LLM Inference cites this paper.

Beyond Greedy Chunking: SLO-Aware Sliding-Window Scheduling for LLM Inference DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:27:06.039657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T23:49:28.318260Z digest=sha256:b4a4950355248c27ac58963d5844063451e789a6c83e83fcbf8b35b3224dd84a

Observation e407d765-5959-43a4-b272-097e3e6cc9fa · inbound

Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents cites this paper.

Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:36:59.463922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T01:07:14.691347Z digest=sha256:3383998c25d186088c904d15402f8ce6b418394314cff686bb1de9c0c82e5231

Observation 338310bb-7333-4cb7-9032-c6507b74679f · inbound

ReMP: Low-Downtime Runtime Model-Parallelism Reconfiguration for LLM Serving cites this paper.

ReMP: Low-Downtime Runtime Model-Parallelism Reconfiguration for LLM Serving DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:39:24.932913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T19:32:31.882977Z digest=sha256:731d8a712de475ba6e3eb06d3a8517d583c5698a984aa559b751bc1d2a54daaa

Observation a6e62147-93fa-408e-88c6-1da34df551c8 · inbound

LiveServe: Interaction-Aware Serving for Real-Time Omni-Modal LLMs cites this paper.

LiveServe: Interaction-Aware Serving for Real-Time Omni-Modal LLMs DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:59:50.585382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T07:26:07.356352Z digest=sha256:fb63a8eccb4ef4a9f2e9f0ff0279d7e8b259e2ff7e3307c32efb4e13d8362b3b

Observation 286d72a8-abb5-474e-8411-d1aa48d0e656 · inbound

SmoothAgent: Efficient Long-Horizon LLM-Based Agent Serving with Lookahead Context Engineering cites this paper.

SmoothAgent: Efficient Long-Horizon LLM-Based Agent Serving with Lookahead Context Engineering DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:14.671196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-02T17:19:53.560039Z digest=sha256:6e558a5d7ad534a1dcdf98e279f5eaae5d3ff0c06f1fadb1e17562572ca83100

Observation 208c00e4-9b51-4ca6-a064-f638bf559b5f · inbound

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving cites this paper.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:ec65faa90827c3383285e0f0b22120d8e66869f9ee3810fac6abc8e01d169649

Observation b03ec011-2f9d-4d74-9362-1893d5dec339 · inbound

BlockServe: Block-Grained Continuous Batching for High-Throughput Diffusion LLM Serving cites this paper.

BlockServe: Block-Grained Continuous Batching for High-Throughput Diffusion LLM Serving DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T05:43:12.690359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:43:12.690359Z digest=sha256:3f70f97c940b064df5972085192ba98c98d92671c225ae0377d9e9fb82fe6446

Observation a0422bac-16ce-46b9-823a-0e3a1899cb82 · inbound

Beyond Storage: State as a Runtime Control Problem in Parallel and Distributed Systems cites this paper.

Beyond Storage: State as a Runtime Control Problem in Parallel and Distributed Systems DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-01T19:51:06.098713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:51:06.098713Z digest=sha256:51635824076990bd6ff476a0d6728b104d8d4e7acba71df6a873ddf480f174f1