Pith. sign in

Paper Citation Record · LEDGER

DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2207.00032.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2207.00032 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:49:09.679785Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T22:59:03.322421Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 39900118-976a-4971-8a9c-d9c4cbe878d9 · inbound

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models cites this paper.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.431176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:a78bb9d7fe16157a1011f5549d2a42deef034b84fd2c4d53d276acf5e8d57556

Observation 08dccdd5-98d0-49e9-8fc2-68424e18710b · inbound

Efficient Memory Management for Large Language Model Serving with PagedAttention cites this paper.

Efficient Memory Management for Large Language Model Serving with PagedAttention DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:03:07.819111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:0267baf3879d4d6ac347254a0f54f9878d06b4477ac2305306bb88afd6c75e32

Observation 5f0de8ea-96d9-4af1-b3e4-258f1bfe3673 · inbound

LLaSA: Large Language and Structured Data Assistant cites this paper.

LLaSA: Large Language and Structured Data Assistant DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T19:24:49.787019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:24:49.787019Z digest=sha256:c7873b0fe65ca107ae52fea85f7ff4fa713a530e0476fff76a9408db8ed5f074

Observation 9bdbfd31-5e9d-4071-8cf2-63ca312898c2 · inbound

FastSwitch: Optimizing Context Switching Efficiency in Fairness-aware Large Language Model Serving cites this paper.

FastSwitch: Optimizing Context Switching Efficiency in Fairness-aware Large Language Model Serving DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T11:17:00.895481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:17:00.895481Z digest=sha256:8dc1d35df5042517259151ded222509919d1b49d107dcef55f0b8cc6c4091d83

Observation d873e317-48e0-4e95-9548-57e999c07e5a · inbound

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference cites this paper.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.202032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.202032Z digest=sha256:c99626a0357e7a66b8fb611daef06dcb2cc1d3a702fa15d4dc5fd3063a1a4bbc

Observation 77777fa4-ed36-491d-ae1c-725ae2020a07 · inbound

iServe: An Intent-based Serving System for LLMs cites this paper.

iServe: An Intent-based Serving System for LLMs DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T21:37:14.126135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:37:14.126135Z digest=sha256:abf9d5576a431fd34ad161092818f1337939bac3ff1e9a89283958aa3542516a

Observation dd5d2662-76b1-44e0-91a3-18f9d06fb1ef · inbound

Aero-LLM: A Distributed Framework for Secure UAV Communication and Intelligent Decision-Making cites this paper.

Aero-LLM: A Distributed Framework for Secure UAV Communication and Intelligent Decision-Making DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T05:18:07.705592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T05:18:07.705592Z digest=sha256:6cd32388945e9a80227dc86ac85a7fed49e1a39d03c4cc8dfcc738f9de14d49f

Observation 8ade0e7c-048e-4156-9adb-9da31078a3e0 · inbound

Taming the Titans: A Survey of Efficient LLM Inference Serving cites this paper.

Taming the Titans: A Survey of Efficient LLM Inference Serving DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T05:49:09.679785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:49:09.679785Z digest=sha256:9822e8b055c1cdd30e65ee8c966cfd7337c4b48044ceb68c9bbb73810ce9c82f

Observation 58157a7a-196f-4952-a0fe-cc8c9c97bf12 · inbound

SCOPE: Compress Mathematical Reasoning Steps for Efficient Automated Process Annotation cites this paper.

SCOPE: Compress Mathematical Reasoning Steps for Efficient Automated Process Annotation DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:16.021201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:39:16.021201Z digest=sha256:c1b7516097491fc42d43f14e67d50b4120a520eaf4ab434bb74f964fb5da6fdf

Observation c8abc7c3-bda3-4139-8f81-e0cb39023140 · inbound

Hardware-Efficient Attention for Fast Decoding cites this paper.

Hardware-Efficient Attention for Fast Decoding DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:29.198408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:29.198408Z digest=sha256:b312e5edf6b3416b5e1715abb217009b9c9e00b31b068bdf14071151c1959872

Observation 2f5b486c-3f15-42ee-8066-5e3bcf0cabab · inbound

EvolveSearch: An Iterative Self-Evolving Search Agent cites this paper.

EvolveSearch: An Iterative Self-Evolving Search Agent DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:55.546676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:55.546676Z digest=sha256:85a9c84bdd923a372ca9958e5f4e4ac4b0c0e5ef8f5fe53cfa1914306bdcc562

Observation ef580638-c827-423f-b193-3e9c54128cf6 · inbound

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes cites this paper.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.417052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.417052Z digest=sha256:e6f8ed13576f1d651cf11c6efedbbac66259194754530bf58ffa1139f657c91c

Observation c331b2c8-9376-400f-81ec-44c44298a02d · inbound

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes cites this paper.

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-15T18:03:15.421740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:03:15.421740Z digest=sha256:5b88e1591f8c66e9af6d1e0b9b1f6de187ac8487194bb2bba818134bebefcfb4

Observation 1bc7a03b-05fc-4f7c-abe1-632d5628a92a · inbound

Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference cites this paper.

Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:51:42.268813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T17:47:58.019030Z digest=sha256:714ea0f6716bae9506a50e5ac38c59b78e781d58c12b545730027ce4253b7f62

Observation de56e279-cca7-41aa-b6f5-8e0a892fad08 · inbound

SweetSpot: An Analytical Model for Predicting Energy Efficiency of LLM Inference cites this paper.

SweetSpot: An Analytical Model for Predicting Energy Efficiency of LLM Inference DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:07:29.589719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T07:05:56.380048Z digest=sha256:bae61981f1f49250ffff1402773243db0a249e5b85445b745fafb2182a5f7549

Observation 48682b07-9fa5-4f44-9259-f8e962a6d74b · inbound

Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding cites this paper.

Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:31:01.439048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T14:50:37.022338Z digest=sha256:1d3a9387ee691d0642b15268bb09d865021a759eabc431dae551bd0c733de31b

Observation 97c1c357-ece0-4476-ad31-83fd25f2782d · inbound

Predict-then-Diffuse: Adaptive Response Length for Compute-Budgeted Inference in Diffusion LLMs cites this paper.

Predict-then-Diffuse: Adaptive Response Length for Compute-Budgeted Inference in Diffusion LLMs DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:41:08.897769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T17:18:05.789417Z digest=sha256:7e65515e650a30860b1554f0ead7feabfa7d0e8dd87b2447db6fdb5cca50e462

Observation 4b7cea5c-8342-4858-b4d1-0379d2ba33f4 · inbound

Predict-then-Diffuse: Adaptive Response Length for Compute-Budgeted Inference in Diffusion LLMs cites this paper.

Predict-then-Diffuse: Adaptive Response Length for Compute-Budgeted Inference in Diffusion LLMs DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:49:48.919062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T06:49:16.271699Z digest=sha256:3fd3c1aec72df747f826bebc9f1df4130d5d9c15efa2dacf82b89ce315871b83

Observation e122ca3f-c6ac-4e43-8450-0c2acab501e7 · inbound

Predict-then-Diffuse: Adaptive Response Length for Compute-Budgeted Inference in Diffusion LLMs cites this paper.

Predict-then-Diffuse: Adaptive Response Length for Compute-Budgeted Inference in Diffusion LLMs DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T02:23:43.898162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:23:43.898162Z digest=sha256:d6846825b070e92cf799403ea7fb445d0119bef4a8ff47e90465467ce6010e4f

Observation bafbf0de-24ce-459a-8c15-d717f12901b7 · inbound

ShardTensor: Domain Parallelism for Scientific Machine Learning cites this paper.

ShardTensor: Domain Parallelism for Scientific Machine Learning DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-13T03:07:09.104038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T03:03:32.316700Z digest=sha256:b49141aa575c643ab0c153787aadcd93a884d2314739f91f0f00348bba9f5774

Observation a83bf6aa-9455-413c-b910-edc84e1334a9 · inbound

A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability cites this paper.

A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:38:09.365091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T07:35:32.225708Z digest=sha256:26117595f1bd61ab5836db6af2d3252e9c2e7b413d238d2021c77eaa8247cf18

Observation 0bf0fca9-42f7-4096-af22-a01062156484 · inbound

A Paired Testing Protocol for Batch-Conditioned Refusal Robustness in LLM Serving cites this paper.

A Paired Testing Protocol for Batch-Conditioned Refusal Robustness in LLM Serving DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:03:47.843565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-29T18:03:19.798168Z digest=sha256:fe532a195743c7c1b3d114a1ffb11e0a67ac89b978c230ae98b32d34e4a6c643

Observation 2ca665c7-8363-4c6b-8dd6-93f5aa9b6272 · inbound

AoiZora: Topology-Aware Auto-Parallel Optimization for Inference of Diffusion Transformers cites this paper.

AoiZora: Topology-Aware Auto-Parallel Optimization for Inference of Diffusion Transformers DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T22:59:03.324627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T23:07:13.368633Z digest=sha256:19f112d344aba84b1267a52e2bb0e3be8296c1c055a3a2c9ce020b959ea3f833

Observation 009c2579-6677-4a94-b3b7-aa8f7fa29de0 · inbound

Optimizing Teacher-Student Partitioning for Scalable Knowledge Distillation on HPC Systems cites this paper.

Optimizing Teacher-Student Partitioning for Scalable Knowledge Distillation on HPC Systems DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T17:45:52.278510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T03:06:29.733729Z digest=sha256:4adfd519bae811625d2f584adc48603685d3de2cc84a0dfe7f34b6129971e349

Observation 46693fd0-7386-4723-be2c-47a2a130340c · inbound

GSRQ: Gain-Shape Residual Quantization for Sub-1-bit KV Cache cites this paper.

GSRQ: Gain-Shape Residual Quantization for Sub-1-bit KV Cache DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T15:57:06.415052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-02T15:55:40.177742Z digest=sha256:25c3f135b39c6ca9f849291646914bffd0faaaaab8ec2ede1440d2cfc638038e