Pith. sign in

Paper Citation Record · LEDGER

Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2401.11181.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.11181 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T10:10:22.887709Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

6
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cd8da1e9-bc32-448d-bd9d-4a7ffbce1000 · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 273

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:39:33.469145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:17755ac1b1507d29d095c7fd95e3f09e42327fcc9b056cc9df7caf1541c4ba00

Observation 7c6ccbed-f64c-4d5b-8fd7-4062c9ac36e2 · inbound

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees cites this paper.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.887709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.887709Z digest=sha256:86f792afffc00e213c6e112d139a5d2c411139e863ffd569ed382751fd5d9b40

Observation bf3a698b-6d10-4bad-bedf-03331e1c72c7 · inbound

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization cites this paper.

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T18:55:04.330334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:55:04.330334Z digest=sha256:73bde2e5d0b058ab026c5c1feed4e73dac46419536df62f60c26dbc9751809e9

Observation d539e8c3-45f1-43ce-a2b0-9593364440c7 · inbound

Beyond the Buzz: A Pragmatic Take on Inference Disaggregation cites this paper.

Beyond the Buzz: A Pragmatic Take on Inference Disaggregation Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:26:32.736179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:26:32.736179Z digest=sha256:8ec248636e6d4e68285d30d289a43aaae97fb7563571645cc31c7777d7324a7c

Observation 5f148349-ce14-4b64-8106-9106552df42e · inbound

Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving cites this paper.

Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:03:53.051327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:03:53.051327Z digest=sha256:7428f299af2edfdc1b8d4828f0c2e8d82ecf4b92efd085791382b439195372a2

Observation f33f8277-8b6a-4991-b6f4-b43908aa318b · inbound

Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference cites this paper.

Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:36.345043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:36.345043Z digest=sha256:993899c3af087e3db30d76e5935cc363d30acc5cdb2576956a7bdd4bb4e3b805

Observation 1f96ef1a-0a70-4f5b-96ed-3cd4aff33490 · inbound

Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism cites this paper.

Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T20:53:31.939770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:53:31.939770Z digest=sha256:d8faae2b5e54bcf98c94400ab1378346a8af24421c411e19b89bca4962fc815a

Observation 26ad8cfa-5076-4b7b-b8bd-67eb2edd7fbc · inbound

FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters cites this paper.

FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:11:03.617983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T07:07:46.228491Z digest=sha256:867ea858642d84bed98648579039743ccf97cef892b289b5f43080c29c6ae204

Observation 1a795543-9eae-427b-8f11-15ae5f4ad8c2 · inbound

STAR: Decode-Phase Rescheduling for LLM Inference cites this paper.

STAR: Decode-Phase Rescheduling for LLM Inference Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T06:12:25.938366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T06:11:28.120860Z digest=sha256:a2732e55ad02527ae0ce50919d9f21332b896afcd681e29bf23ceb7358b752e4

Observation 0af25593-fd88-4ed5-954b-b2bef949192e · inbound

ODMA: On-Demand Memory Allocation Strategy for LLM Serving on LPDDR-Class Accelerators cites this paper.

ODMA: On-Demand Memory Allocation Strategy for LLM Serving on LPDDR-Class Accelerators Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:58:43.005876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T23:54:08.052482Z digest=sha256:87d2eecffff41c1f2b437b5e26e7f52ddd10a47af8c08b5089398088e676e6d6

Observation d20fdeaf-85f2-4752-af64-7f07217c58e6 · inbound

Efficient Multi-round LLM Inference over Disaggregated Serving cites this paper.

Efficient Multi-round LLM Inference over Disaggregated Serving Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 1994

Resolution
unresolved
no resolver link, observed 2026-08-02T23:16:44.283021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:16:44.283021Z digest=sha256:b908e1e7e63cd50ab16c0db4619a4c6f02fadd35f9c270b2544c8eb9f60b07da

Observation 1ce18c33-284a-4fde-b632-ee06b6f0368e · inbound

Coral: Cost-Efficient Multi-LLM Serving over Heterogeneous Cloud GPUs cites this paper.

Coral: Cost-Efficient Multi-LLM Serving over Heterogeneous Cloud GPUs Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:01:07.014104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T16:45:39.853764Z digest=sha256:3e13d67f6be7c5e8be033f9643cbcb6de9f3a8f18afd09d15635011d3a877a24

Observation 193c6fde-95d1-4e3c-bdc9-5cfc75a66375 · inbound

VeriCache: Turning Lossy KV Cache into Lossless LLM Inference cites this paper.

VeriCache: Turning Lossy KV Cache into Lossless LLM Inference Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-19T22:12:51.167065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T22:09:15.782104Z digest=sha256:14080789960b17bf63bc9f092eaccdd0daf7d1267b41bdeb53c5457f67a93447

Observation 93fbeaed-f880-4c22-88d8-a296d194c9da · inbound

AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System cites this paper.

AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:15:17.288802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T03:12:49.028342Z digest=sha256:7b6896998b6a707307102366e0953a6c92f5acb289f02748371ec39d20b6074c

Observation 34fb4efc-43e1-4098-ae80-ff6ca91927d7 · inbound

Human-Less LLM Serving: Quantifying the Human Tax on Throughput cites this paper.

Human-Less LLM Serving: Quantifying the Human Tax on Throughput Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:45:11.801306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T00:44:55.517266Z digest=sha256:ed8baf30806cad11eb035372e3fafc8def591061e734e85bd54fd2f90b2432d7

Observation 66c5f74e-e132-435a-8a50-d755ca4e86e2 · inbound

Geometry-Aware Online Scheduling for LLM Serving: From Theoretical Bound to System Practice cites this paper.

Geometry-Aware Online Scheduling for LLM Serving: From Theoretical Bound to System Practice Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:39:41.985966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T11:17:35.860286Z digest=sha256:55f3c3853c2a94ed2e03a2467b0e935a2d61294e272e4706cb7135ab56fd2c3e

Observation 91619bb5-87eb-472e-8b08-6cf731256fcb · inbound

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving cites this paper.

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T14:04:45.341416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T05:37:13.211613Z digest=sha256:a2d997aab71ae891fb22a47de0a4dc7d4d054af02b89b7a516fb6fc3175253fc

Observation 04af9445-3e67-4eb8-b5ea-8d59b1fe8549 · inbound

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving cites this paper.

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T08:55:35.143870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T07:06:53.318182Z digest=sha256:668af9096365cdc925bd59e49109f05ed97ab597e7e3b969583643daebea73d0

Observation 9045a234-d428-4ed8-947a-61dad41f26e5 · inbound

Towards Load-Aware Prefill Deflection for Disaggregated LLM Serving cites this paper.

Towards Load-Aware Prefill Deflection for Disaggregated LLM Serving Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-03T06:07:40.890885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-03T06:05:06.467649Z digest=sha256:17075894136c021ad00e7bf3abc6aaec0d261db4020006e8c51d1c0bd3823b00

Observation ac815cd7-9513-4cb1-b790-b1cb655a2309 · inbound

CoCoScale: Leveraging Layer-wise Scaling to Unlock the Potential of Online LLM Serving cites this paper.

CoCoScale: Leveraging Layer-wise Scaling to Unlock the Potential of Online LLM Serving Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-11T21:08:24.159706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:08:24.159706Z digest=sha256:b3691ca6be4635b4265d14a97a3bb97b53c6259b480e08421dbaf9cec3479276

Observation ebab7f57-b3e2-4028-9376-3b53766a92b9 · inbound

Sangam: Efficiently Serving Diffusion LLMs with the AR Stack cites this paper.

Sangam: Efficiently Serving Diffusion LLMs with the AR Stack Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-11T20:58:04.182764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:58:04.182764Z digest=sha256:864eb2458b5ef84e97415d3c56008b1586ddf12f2e06c3fa3c60fa599cfe0b6a

Observation 25516814-d561-4c3a-aa0a-af61bd4072d0 · inbound

AutoSLO: Practical Latency SLOs on Cloud Data Warehouses -- Extended Version cites this paper.

AutoSLO: Practical Latency SLOs on Cloud Data Warehouses -- Extended Version Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T03:18:47.408836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T03:18:47.408836Z digest=sha256:31b3965046b66e9256acae19d7ea92488c712c9f71811ef507eb8712098fcbce

Observation 88f9298f-b660-4f3d-8b42-61c58cd035e9 · inbound

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer cites this paper.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:14.862617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:14.862617Z digest=sha256:db327048d4a31beb38ed992ba841939bc96a6155b08c848c3c9c1e253c8e7aa0

Observation b49bd378-5512-4d7b-b8f6-08a197292084 · inbound

Rethinking AI Cloud Infrastructure for Agentic Serving Systems with the Aries Experimentation Framework cites this paper.

Rethinking AI Cloud Infrastructure for Agentic Serving Systems with the Aries Experimentation Framework Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T14:14:10.457557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:14:10.457557Z digest=sha256:c3e92771ceeb5bf948a4b9553fe120d6aebbc72223b34febeeb946d0cc814411

Observation 7d2fc877-c3e9-4a5b-9681-8eb28e071b5f · inbound

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling cites this paper.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:22.951059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:22.951059Z digest=sha256:bbd105747f79fe80dab13e1c2ab899a4e09d752ebe15b647a534128b5e399a5a

Observation ea77d351-5772-4605-b59d-d2342eac07e3 · inbound

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure cites this paper.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.353271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.353271Z digest=sha256:f628af896cc9ba92ca3135e96eed0beb9d3bcaf91ed89d9a9765d9618239c2eb