Pith. sign in

Paper Citation Record · LEDGER

Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2401.11181.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.11181 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:26:32.736179Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

6
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cd8da1e9-bc32-448d-bd9d-4a7ffbce1000 · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 273

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:39:33.469145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:defd7ce00d3ab70999392f33f11b0611a9514022daa430c8ab84ec53ba136342

Observation d539e8c3-45f1-43ce-a2b0-9593364440c7 · inbound

Beyond the Buzz: A Pragmatic Take on Inference Disaggregation cites this paper.

Beyond the Buzz: A Pragmatic Take on Inference Disaggregation Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:26:32.736179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:26:32.736179Z digest=sha256:8ec248636e6d4e68285d30d289a43aaae97fb7563571645cc31c7777d7324a7c

Observation 5f148349-ce14-4b64-8106-9106552df42e · inbound

Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving cites this paper.

Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:03:53.051327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:03:53.051327Z digest=sha256:7428f299af2edfdc1b8d4828f0c2e8d82ecf4b92efd085791382b439195372a2

Observation f33f8277-8b6a-4991-b6f4-b43908aa318b · inbound

Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference cites this paper.

Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:36.345043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:36.345043Z digest=sha256:6c2c3dabd8db2a10be65e2a91a026999cbcd6c078f6e8a11d08bd5a720c83097

Observation 1f96ef1a-0a70-4f5b-96ed-3cd4aff33490 · inbound

Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism cites this paper.

Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T20:53:31.939770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:53:31.939770Z digest=sha256:d8faae2b5e54bcf98c94400ab1378346a8af24421c411e19b89bca4962fc815a

Observation 26ad8cfa-5076-4b7b-b8bd-67eb2edd7fbc · inbound

FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters cites this paper.

FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:11:03.617983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T07:07:46.228491Z digest=sha256:e3d88fb15db5f70e0b63b21a92b2533fff5558631c7d008527bcf62a011690e2

Observation 1a795543-9eae-427b-8f11-15ae5f4ad8c2 · inbound

STAR: Decode-Phase Rescheduling for LLM Inference cites this paper.

STAR: Decode-Phase Rescheduling for LLM Inference Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T06:12:25.938366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T06:11:28.120860Z digest=sha256:93c1067d019cc6c8db48ee0ace262d2c4bb0e47952b72b002f8baf2242ad8327

Observation 0af25593-fd88-4ed5-954b-b2bef949192e · inbound

ODMA: On-Demand Memory Allocation Strategy for LLM Serving on LPDDR-Class Accelerators cites this paper.

ODMA: On-Demand Memory Allocation Strategy for LLM Serving on LPDDR-Class Accelerators Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:58:43.005876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T23:54:08.052482Z digest=sha256:35c5cb6fa946f03b7744061e6130cbd06a46d4fd25be11db060b71a5e0b19582

Observation d20fdeaf-85f2-4752-af64-7f07217c58e6 · inbound

Efficient Multi-round LLM Inference over Disaggregated Serving cites this paper.

Efficient Multi-round LLM Inference over Disaggregated Serving Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 1994

Resolution
unresolved
no resolver link, observed 2026-08-02T23:16:44.283021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:16:44.283021Z digest=sha256:b908e1e7e63cd50ab16c0db4619a4c6f02fadd35f9c270b2544c8eb9f60b07da

Observation 1ce18c33-284a-4fde-b632-ee06b6f0368e · inbound

Coral: Cost-Efficient Multi-LLM Serving over Heterogeneous Cloud GPUs cites this paper.

Coral: Cost-Efficient Multi-LLM Serving over Heterogeneous Cloud GPUs Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:01:07.014104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T16:45:39.853764Z digest=sha256:9dfbfee199448e09793e3efc7abbb152bbb39aedaeeb8aed1f209b63b8623cdb

Observation 193c6fde-95d1-4e3c-bdc9-5cfc75a66375 · inbound

VeriCache: Turning Lossy KV Cache into Lossless LLM Inference cites this paper.

VeriCache: Turning Lossy KV Cache into Lossless LLM Inference Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-19T22:12:51.167065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T22:09:15.782104Z digest=sha256:416fe924cd35ad550621f4f6dee21193135f3bb10613ba1abd8d43a50518e913

Observation 93fbeaed-f880-4c22-88d8-a296d194c9da · inbound

AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System cites this paper.

AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:15:17.288802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T03:12:49.028342Z digest=sha256:146df158ea7a0aacff83439746240ede729d130c3cdcedad435fbb7af36309fe

Observation 34fb4efc-43e1-4098-ae80-ff6ca91927d7 · inbound

Human-Less LLM Serving: Quantifying the Human Tax on Throughput cites this paper.

Human-Less LLM Serving: Quantifying the Human Tax on Throughput Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:45:11.801306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T00:44:55.517266Z digest=sha256:85bbe822a8ad9266c0a4b7db3225e71780d62c81a892ea8f201c2493ab5bcb70

Observation 66c5f74e-e132-435a-8a50-d755ca4e86e2 · inbound

Geometry-Aware Online Scheduling for LLM Serving: From Theoretical Bound to System Practice cites this paper.

Geometry-Aware Online Scheduling for LLM Serving: From Theoretical Bound to System Practice Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:39:41.985966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T11:17:35.860286Z digest=sha256:4c769fdecd177383760fbd91d95b41f4417c662c20e35515c794990a7a6f10bf

Observation 91619bb5-87eb-472e-8b08-6cf731256fcb · inbound

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving cites this paper.

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T14:04:45.341416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T05:37:13.211613Z digest=sha256:119b4eade36684510f211181a76df44a236d8050021913f411ef9b03d2dcc5a2

Observation 04af9445-3e67-4eb8-b5ea-8d59b1fe8549 · inbound

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving cites this paper.

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T08:55:35.143870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T07:06:53.318182Z digest=sha256:f93de88399458d9539af12745f5656d14bc095c8872efeadb30809abd6320aa0

Observation 9045a234-d428-4ed8-947a-61dad41f26e5 · inbound

Towards Load-Aware Prefill Deflection for Disaggregated LLM Serving cites this paper.

Towards Load-Aware Prefill Deflection for Disaggregated LLM Serving Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-03T06:07:40.890885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T06:05:06.467649Z digest=sha256:a76b2d407abd2c55166cbdf255820bbd997ade84b6a00b84e5cee4eeb9550284

Observation ac815cd7-9513-4cb1-b790-b1cb655a2309 · inbound

CoCoScale: Leveraging Layer-wise Scaling to Unlock the Potential of Online LLM Serving cites this paper.

CoCoScale: Leveraging Layer-wise Scaling to Unlock the Potential of Online LLM Serving Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-11T21:08:24.159706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:08:24.159706Z digest=sha256:b3691ca6be4635b4265d14a97a3bb97b53c6259b480e08421dbaf9cec3479276

Observation ebab7f57-b3e2-4028-9376-3b53766a92b9 · inbound

Sangam: Efficiently Serving Diffusion LLMs with the AR Stack cites this paper.

Sangam: Efficiently Serving Diffusion LLMs with the AR Stack Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-11T20:58:04.182764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:58:04.182764Z digest=sha256:8945bb39fdc53c404366cf876fc552cb3005b9b0755c234ba267009132445651

Observation 25516814-d561-4c3a-aa0a-af61bd4072d0 · inbound

AutoSLO: Practical Latency SLOs on Cloud Data Warehouses -- Extended Version cites this paper.

AutoSLO: Practical Latency SLOs on Cloud Data Warehouses -- Extended Version Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T03:18:47.408836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T03:18:47.408836Z digest=sha256:31b3965046b66e9256acae19d7ea92488c712c9f71811ef507eb8712098fcbce

Observation 88f9298f-b660-4f3d-8b42-61c58cd035e9 · inbound

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer cites this paper.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:14.862617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:14.862617Z digest=sha256:db327048d4a31beb38ed992ba841939bc96a6155b08c848c3c9c1e253c8e7aa0

Observation b49bd378-5512-4d7b-b8f6-08a197292084 · inbound

Rethinking AI Cloud Infrastructure for Agentic Serving Systems with the Aries Experimentation Framework cites this paper.

Rethinking AI Cloud Infrastructure for Agentic Serving Systems with the Aries Experimentation Framework Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T14:14:10.457557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:14:10.457557Z digest=sha256:c3e92771ceeb5bf948a4b9553fe120d6aebbc72223b34febeeb946d0cc814411

Observation 7d2fc877-c3e9-4a5b-9681-8eb28e071b5f · inbound

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling cites this paper.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:22.951059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:22.951059Z digest=sha256:bbd105747f79fe80dab13e1c2ab899a4e09d752ebe15b647a534128b5e399a5a