Pith. sign in

Paper Citation Record · LEDGER

On Evaluating Performance of LLM Inference Serving Systems

As of 7 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 5 inbound Pith citation observations for arXiv:2507.09019.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09019 v1

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:11:16.388804Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T15:10:12.647037Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T23:52:49.103329Z

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved10
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 28100343-384c-4f3d-b8fb-69ae1e05919f · outbound

This paper cites MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving.

On Evaluating Performance of LLM Inference Serving Systems MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:11:15.829163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:11:15.829163Z digest=sha256:c7636af8f1d82d745c2594e4eef29b6d5ba5ef9f21add20b0809272707c4e545

Observation aa5dc71d-861b-48c0-9a3a-05f55e5e1dc4 · outbound

This paper cites CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving.

On Evaluating Performance of LLM Inference Serving Systems CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:11:15.906010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:11:15.906010Z digest=sha256:196ada9cf6bc71df4652c41d7746a2b4920106743bdb4e89ca203462f0bbdca5

Observation 9ebb2090-66ea-44ff-9cb0-72c5a1ba1f84 · outbound

This paper cites Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving.

On Evaluating Performance of LLM Inference Serving Systems Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:11:15.969489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:11:15.969489Z digest=sha256:fb186888ba58ca8f27b825dc0d0e730b2e0810e5dc02ea73b04f037cf7c942bb

Observation dde8cc21-f904-417b-b360-f882410d750e · outbound

This paper cites Guan Wang, Sijie Cheng, Xianyuan Zhan, Xiangang Li, Sen Song, and Yang Liu.

On Evaluating Performance of LLM Inference Serving Systems Guan Wang, Sijie Cheng, Xianyuan Zhan, Xiangang Li, Sen Song, and Yang Liu

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:11:16.246971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:11:16.246971Z digest=sha256:f30a2670d8273d2aa109c83ae85f7b7beb65b7bb0ddd7c3ca5eb609651d5a872

Observation 9182bd48-f2ae-424d-a307-cd01d9864f44 · outbound

This paper cites LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism.

On Evaluating Performance of LLM Inference Serving Systems LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:11:16.254256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:11:16.254256Z digest=sha256:205c15b76be0f5ff330d5eb5bd6ed905b1b2ab4a5f0884aabe5b72ce1c07c55a

Observation 11274d95-91cf-475e-b1a9-373b85f355a5 · outbound

This paper cites NanoFlow: Towards Optimal Large Language Model Serving Throughput.

On Evaluating Performance of LLM Inference Serving Systems NanoFlow: Towards Optimal Large Language Model Serving Throughput

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:11:16.262684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:11:16.262684Z digest=sha256:8c40cea19d5ee067a58ec27d62ea4d73f2b5574f8247b36f4ca98085fdeab47d

Observation 55532324-5a12-45f1-b6fd-75421016dbd9 · outbound

This paper cites NanoFlow: Towards Optimal Large Language Model Serving Throughput.

On Evaluating Performance of LLM Inference Serving Systems NanoFlow: Towards Optimal Large Language Model Serving Throughput

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:11:16.322242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:11:16.322242Z digest=sha256:8182292d6811e40188c146ff069ac5dec304270b05c5a391c33193e13c2a1f5c

Observation a9a0e147-24c3-4da7-9214-a9280e669f1b · outbound

This paper cites Parameter T uning (✓) In our judgment, the baselines don’t need parameter tuning.

On Evaluating Performance of LLM Inference Serving Systems Parameter T uning (✓) In our judgment, the baselines don’t need parameter tuning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:11:16.740082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:11:16.388804Z digest=sha256:57fec45c380b92b5b30a36e5100ffe4f994dcd21f2eb8f4359c8404939334aff

Observation c5d486c1-2b71-4d34-973c-666392469670 · outbound

This paper cites S-LoRA: Serving Thousands of Concurrent LoRA Adapters.

On Evaluating Performance of LLM Inference Serving Systems S-LoRA: Serving Thousands of Concurrent LoRA Adapters

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T18:11:16.146190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:11:16.146190Z digest=sha256:ebf2191846132857b3ca4c1cb406a073c6fa09f9fbda8f5b44239855a64be875

Observation 6861722a-35ce-4626-ab63-0f76e1044013 · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

On Evaluating Performance of LLM Inference Serving Systems Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T18:11:16.045289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:11:16.045289Z digest=sha256:6a3caa9ff79121438490010e5e3e1401f1fd1936515649b0724c8b80ead1233a

Observation 284a4b5b-dfb1-44f0-a2bb-a24dd331d451 · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

On Evaluating Performance of LLM Inference Serving Systems Accelerating Large Language Model Decoding with Speculative Sampling

Reference 2023

Resolution
malformed identifier
no resolver link, observed 2026-08-06T18:11:15.769967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:11:15.769967Z digest=sha256:24252dc9e077817ed8e251e4d7ca16da77cfd694c1a8496c8a252cb1ea207096

Observation 9afccbaa-727d-41b4-8f2c-29152b6b8767 · outbound

This paper cites Efficient LLM Scheduling by Learning to Rank.

On Evaluating Performance of LLM Inference Serving Systems Efficient LLM Scheduling by Learning to Rank

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T18:11:15.858434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:11:15.858434Z digest=sha256:7350ae4e81fd97e0e38bbfe2d8b62f57e8ca5d7b748d8e91d67c7e2678bf7fcb

Pith citing papers

Observation d3452ba6-8358-48f7-8b0e-968225f2155d · inbound

Specification and Detection of LLM Code Smells cites this paper.

Specification and Detection of LLM Code Smells On Evaluating Performance of LLM Inference Serving Systems

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T15:10:12.647037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:10:12.647037Z digest=sha256:c92073733fdde0333921819e31081cdfcc9d25735b1a798a921d3332d82a37a2

Observation e853f491-9ddf-45e8-957b-e4d986b6e523 · inbound

CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding cites this paper.

CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding On Evaluating Performance of LLM Inference Serving Systems

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:32:36.437742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T08:30:50.984873Z digest=sha256:32221c8578c25bbdbc6554af4cd50f84ab544adbfe0f77905b8be0401cf1a775

Observation 9ad232a9-fc3f-4fd8-b1f5-c2238c852d2a · inbound

Characterizing Performance-Energy Trade-offs of Large Language Models in Multi-Request Workflows cites this paper.

Characterizing Performance-Energy Trade-offs of Large Language Models in Multi-Request Workflows On Evaluating Performance of LLM Inference Serving Systems

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:30:00.285701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:28:38.809950Z digest=sha256:10fe79713e8ae8b7296f481ab14f0326232817a5854f0f08c9bee3c29ebae322

Observation 8598cbcc-8abd-43f3-973d-68d1c2d75d7f · inbound

SAGE: Selective Attention-Guided Extraction for Token-Efficient Document Indexing cites this paper.

SAGE: Selective Attention-Guided Extraction for Token-Efficient Document Indexing On Evaluating Performance of LLM Inference Serving Systems

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:32:52.126778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T08:32:02.222528Z digest=sha256:3b5888783d6b611f277c0357a0acb49bec2a8e7eea2a76daa157fd66c1d5aa68

Observation 9f3aa297-b3ea-47f7-8508-e3d155580047 · inbound

RTP-LLM: High-Performance Alibaba LLM Inference Engine cites this paper.

RTP-LLM: High-Performance Alibaba LLM Inference Engine On Evaluating Performance of LLM Inference Serving Systems

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:52:49.104862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T23:52:40.763228Z digest=sha256:2de2787d6fcf4a80993280bbc11e38ac046b0d782cefb8776ceb24afc499e5ce