Pith. sign in

Paper Citation Record · LEDGER

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving

As of 19 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2608.08382.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08382 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:42:24.754153Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d9a16fe4-8913-40d9-9507-fc9a7cbf5576 · outbound

This paper cites Amazon EC2 Instance Types – Burstable Performance Instances (T3).

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Amazon EC2 Instance Types – Burstable Performance Instances (T3)

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:42:25.695693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-14T04:42:24.558207Z digest=sha256:703890186390b7bd3559d7577559e72f8b044a06ec693cad1e787aed8b2512e0

Observation b16004a4-1535-4515-8710-48f1cd614722 · outbound

This paper cites Anthropic API.https://www.anthropic.com/api, 2025.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Anthropic API.https://www.anthropic.com/api, 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:42:25.534772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-14T04:42:24.563235Z digest=sha256:39083dd8ab07c3b968d558b541326cbd509229147539bd7549d559cabc9ae9d3

Observation e963fe0b-4c44-4bec-8014-3b1f40675fd7 · outbound

This paper cites Xen and the art of virtualization.ACM SIGOPS operating systems review, 37(5):164–177, 2003.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Xen and the art of virtualization.ACM SIGOPS operating systems review, 37(5):164–177, 2003

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:42:25.449733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-14T04:42:24.567834Z digest=sha256:38179be4adede60aa850cf6fa353dd08ebed5afa5143c885bf75e9a8164aa993

Observation de6244fb-1abf-4d20-b503-d4a1b53c0793 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving On the Opportunities and Risks of Foundation Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T04:42:24.571959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:42:24.571959Z digest=sha256:ed7a3132c1a67449dc29580635c33ef097b766b44170991129fc80875608e576

Observation 98058cc8-0fe4-494f-8816-2ade85e32230 · outbound

This paper cites GPT-4o System Card.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving GPT-4o System Card

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-14T04:42:24.577605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:42:24.577605Z digest=sha256:dbcba7cb832eb5caa0010d7be0630739bedf6a60d828891628c2d815447fc92b

Observation ed21b4eb-dfbf-493c-b345-40da02c721f2 · outbound

This paper cites Predicting llm inference latency: A roofline-driven ml method.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Predicting llm inference latency: A roofline-driven ml method

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:42:25.430409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-14T04:42:24.583654Z digest=sha256:3625a1d54b4016ac78ba3cf4de63eb6700679d55086764af863923eb5b22f6c3

Observation 41ddca52-034c-4ff7-beaf-bc799f2b5df1 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-14T04:42:24.599303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:42:24.599303Z digest=sha256:44635dfe8d370f7c1613e590af5c166e357a80d96a3ef0920a635b21d4c17a44

Observation 03a6f21d-fa99-4fab-ab9d-a605b9afa755 · outbound

This paper cites Plato: Plan to efficient decode for large language model inference.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Plato: Plan to efficient decode for large language model inference

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:42:25.353397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-14T04:42:24.611338Z digest=sha256:266bdcf70586b445b41ec3b864e6dd86cc9400485f1a4442d03e34002987913e

Observation f65ed111-c9b3-4c86-bf3e-cc00fed2f664 · outbound

This paper cites Compute Or Load KV Cache? Why Not Both?.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Compute Or Load KV Cache? Why Not Both?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T04:42:24.617477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:42:24.617477Z digest=sha256:f98cbf9f748d1da42ef4d698e5e8f348974dc18f782cf4d5c14942f9d0cc27d3

Observation 92a87986-2028-40cb-b234-198314006b62 · outbound

This paper cites s3: Increasing gpu utilization during generative inference for higher throughput.Advances in Neural Information Processing Systems, 36:18015–18027, 2023.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving s3: Increasing gpu utilization during generative inference for higher throughput.Advances in Neural Information Processing Systems, 36:18015–18027, 2023

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:42:25.340557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-14T04:42:24.627354Z digest=sha256:1b27e702de95311b840644f472b6b94da00b9f21359a913eeb819b2e6e9d5024

Observation 7687dbd3-5195-445c-a1a1-5564dab4481d · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T04:42:24.655668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:42:24.655668Z digest=sha256:f903f690dbb5a2ab15a7c230cc695433b5cd06463bd1164508993c96397f4a8b

Observation bcffb26d-7b89-4587-9b70-57120cafd825 · outbound

This paper cites Cgroups.Available on-line at: http://www.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Cgroups.Available on-line at: http://www

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:42:25.294760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-14T04:42:24.660374Z digest=sha256:41af44c878aa824d38d1b96ae99f5a4de6e3136a7509a38ce25b25f7989162f9

Observation d6eaa769-159d-467f-b7b3-6e240984dc41 · outbound

This paper cites Docker: lightweight linux containers for consistent development and deployment.Linux j, 239(2):2, 2014.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Docker: lightweight linux containers for consistent development and deployment.Linux j, 239(2):2, 2014

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:42:25.239920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-14T04:42:24.664605Z digest=sha256:1d794bde7a79bb716cc5b16d40b39303b168056c67ad3cc852a30258d8834f54

Observation 8c9f05d0-ed81-4908-a366-f7cc6d4c337d · outbound

This paper cites Introducing llama 3.1: The next generation of open models.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Introducing llama 3.1: The next generation of open models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:42:25.174750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-14T04:42:24.668261Z digest=sha256:8668261b8dac9ab0a93f5d3c9e7c751adcedd797b632dab58df108369949a4d7

Observation d9f7c4b6-7752-4e4b-936e-b4bbd6d43ca0 · outbound

This paper cites NVIDIA Multi-Instance GPU (MIG) User Guide.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving NVIDIA Multi-Instance GPU (MIG) User Guide

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:42:25.134749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-14T04:42:24.671996Z digest=sha256:82dad78698dabe6dbcb6961711e910b008ee61119e1e0c5767da1357bf0a6fc3

Observation ea971a8b-95b7-4af9-993e-95a112d112c3 · outbound

This paper cites OpenAI API.https://openai.com/api/, 2025.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving OpenAI API.https://openai.com/api/, 2025

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:42:25.080487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-14T04:42:24.676161Z digest=sha256:6f45feb433a7bae4b45e8a1e45690a1a1b726701fbfaccaa961a178c0ed50aac

Observation d9a830fd-438d-4493-a156-12e5b7c64d9f · outbound

This paper cites Fairness in serving large language models.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Fairness in serving large language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:42:25.059925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-14T04:42:24.680943Z digest=sha256:3a91fbfa62ee23ff7ec2d42b4b4a3f17ec6c51d2d0247d0dfa5a4e7c8abdec3d

Observation 0ea71899-891b-4a5c-90a5-00c332b866c0 · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Qwen2.5: A party of foundation models, September 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T04:42:24.685389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:42:24.685389Z digest=sha256:8a2983dade255e596ab31bd7fb1d2587b9010dd2e09a6e7d74d19a160739b6ac

Observation 242832dd-cf0d-41fc-8950-beaf09c15ab0 · outbound

This paper cites AIBrix: Towards Scalable, Cost-Effective Large Language Model Inference Infrastructure.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving AIBrix: Towards Scalable, Cost-Effective Large Language Model Inference Infrastructure

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-14T04:42:24.689292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:42:24.689292Z digest=sha256:0ff1ade207ef9ad2c3ab596cb2b8a05295c74c9c20e8a697f0373a379a7ad7e3

Observation 4bbb7180-67a2-4bba-8b33-34fd237eb44a · outbound

This paper cites HeterMoE: Efficient Training of Mixture-of-Experts Models on Heterogeneous GPUs.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving HeterMoE: Efficient Training of Mixture-of-Experts Models on Heterogeneous GPUs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-14T04:42:24.697781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:42:24.697781Z digest=sha256:959b3ddf45d61fadedb86ac497048fa54a23fd616e06a63a6acb20ee5105a85f

Observation 884f6640-65be-4964-9c1f-d5909c70f64a · outbound

This paper cites RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-14T04:42:24.724753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:42:24.724753Z digest=sha256:295829516a5ed15ac3f8c48066df744c38b69a1065c4d227855e55b8508253e5

Observation 768fbee0-c871-44cd-9440-f7b38a176288 · outbound

This paper cites LLM Inference Unveiled: Survey and Roofline Model Insights.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-14T04:42:24.748447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:42:24.748447Z digest=sha256:e40a5c0a532cb63b73b7f0b5bc5768cba68fe7dbf07c07182e00dae4f69b6713

Observation 759bacf0-3e8e-4c02-b100-4947e78a16ef · outbound

This paper cites Eagle: Efficient Training-Free Router for Multi-LLM Inference.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Eagle: Efficient Training-Free Router for Multi-LLM Inference

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-14T04:42:24.754153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:42:24.754153Z digest=sha256:9c6f84ee36ad5aaacbb0652e9fd6dc6329e96eb997e126e7fc8d368a954be233

Pith citing papers

No inbound Pith citation observations are available.