Pith. sign in

Paper Citation Record · LEDGER

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving

As of 19 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2608.08382.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08382 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:42:24.754153Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d9a16fe4-8913-40d9-9507-fc9a7cbf5576 · outbound

This paper cites Amazon EC2 Instance Types – Burstable Performance Instances (T3).

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Amazon EC2 Instance Types – Burstable Performance Instances (T3)

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:42:25.695693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-14T04:42:24.558207Z digest=sha256:b18009480be1aa9f16129bdb8605a3343a076f152a7f6930117a2faf1cff1e63

Observation b16004a4-1535-4515-8710-48f1cd614722 · outbound

This paper cites Anthropic API.https://www.anthropic.com/api, 2025.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Anthropic API.https://www.anthropic.com/api, 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:42:25.534772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-14T04:42:24.563235Z digest=sha256:afc104693a0cbbb5baeeb07cf55a4ed6eee4aeb266e9d97851b285a08e3f2fdd

Observation e963fe0b-4c44-4bec-8014-3b1f40675fd7 · outbound

This paper cites Xen and the art of virtualization.ACM SIGOPS operating systems review, 37(5):164–177, 2003.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Xen and the art of virtualization.ACM SIGOPS operating systems review, 37(5):164–177, 2003

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:42:25.449733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-14T04:42:24.567834Z digest=sha256:43313edff5ac38ef9cb41acedf8191bd23623a1c5998b16bf9acddfbb48ce0e4

Observation de6244fb-1abf-4d20-b503-d4a1b53c0793 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving On the Opportunities and Risks of Foundation Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T04:42:24.571959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:42:24.571959Z digest=sha256:4ed2fad7087a236c7c17216e8cf415afbbf93f9423e0b60205e7022423059ec2

Observation 98058cc8-0fe4-494f-8816-2ade85e32230 · outbound

This paper cites GPT-4o System Card.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving GPT-4o System Card

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-14T04:42:24.577605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:42:24.577605Z digest=sha256:5eb2a38046284a8b5726bc0a5a0c38c2e213d4088bca79aaaa337b1657066be8

Observation ed21b4eb-dfbf-493c-b345-40da02c721f2 · outbound

This paper cites Predicting llm inference latency: A roofline-driven ml method.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Predicting llm inference latency: A roofline-driven ml method

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:42:25.430409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-14T04:42:24.583654Z digest=sha256:5549498df85d35056e8337cf8ed05e6272a30b42582ac051394a86084f179709

Observation 41ddca52-034c-4ff7-beaf-bc799f2b5df1 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-14T04:42:24.599303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:42:24.599303Z digest=sha256:1ca8eb759ec4d8d643ffd7088675f6d49e1fc31d35fc3d804eab42d7e0a34006

Observation 03a6f21d-fa99-4fab-ab9d-a605b9afa755 · outbound

This paper cites Plato: Plan to efficient decode for large language model inference.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Plato: Plan to efficient decode for large language model inference

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:42:25.353397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-14T04:42:24.611338Z digest=sha256:4221cef6cdaec1fd8dcfa2b5ce961cbbb2ce497ce11732460246acf6558e0984

Observation f65ed111-c9b3-4c86-bf3e-cc00fed2f664 · outbound

This paper cites Compute Or Load KV Cache? Why Not Both?.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Compute Or Load KV Cache? Why Not Both?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T04:42:24.617477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:42:24.617477Z digest=sha256:91446f76c34192c43e0f27d6ee970f5252c5f2fac3e56c2a33c6d72321f58d87

Observation 92a87986-2028-40cb-b234-198314006b62 · outbound

This paper cites s3: Increasing gpu utilization during generative inference for higher throughput.Advances in Neural Information Processing Systems, 36:18015–18027, 2023.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving s3: Increasing gpu utilization during generative inference for higher throughput.Advances in Neural Information Processing Systems, 36:18015–18027, 2023

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:42:25.340557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-14T04:42:24.627354Z digest=sha256:6f6f8a365ce68400e32f096574ecdfff2ac2fb9c5d03e665a563f0702ae8d208

Observation 7687dbd3-5195-445c-a1a1-5564dab4481d · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T04:42:24.655668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:42:24.655668Z digest=sha256:828fc145cbfcd1c000bed6df726d8b53ed9f16c7222515a6bf2db743b44ea3b3

Observation bcffb26d-7b89-4587-9b70-57120cafd825 · outbound

This paper cites Cgroups.Available on-line at: http://www.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Cgroups.Available on-line at: http://www

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:42:25.294760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-14T04:42:24.660374Z digest=sha256:d6b5920b88807444d2e46de01d8f7f274334e3fa05728ec489f4e63d0f0f46e3

Observation d6eaa769-159d-467f-b7b3-6e240984dc41 · outbound

This paper cites Docker: lightweight linux containers for consistent development and deployment.Linux j, 239(2):2, 2014.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Docker: lightweight linux containers for consistent development and deployment.Linux j, 239(2):2, 2014

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:42:25.239920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-14T04:42:24.664605Z digest=sha256:f9bae9a4d1b7421731841c57a7806353fcd827144ecc79ca8b9e88a0efa0fcc4

Observation 8c9f05d0-ed81-4908-a366-f7cc6d4c337d · outbound

This paper cites Introducing llama 3.1: The next generation of open models.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Introducing llama 3.1: The next generation of open models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:42:25.174750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-14T04:42:24.668261Z digest=sha256:b1366bd0dac41c2b891be50763f0436473f22d530f716409a9e49bcb62b9ee7b

Observation d9f7c4b6-7752-4e4b-936e-b4bbd6d43ca0 · outbound

This paper cites NVIDIA Multi-Instance GPU (MIG) User Guide.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving NVIDIA Multi-Instance GPU (MIG) User Guide

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:42:25.134749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-14T04:42:24.671996Z digest=sha256:acc5ee8b3be4a70b5906d39532f5fd130fc527f4596f471c8c8db297ebcbfb02

Observation ea971a8b-95b7-4af9-993e-95a112d112c3 · outbound

This paper cites OpenAI API.https://openai.com/api/, 2025.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving OpenAI API.https://openai.com/api/, 2025

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:42:25.080487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-14T04:42:24.676161Z digest=sha256:86143cffd85d2b45ac71315898c83f601f2895ce9f125c8f246cb7dadb5650ac

Observation d9a830fd-438d-4493-a156-12e5b7c64d9f · outbound

This paper cites Fairness in serving large language models.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Fairness in serving large language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:42:25.059925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-14T04:42:24.680943Z digest=sha256:1263ad8eddd5f256a6db74d0b29b00c88d0918c1aa26adc3e0844fe0eb87e1b9

Observation 0ea71899-891b-4a5c-90a5-00c332b866c0 · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Qwen2.5: A party of foundation models, September 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T04:42:24.685389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:42:24.685389Z digest=sha256:06b6e646fde387ed377124293c2957b836168f8d20b34089558c056528073052

Observation 242832dd-cf0d-41fc-8950-beaf09c15ab0 · outbound

This paper cites AIBrix: Towards Scalable, Cost-Effective Large Language Model Inference Infrastructure.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving AIBrix: Towards Scalable, Cost-Effective Large Language Model Inference Infrastructure

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-14T04:42:24.689292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:42:24.689292Z digest=sha256:649861098a7567a1599224fb6fe0a51a023997633c2201d12590d928e3ecee7a

Observation 4bbb7180-67a2-4bba-8b33-34fd237eb44a · outbound

This paper cites HeterMoE: Efficient Training of Mixture-of-Experts Models on Heterogeneous GPUs.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving HeterMoE: Efficient Training of Mixture-of-Experts Models on Heterogeneous GPUs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-14T04:42:24.697781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:42:24.697781Z digest=sha256:a78bf85b5a26be91961fa23d794fa6f09e9b82ae8adee1888b81729ff7ff8a5f

Observation 884f6640-65be-4964-9c1f-d5909c70f64a · outbound

This paper cites RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-14T04:42:24.724753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:42:24.724753Z digest=sha256:dcf0db06584e518bc58fc622e0f21704d64e924a108218859bf891b4078e92d6

Observation 768fbee0-c871-44cd-9440-f7b38a176288 · outbound

This paper cites LLM Inference Unveiled: Survey and Roofline Model Insights.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-14T04:42:24.748447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:42:24.748447Z digest=sha256:40296ea3ccf86dc442d1f6ae0f66ccf9e603571bbf56184b0f800ecd02efb9ba

Observation 759bacf0-3e8e-4c02-b100-4947e78a16ef · outbound

This paper cites Eagle: Efficient Training-Free Router for Multi-LLM Inference.

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Eagle: Efficient Training-Free Router for Multi-LLM Inference

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-14T04:42:24.754153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:42:24.754153Z digest=sha256:4f73ed6232b1f1ef7d67f06760612d1b12b955da2bd0306ef71ab5cfcb525d2a

Pith citing papers

No inbound Pith citation observations are available.