Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T04:42:24.754153Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2608.08382.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T04:42:24.754153Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
23 of 23 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d9a16fe4-8913-40d9-9507-fc9a7cbf5576 · outbound
LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Amazon EC2 Instance Types – Burstable Performance Instances (T3)
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b16004a4-1535-4515-8710-48f1cd614722 · outbound
LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Anthropic API.https://www.anthropic.com/api, 2025
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e963fe0b-4c44-4bec-8014-3b1f40675fd7 · outbound
LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Xen and the art of virtualization.ACM SIGOPS operating systems review, 37(5):164–177, 2003
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation de6244fb-1abf-4d20-b503-d4a1b53c0793 · outbound
LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving On the Opportunities and Risks of Foundation Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98058cc8-0fe4-494f-8816-2ade85e32230 · outbound
LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving GPT-4o System Card
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed21b4eb-dfbf-493c-b345-40da02c721f2 · outbound
LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Predicting llm inference latency: A roofline-driven ml method
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 41ddca52-034c-4ff7-beaf-bc799f2b5df1 · outbound
LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03a6f21d-fa99-4fab-ab9d-a605b9afa755 · outbound
LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Plato: Plan to efficient decode for large language model inference
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f65ed111-c9b3-4c86-bf3e-cc00fed2f664 · outbound
LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Compute Or Load KV Cache? Why Not Both?
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92a87986-2028-40cb-b234-198314006b62 · outbound
LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving s3: Increasing gpu utilization during generative inference for higher throughput.Advances in Neural Information Processing Systems, 36:18015–18027, 2023
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7687dbd3-5195-445c-a1a1-5564dab4481d · outbound
LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcffb26d-7b89-4587-9b70-57120cafd825 · outbound
LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Cgroups.Available on-line at: http://www
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d6eaa769-159d-467f-b7b3-6e240984dc41 · outbound
LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Docker: lightweight linux containers for consistent development and deployment.Linux j, 239(2):2, 2014
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8c9f05d0-ed81-4908-a366-f7cc6d4c337d · outbound
LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Introducing llama 3.1: The next generation of open models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d9f7c4b6-7752-4e4b-936e-b4bbd6d43ca0 · outbound
LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving NVIDIA Multi-Instance GPU (MIG) User Guide
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ea971a8b-95b7-4af9-993e-95a112d112c3 · outbound
LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving OpenAI API.https://openai.com/api/, 2025
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d9a830fd-438d-4493-a156-12e5b7c64d9f · outbound
LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Fairness in serving large language models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0ea71899-891b-4a5c-90a5-00c332b866c0 · outbound
LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Qwen2.5: A party of foundation models, September 2024
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 242832dd-cf0d-41fc-8950-beaf09c15ab0 · outbound
LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving AIBrix: Towards Scalable, Cost-Effective Large Language Model Inference Infrastructure
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bbb7180-67a2-4bba-8b33-34fd237eb44a · outbound
LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving HeterMoE: Efficient Training of Mixture-of-Experts Models on Heterogeneous GPUs
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 884f6640-65be-4964-9c1f-d5909c70f64a · outbound
LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 768fbee0-c871-44cd-9440-f7b38a176288 · outbound
LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving LLM Inference Unveiled: Survey and Roofline Model Insights
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 759bacf0-3e8e-4c02-b100-4947e78a16ef · outbound
LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving Eagle: Efficient Training-Free Router for Multi-LLM Inference
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.