Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2404.08509.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T05:49:10.079027Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T08:39:41.968159Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation d407896f-c54b-482a-ba84-0ef671fcfe52 · inbound
TimelyLLM: Segmented LLM Serving System for Time-sensitive Robotic Applications Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e03dc80-95e4-4b93-a4f1-f6d587f7711a · inbound
KVDirect: Distributed Disaggregated LLM Inference Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af5cfa02-45bd-4751-ba5f-ae51ac181353 · inbound
Taming the Titans: A Survey of Efficient LLM Inference Serving Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 580c0349-75b3-49e0-9e6f-96378365680c · inbound
Efficient Serving of LLM Applications with Probabilistic Demand Modeling Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7831e371-0d79-4ee5-8c34-0a809bbc6b12 · inbound
Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f8399dc-ff72-4234-bc1e-d302dff8fba4 · inbound
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f82b0bb-98e2-4a74-bc9e-00fef6d80ae5 · inbound
Make a Video Call with LLM: A Measurement Campaign over Six Mainstream Apps Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d527a7b-b297-408f-9e99-36cf7edf74f4 · inbound
STAR: Decode-Phase Rescheduling for LLM Inference Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5bf92d1d-3317-4373-90da-738cec23475a · inbound
CascadeInfer: Length-Aware Scheduling of LLM Serving with Low Latency and Load Balancing Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b887a3d0-c840-4cb7-8a41-1053673f8a5d · inbound
SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 480d42c1-e30b-4030-9696-52950dec518b · inbound
Parallelizing Tool Execution and LLM Generation for Low-Latency Agent Serving Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e2d4b04-a066-4c30-8278-ae2f3481830f · inbound
Autopoiesis: A Self-Evolving System Paradigm for LLM Serving Under Runtime Dynamics Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 194f3840-2a3d-4c94-acab-8328da2be2c4 · inbound
Robust Length Prediction: A Perspective from Heavy-Tailed Prompt-Conditioned Distributions Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ebf300e4-9ae2-4898-bf40-9b833550ff0d · inbound
A Queueing-Theoretic Framework for Stability Analysis of LLM Inference with KV Cache Memory Constraints Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 087b1463-2ff7-4def-a36b-49e8544a58a0 · inbound
Tackling the Data-Parallel Load Balancing Bottleneck in LLM Serving: Practical Online Routing at Scale Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation cedacd78-8a15-45a4-937f-9b3b6698bff6 · inbound
Tackling the Data-Parallel Load Balancing Bottleneck in LLM Serving: Practical Online Routing at Scale Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2d07ea33-387c-4801-b984-d4ca49b5a361 · inbound
Clairvoyant: Predictive Shortest-Job-First Admission for Serial LLM Inference Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7212223b-05ec-4998-891a-2444fca7ee67 · inbound
Clairvoyant: Predictive Shortest-Job-First Admission for Serial LLM Inference Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b8bedb0-2a8d-408a-97ba-dfba847d6ee1 · inbound
Beyond Prediction: Tail-Aware Scheduling for LLM Inference Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 98c7d8f5-f289-4f2c-902f-337f9da42e53 · inbound
Geometry-Aware Online Scheduling for LLM Serving: From Theoretical Bound to System Practice Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 283d4440-2c2a-4bf6-8fd9-aeb4d21e7e00 · inbound
Energy-Aware Scheduling for Serverless LLM Serving on Shared GPUs Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 144f9d3b-2a0f-4c9f-ae40-97a4e2c6276b · inbound
Online Linear Programming for Multi-Objective Routing in LLM Serving Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6074f733-8b52-4ad7-8e45-7b659fd43a4f · inbound
General Non-Clairvoyant KV-Cache Scheduling via Regime-Aware Routing Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42c23f12-c4f6-477f-81c8-b3cba8af8082 · inbound
Auto-Scaling Heterogeneous Neural Processing Units for Energy and Cost-Efficient LLM Serving Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be2c78d2-20ee-4db1-91b9-59affbe87787 · inbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 278d6a21-d664-440c-900d-9e99ab57a0c5 · inbound
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.