Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T05:50:05.925477Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 0 inbound Pith citation observations for arXiv:2608.13499.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T05:50:05.925477Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
71 of 71 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d5b0165d-3fde-44d8-a308-1c4a95bd7cbf · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5f6fc802-b4cf-4773-a999-bf8eac2017e3 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Medha: Efficiently serving multi-million context length LLM inference requests without approximations.arXiv preprint arXiv:2409.17264, 2024
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd6e8b85-6033-4dbc-982f-e3bb0f6b1970 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Look Ma, No Bubbles! Designing a Low-Latency Megakernel for Llama-1B
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 16cda042-bf71-4f5f-b21b-40049ac7ae63 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Internet and the Erlang formula.ACM SIGCOMM Computer Communication Review, 42(1):23–30, 2012
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ad8f9766-7a47-4bef-b45e-a748f1e1e46b · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Stability, queue length, and delay of deterministic and stochastic queueing networks
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 68dbfb04-249d-4ec9-b604-d4cd9f6d3090 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving TVM: An automated end-to-end optimizing compiler for deep learning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0fefdded-8a8b-4e27-905e-e3500f08e17f · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Towards high-goodput LLM serving with prefill-decode multiplexing
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2baf0e86-eecd-4c90-afb2-57ff39e6db10 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Mirage Persistent Kernel: A Compiler and Runtime for Mega-Kernelizing Tensor Programs.arXiv preprint arXiv:2512.22219, 2025
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cd6590e-3aac-4d18-a443-5bb3e3cab5f5 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Serving heterogeneous machine learning models on multi-GPU servers with spatio-temporal sharing
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0d62dfb8-90d2-48b0-aaf8-8cff978ee44e · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving PaLM: Scaling language modeling with Pathways.Journal of Machine Learning Research, 24(240):1–113, 2023
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6c5e8fba-a593-4d5d-9b4c-e2f2e464a0a4 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving LithOS: An operating system for efficient machine learning on GPUs
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a9f31d62-9979-44a0-91d1-1825d6ad419a · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving FlashAttention: Fast and memory-efficient exact attention with IO-awareness
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 95efc589-7ff8-4255-9460-c32d9e07069c · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving GSLICE: controlled spatial sharing of GPUs for a scalable inference platform
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0a27caf4-6d0b-403a-b2e3-b570a1d146db · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving HydraInfer: Hybrid disaggregated scheduling for multimodal large language model serving.arXiv preprint arXiv:2505.12658, 2025
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c21e7f3-3e67-45ea-b50e-a33072e1a513 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6642cb48-6559-4fd3-8ae6-560efea4baf9 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving ServerlessLLM:low-latency serverless inference for large language models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4294a411-ecdb-46f5-a779-cc993332638b · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving ATOM: Model-driven autoscaling for microservices
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 938876dd-505b-4dd3-99e0-3fe6c9225c4c · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Nano-vLLM.https: //github.com/GeeeekExplorer/nano-vllm, 2025
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f60a81f9-0807-462d-9bb5-cb62b20f2324 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving NVIDIA Dynamo
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 43697b43-a2b8-49ae-92f0-db0b84ad1d31 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving vLLM Production Stack.https: //github.com/vllm-project/production-stack, 2025
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3c53029e-3559-4473-98af-3781038d5753 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28ec9bbf-799b-4bea-917e-ed09ae1cef7d · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving DEEPSERVE: Serverless large language model serving at scale
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b65cc737-39d8-4e89-85c3-5b3c43cb3795 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving DDiT: Dynamic Resource Allocation for Diffusion Transformer Model Serving
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 499f5144-81ba-401c-b2fa-2ef8e3a065e2 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving In 2026 IEEE International Symposium on High Performance Computer Architecture (HPCA 2026), pages 1–14
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 29fa7ea1-7334-464c-a2f2-b9314d28960a · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Llama-3-8b.https://huggingface
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 24b94bb3-2292-4c52-a28c-9c24e6a8b3bf · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Mixtral-8x7B-v0.1.https:// huggingface.co/mistralai/Mixtral-8x7B-v0.1, 2025
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e672d2cb-277e-4f5b-b8ef-9be6f7661577 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Qwen2-57B-A14B
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1b2ccc7f-1231-4695-855f-f2444195f0b6 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving QWen2-7B-Instruct.https: //huggingface.co/Qwen/Qwen2-7B-Instruct, 2025
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 771dd334-511b-4a90-b6e3-134abeed11a4 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Qwen2.5-VL-32B
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 924bdd43-ab67-4ad1-a3e8-8567fbd67efc · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Amant, Chetan Bansal, Victor Ruhle, Anoop Kulkarni, Steve Kofsky, and Saravan Rajmohan
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1be6b991-b591-4ecf-9a07-3c83f133207c · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Pod-Attention: Unlocking full prefill-decode overlap for faster LLM inference
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9df997e1-05d3-4d02-b715-96b58486c268 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving A simulation analysis of sojourn times in a Jackson network
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0933a7e9-53cd-46fa-94a9-99b1c2001f57 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Horizontal Pod Autoscaling.http: //kubernetes.io/docs/concepts/workloads/ autoscaling/horizontal-pod-autoscale, 2026
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation cf63d73f-90b7-4d69-b425-1e235e89558a · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving AlpaServe: Statistical multiplexing with model parallelism for deep learning serving
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 40f2fa55-c712-490a-90a2-19b5dfa05470 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Bullet: Boosting GPU utilization for LLM serving via dynamic spatial-temporal orchestration.arXiv preprint arXiv:2504.19516, 2025
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 596d5588-e44b-4d1a-aa4e-c4eabcf73a61 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Expert-as-a-service: Towards efficient, scalable, and robust large-scale MoE serving.arXiv preprint arXiv:2509.17863, 2025
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1dbf675-28a5-47e3-b64b-87848916c88f · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Azure VM NDm-A100-v4 sizes series.https://learn.microsoft.com/en-us/ azure/virtual-machines/sizes/ gpu-accelerated/ndma100v4-series, 2024
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation efb460e8-7530-4193-b535-370850ecdbc9 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Azure VM ND GB200-v6 sizes series.https://learn.microsoft.com/en-us/ azure/virtual-machines/sizes/ gpu-accelerated/nd-gb200-v6-series, 2026
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3641b40c-e348-4ac7-8ca9-790f26fa96d8 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Documentation on NVIDIA Multi-Instance GPU (MIG).https://www.nvidia.com/en-us/ technologies/multi-instance-gpu/, 2025
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e86991f8-af73-4fe2-a167-f491c9dcca21 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Documentation on NVIDIA Multi-Process Service (MPS).https: //docs.nvidia.com/deploy/mps/index.html, 2025
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0bd48cb5-f5fc-42b6-8193-4375d847a04f · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Nsight Systems.https: //developer.nvidia.com/nsight-systems, 2025
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e116b5e0-47e3-4200-a5de-75bc36638341 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving NVIDIA DCGM
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d6a147d4-c168-4e68-9113-0e72f68209e9 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving NVIDIA Green Context Documentation
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8961f5ca-e16c-4bdc-831e-0c389ce7d350 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Introducing ChatGPT
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 224920e1-81ed-4248-a34b-74b856d9128d · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving ChatGPT Codex
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation baacd22e-c7a1-4486-845d-de686a0bc8be · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Introducing Deep Research
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3837b475-2097-41fa-ae2e-c81018b4ab17 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Measuring Agents in Production
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b9ac719-4b3d-4f6a-b323-812a45d49b10 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Splitwise: Efficient generative LLM inference using phase splitting
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2fbffbea-b1bd-4ee0-9093-4e8370d75c5d · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Hierarchical Autoscaling for Large Language Model Serving with Chiron
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3341333c-74d4-42ea-a23c-6251c0123749 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Gonzalez, Ion Stoica, and Harry Xu
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d3859c9b-3093-4662-b8c0-68de115c3106 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bbe73cc-c469-4f6d-908c-af50bd5574b0 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving FIRM: An intelligent fine-grained resource management framework for SLO-oriented microservices
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c5fb8028-6a9e-41e4-89fd-e945f150bfc9 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving ModServe: Scalable and resource-efficient large multimodal model serving.arXiv preprint arXiv:2502.00937, 2025
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e409514e-9414-42e2-a094-11fa45ee5b8b · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Power-aware deep learning model serving with µ-Serve
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation dc9df5f0-ac90-482d-aef0-696a9fcb3403 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving USHER: Holistic interference avoidance for resource optimized ML inference
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c9ad68a2-064b-453d-9a35-568e673923f5 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Efficiently Serving Large Multimodal Models Using EPD Disaggregation
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c846edf-bf77-4054-80c6-90ea870afdf4 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving DynamoLLM: Designing LLM inference clusters for performance and energy efficiency
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 434b7763-0b9d-44ff-bb37-b8b3c97e6e35 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Orion: Interference-aware, fine-grained GPU sharing for ML applications
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 005c611f-6144-43ca-b66f-cb6d1b47f03d · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving AIBrix: Towards Scalable, Cost-Effective Large Language Model Inference Infrastructure
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31a91211-922b-4548-aba9-dd83a56be30a · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Distributed Inference and Serving
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e08d7657-36b4-4a3c-8e1a-2631e020b52e · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving vLLM Profiler.https://docs.vllm.ai/en/ stable/contributing/profiling/, 2025
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3f41d75b-28c2-43c8-920c-4d854305b720 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Step-3 is large yet affordable: Model-system co-design for cost-effective decoding.arXiv preprint arXiv:2507.19427, 2025
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cf61f99-b599-4973-bf5f-447bf464003c · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Autothrottle: A practical bi-level approach to resource management for SLO-targeted microservices
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c200162e-ce8a-440c-86a6-587522659815 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving DeepScaling: microservices autoscaling for stable cpu utilization in large scale cloud systems
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 623bbe61-556d-4498-be9c-999e8dba87dc · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Aegaeon: Effective GPU pooling for concurrent LLM serving on the market
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b6bf6a03-1957-45cf-af65-9246a48b9bd8 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Towards Efficient and Practical GPU Multitasking in the Era of LLM
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7174262-3b1c-40b6-9d86-119ebe8cdb6e · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Prism: Cost-Efficient Multi-LLM Serving via GPU Memory Ballooning
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c956474c-8c6b-4762-b6d5-434fa09e7650 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e268cb21-855f-4aa2-b214-de1acc1dcedf · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving NanoFlow: Towards optimal large language model serving throughput
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6fd1a60d-da8f-4b10-9a6f-d793d5a53838 · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e3af8a0-484b-4851-98c2-04187c6ab46c · outbound
OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Serving Large Language Models on Huawei CloudMatrix384
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.