Pith. sign in

Paper Citation Record · LEDGER

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving

As of 15 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 0 inbound Pith citation observations for arXiv:2608.13499.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.13499 v1

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T05:50:05.925477Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

71 of 71 outbound references displayed

  • verified exact0
  • verified fuzzy52
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d5b0165d-3fde-44d8-a308-1c4a95bd7cbf · outbound

This paper cites Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:07.558975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.568677Z digest=sha256:c1e2f12eab369d9ff32e5b2337719eb8b0909a8193703e147599bc861dc23810

Observation 5f6fc802-b4cf-4773-a999-bf8eac2017e3 · outbound

This paper cites Medha: Efficiently serving multi-million context length LLM inference requests without approximations.arXiv preprint arXiv:2409.17264, 2024.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Medha: Efficiently serving multi-million context length LLM inference requests without approximations.arXiv preprint arXiv:2409.17264, 2024

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-14T05:50:05.574731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T05:50:05.574731Z digest=sha256:59702b68ddca69cd0c9f61459b9636f63a6e03ebb1624d6c4bc65d89055e8dc3

Observation cd6e8b85-6033-4dbc-982f-e3bb0f6b1970 · outbound

This paper cites Look Ma, No Bubbles! Designing a Low-Latency Megakernel for Llama-1B.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Look Ma, No Bubbles! Designing a Low-Latency Megakernel for Llama-1B

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:07.542109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.580018Z digest=sha256:a9ab9ef3c89377efb10c709aca7d00f0d68f20739651e56da04b68566f678832

Observation 16cda042-bf71-4f5f-b21b-40049ac7ae63 · outbound

This paper cites Internet and the Erlang formula.ACM SIGCOMM Computer Communication Review, 42(1):23–30, 2012.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Internet and the Erlang formula.ACM SIGCOMM Computer Communication Review, 42(1):23–30, 2012

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:07.525517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.585184Z digest=sha256:05a54311fe73c20d7c748b3a32ac0a6fca5aaaa57c5398576d5f1f945443fe37

Observation ad8f9766-7a47-4bef-b45e-a748f1e1e46b · outbound

This paper cites Stability, queue length, and delay of deterministic and stochastic queueing networks.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Stability, queue length, and delay of deterministic and stochastic queueing networks

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:07.509278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.590364Z digest=sha256:3a0a61cb111a5d2e488d797be13ee107161f11561ab8d476bd7759ca52919e0f

Observation 68dbfb04-249d-4ec9-b604-d4cd9f6d3090 · outbound

This paper cites TVM: An automated end-to-end optimizing compiler for deep learning.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving TVM: An automated end-to-end optimizing compiler for deep learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:07.493059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.595384Z digest=sha256:249f1e174e64fb771abff1c33da7a9a3ba03288aa9cb3944cb3b7a9a9f977798

Observation 0fefdded-8a8b-4e27-905e-e3500f08e17f · outbound

This paper cites Towards high-goodput LLM serving with prefill-decode multiplexing.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Towards high-goodput LLM serving with prefill-decode multiplexing

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:07.475563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.601314Z digest=sha256:415829f83c4c2d625a28bb68cd55337d2a729f827fe809dfa76ca758f6d6ef17

Observation 2baf0e86-eecd-4c90-afb2-57ff39e6db10 · outbound

This paper cites Mirage Persistent Kernel: A Compiler and Runtime for Mega-Kernelizing Tensor Programs.arXiv preprint arXiv:2512.22219, 2025.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Mirage Persistent Kernel: A Compiler and Runtime for Mega-Kernelizing Tensor Programs.arXiv preprint arXiv:2512.22219, 2025

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T05:50:05.607051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T05:50:05.607051Z digest=sha256:45a9e648a6a2a63325a3863de53819f4b3ca33ebb91aad3fa190a88a88fee563

Observation 1cd6590e-3aac-4d18-a443-5bb3e3cab5f5 · outbound

This paper cites Serving heterogeneous machine learning models on multi-GPU servers with spatio-temporal sharing.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Serving heterogeneous machine learning models on multi-GPU servers with spatio-temporal sharing

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:07.460293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.612248Z digest=sha256:240232a2b4ded4872064c6314084f831e542bfec0f488072cf4c7dd466d8ef0f

Observation 0d62dfb8-90d2-48b0-aaf8-8cff978ee44e · outbound

This paper cites PaLM: Scaling language modeling with Pathways.Journal of Machine Learning Research, 24(240):1–113, 2023.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving PaLM: Scaling language modeling with Pathways.Journal of Machine Learning Research, 24(240):1–113, 2023

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:07.443604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.617457Z digest=sha256:27600cbb3c0e7d5da322fa86838599f5e6c40dd38c801ed04a20c68f5bf3ac22

Observation 6c5e8fba-a593-4d5d-9b4c-e2f2e464a0a4 · outbound

This paper cites LithOS: An operating system for efficient machine learning on GPUs.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving LithOS: An operating system for efficient machine learning on GPUs

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:07.428119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.622750Z digest=sha256:f17a3e6c2284a0a85f959a1deb70563a19c76f8589ca355e359647ad531e8129

Observation a9f31d62-9979-44a0-91d1-1825d6ad419a · outbound

This paper cites FlashAttention: Fast and memory-efficient exact attention with IO-awareness.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving FlashAttention: Fast and memory-efficient exact attention with IO-awareness

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:07.413635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.628242Z digest=sha256:7e60a0ca33e165499c16698ca78b8f858dda39dbc1886e87b00f150a8cff4d11

Observation 95efc589-7ff8-4255-9460-c32d9e07069c · outbound

This paper cites GSLICE: controlled spatial sharing of GPUs for a scalable inference platform.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving GSLICE: controlled spatial sharing of GPUs for a scalable inference platform

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:07.396879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.633194Z digest=sha256:fb6fb48aa77c8670855169d8dad145e889f93873203972a412d0267e49f54979

Observation 0a27caf4-6d0b-403a-b2e3-b570a1d146db · outbound

This paper cites HydraInfer: Hybrid disaggregated scheduling for multimodal large language model serving.arXiv preprint arXiv:2505.12658, 2025.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving HydraInfer: Hybrid disaggregated scheduling for multimodal large language model serving.arXiv preprint arXiv:2505.12658, 2025

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T05:50:05.638466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T05:50:05.638466Z digest=sha256:9d868de43ff6d12aabe4f84a0f673e822592b6274ee94b1b994daee14548f165

Observation 3c21e7f3-3e67-45ea-b50e-a33072e1a513 · outbound

This paper cites MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T05:50:05.643366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T05:50:05.643366Z digest=sha256:0f05c2862b43402d681262f12e3905e1dd229c7e73a0a03cf77518aaf95d150f

Observation 6642cb48-6559-4fd3-8ae6-560efea4baf9 · outbound

This paper cites ServerlessLLM:low-latency serverless inference for large language models.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving ServerlessLLM:low-latency serverless inference for large language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:07.381562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.648448Z digest=sha256:d27e579db90e7df4765064c57c421f6d49380fe9a58893273b190e7f8a7e6cc7

Observation 4294a411-ecdb-46f5-a779-cc993332638b · outbound

This paper cites ATOM: Model-driven autoscaling for microservices.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving ATOM: Model-driven autoscaling for microservices

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:07.365712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.653791Z digest=sha256:8838e3ad0cdbaca44f0ff737b2259e1ee1475904280b4ccf7c696a04ab49bcad

Observation 938876dd-505b-4dd3-99e0-3fe6c9225c4c · outbound

This paper cites Nano-vLLM.https: //github.com/GeeeekExplorer/nano-vllm, 2025.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Nano-vLLM.https: //github.com/GeeeekExplorer/nano-vllm, 2025

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:07.349857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.658523Z digest=sha256:7f77c97774053f6072fef30e2a87b6e45fa5027cc041c6963cafa7ea4bf3af08

Observation f60a81f9-0807-462d-9bb5-cb62b20f2324 · outbound

This paper cites NVIDIA Dynamo.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving NVIDIA Dynamo

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:07.334485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.664438Z digest=sha256:6af784080ed30b34268b7b73022d092e45e97a40eed1da5b07a3039f31b2fc8b

Observation 43697b43-a2b8-49ae-92f0-db0b84ad1d31 · outbound

This paper cites vLLM Production Stack.https: //github.com/vllm-project/production-stack, 2025.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving vLLM Production Stack.https: //github.com/vllm-project/production-stack, 2025

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:07.315794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.670158Z digest=sha256:4b3a71651ea29cd5988b1730329c3c189047a727f2fbc01fe5c641d45de0946d

Observation 3c53029e-3559-4473-98af-3781038d5753 · outbound

This paper cites semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-14T05:50:05.675450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T05:50:05.675450Z digest=sha256:e34a37061ae98e5d1fb6f2f4fd517e4f2f413201d886ea97aa4f8027b5fcb552

Observation 28ec9bbf-799b-4bea-917e-ed09ae1cef7d · outbound

This paper cites DEEPSERVE: Serverless large language model serving at scale.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving DEEPSERVE: Serverless large language model serving at scale

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:07.298852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.680591Z digest=sha256:0e76d23186676a3e7c47a23c5b2d06c8aa602938795f21919d3427ec775aebf3

Observation b65cc737-39d8-4e89-85c3-5b3c43cb3795 · outbound

This paper cites DDiT: Dynamic Resource Allocation for Diffusion Transformer Model Serving.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving DDiT: Dynamic Resource Allocation for Diffusion Transformer Model Serving

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-14T05:50:05.685350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T05:50:05.685350Z digest=sha256:86a56164d36d25be7d4586de8bf1c0a0f0b0df8213eb2ba51996d71060757769

Observation 499f5144-81ba-401c-b2fa-2ef8e3a065e2 · outbound

This paper cites In 2026 IEEE International Symposium on High Performance Computer Architecture (HPCA 2026), pages 1–14.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving In 2026 IEEE International Symposium on High Performance Computer Architecture (HPCA 2026), pages 1–14

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:07.282903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.691663Z digest=sha256:77111944d7a2d0974899271cc59814096d5bf0cf1a930e7d7608d94e805befb2

Observation 29fa7ea1-7334-464c-a2f2-b9314d28960a · outbound

This paper cites Llama-3-8b.https://huggingface.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Llama-3-8b.https://huggingface

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:07.265449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.696322Z digest=sha256:8cebafbaf693b4105d38d471b366fd8c05c23ca67680f687dcc866095f052cf8

Observation 24b94bb3-2292-4c52-a28c-9c24e6a8b3bf · outbound

This paper cites Mixtral-8x7B-v0.1.https:// huggingface.co/mistralai/Mixtral-8x7B-v0.1, 2025.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Mixtral-8x7B-v0.1.https:// huggingface.co/mistralai/Mixtral-8x7B-v0.1, 2025

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:07.246809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.701113Z digest=sha256:3da2051b65e12fe1eb5719f07e263adea3c42d1515c788ba8bb3070317bb7007

Observation e672d2cb-277e-4f5b-b8ef-9be6f7661577 · outbound

This paper cites Qwen2-57B-A14B.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Qwen2-57B-A14B

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:07.230407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.706060Z digest=sha256:ccab1ee70c16fcc8419a3bf4f2f6620d6f6c6ad21333390809e18eaf1a5946fd

Observation 1b2ccc7f-1231-4695-855f-f2444195f0b6 · outbound

This paper cites QWen2-7B-Instruct.https: //huggingface.co/Qwen/Qwen2-7B-Instruct, 2025.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving QWen2-7B-Instruct.https: //huggingface.co/Qwen/Qwen2-7B-Instruct, 2025

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:07.214383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.710659Z digest=sha256:ac1bc787281f7ef7389136f15a6d575a95346887954d1f5fb92bbcc56ce10fae

Observation 771dd334-511b-4a90-b6e3-134abeed11a4 · outbound

This paper cites Qwen2.5-VL-32B.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Qwen2.5-VL-32B

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:07.200036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.715410Z digest=sha256:4f07bb30d21072de000f81e4d6c6384d0e0e4b8da16571e29f64012a1a0e6a3c

Observation 924bdd43-ab67-4ad1-a3e8-8567fbd67efc · outbound

This paper cites Amant, Chetan Bansal, Victor Ruhle, Anoop Kulkarni, Steve Kofsky, and Saravan Rajmohan.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Amant, Chetan Bansal, Victor Ruhle, Anoop Kulkarni, Steve Kofsky, and Saravan Rajmohan

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:07.185431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.720125Z digest=sha256:4605dc1b5b72602d87a3d1684673ad7b6c0fa9e96ce69b2a5682d839fad0fedb

Observation 1be6b991-b591-4ecf-9a07-3c83f133207c · outbound

This paper cites Pod-Attention: Unlocking full prefill-decode overlap for faster LLM inference.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Pod-Attention: Unlocking full prefill-decode overlap for faster LLM inference

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:07.169644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.724751Z digest=sha256:723397e32be2d908ca4ba0c78579b3d5ea751b1cabc8ab7d78d7f82056b4dbeb

Observation 9df997e1-05d3-4d02-b715-96b58486c268 · outbound

This paper cites A simulation analysis of sojourn times in a Jackson network.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving A simulation analysis of sojourn times in a Jackson network

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:07.154096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.729378Z digest=sha256:22fd8b4cf729100df87cadf881f847888fe0e664925e4504ad4c24fdee947870

Observation 0933a7e9-53cd-46fa-94a9-99b1c2001f57 · outbound

This paper cites Horizontal Pod Autoscaling.http: //kubernetes.io/docs/concepts/workloads/ autoscaling/horizontal-pod-autoscale, 2026.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Horizontal Pod Autoscaling.http: //kubernetes.io/docs/concepts/workloads/ autoscaling/horizontal-pod-autoscale, 2026

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:07.136303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.733920Z digest=sha256:622006a1aab3a1c566a6469ec77175c61cefbf90d81aef850c8e34b4fd97db9d

Observation cf63d73f-90b7-4d69-b425-1e235e89558a · outbound

This paper cites AlpaServe: Statistical multiplexing with model parallelism for deep learning serving.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving AlpaServe: Statistical multiplexing with model parallelism for deep learning serving

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:07.120020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.738575Z digest=sha256:78053f87d5f2c69460ba6f08a50f484fd42ae2ba2af452602baa4a5c210792e6

Observation 40f2fa55-c712-490a-90a2-19b5dfa05470 · outbound

This paper cites Bullet: Boosting GPU utilization for LLM serving via dynamic spatial-temporal orchestration.arXiv preprint arXiv:2504.19516, 2025.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Bullet: Boosting GPU utilization for LLM serving via dynamic spatial-temporal orchestration.arXiv preprint arXiv:2504.19516, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-14T05:50:05.743399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T05:50:05.743399Z digest=sha256:d76d92790703104a1d11e36e0571b450d85feb23e90910d865d1ee8c4c948034

Observation 596d5588-e44b-4d1a-aa4e-c4eabcf73a61 · outbound

This paper cites Expert-as-a-service: Towards efficient, scalable, and robust large-scale MoE serving.arXiv preprint arXiv:2509.17863, 2025.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Expert-as-a-service: Towards efficient, scalable, and robust large-scale MoE serving.arXiv preprint arXiv:2509.17863, 2025

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-14T05:50:05.747912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T05:50:05.747912Z digest=sha256:d0efb98b23c093034717a3d77bbbb512257f656d24db0f2a9ceebf02ac4a3681

Observation d1dbf675-28a5-47e3-b64b-87848916c88f · outbound

This paper cites Azure VM NDm-A100-v4 sizes series.https://learn.microsoft.com/en-us/ azure/virtual-machines/sizes/ gpu-accelerated/ndma100v4-series, 2024.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Azure VM NDm-A100-v4 sizes series.https://learn.microsoft.com/en-us/ azure/virtual-machines/sizes/ gpu-accelerated/ndma100v4-series, 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:07.104092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.752956Z digest=sha256:78abf7e3412b39d039284d68215ae68e668980c8e656e38eab3359822faa906a

Observation efb460e8-7530-4193-b535-370850ecdbc9 · outbound

This paper cites Azure VM ND GB200-v6 sizes series.https://learn.microsoft.com/en-us/ azure/virtual-machines/sizes/ gpu-accelerated/nd-gb200-v6-series, 2026.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Azure VM ND GB200-v6 sizes series.https://learn.microsoft.com/en-us/ azure/virtual-machines/sizes/ gpu-accelerated/nd-gb200-v6-series, 2026

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:07.088207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.757849Z digest=sha256:260cae03853c4841a4d1697a46e48314fc6e2cfbb586fbd3373a9b2156b38ac6

Observation 3641b40c-e348-4ac7-8ca9-790f26fa96d8 · outbound

This paper cites Documentation on NVIDIA Multi-Instance GPU (MIG).https://www.nvidia.com/en-us/ technologies/multi-instance-gpu/, 2025.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Documentation on NVIDIA Multi-Instance GPU (MIG).https://www.nvidia.com/en-us/ technologies/multi-instance-gpu/, 2025

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:07.070755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.762812Z digest=sha256:7d4f275b745bef41b0722e02cd0e26d39a8ddda5891721fab0cdcbf2f5a28bd5

Observation e86991f8-af73-4fe2-a167-f491c9dcca21 · outbound

This paper cites Documentation on NVIDIA Multi-Process Service (MPS).https: //docs.nvidia.com/deploy/mps/index.html, 2025.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Documentation on NVIDIA Multi-Process Service (MPS).https: //docs.nvidia.com/deploy/mps/index.html, 2025

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:07.053891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.767872Z digest=sha256:78011e8529cdf77a02a7b0e3ba6242a8b5ccedf418d17adfcd7963a582cbc2c0

Observation 0bd48cb5-f5fc-42b6-8193-4375d847a04f · outbound

This paper cites Nsight Systems.https: //developer.nvidia.com/nsight-systems, 2025.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Nsight Systems.https: //developer.nvidia.com/nsight-systems, 2025

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:07.034860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.773012Z digest=sha256:0c9ed1e41f970206025533b2328a6de147ed29ea875322524d9af819780432d9

Observation e116b5e0-47e3-4200-a5de-75bc36638341 · outbound

This paper cites NVIDIA DCGM.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving NVIDIA DCGM

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:07.016725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.778394Z digest=sha256:d824b526c69626ec98c50e37cd60bea0ec2dcea08c5f4e666bc82f4513e163af

Observation d6a147d4-c168-4e68-9113-0e72f68209e9 · outbound

This paper cites NVIDIA Green Context Documentation.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving NVIDIA Green Context Documentation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:07.001372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.783729Z digest=sha256:248a24d0ec119b2081bde272c680e4ec5f4257617651f0a9114604ae221de457

Observation 8961f5ca-e16c-4bdc-831e-0c389ce7d350 · outbound

This paper cites Introducing ChatGPT.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Introducing ChatGPT

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:06.985505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.788270Z digest=sha256:364efde445acdd045073c826dc2039bf5cd1442786d0f084c5f4b34644fbbbf6

Observation 224920e1-81ed-4248-a34b-74b856d9128d · outbound

This paper cites ChatGPT Codex.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving ChatGPT Codex

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:06.969116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.792789Z digest=sha256:515e15329de976da47a5d3aa3b5db8438a65127947c2271bd120b4fa4271d989

Observation baacd22e-c7a1-4486-845d-de686a0bc8be · outbound

This paper cites Introducing Deep Research.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Introducing Deep Research

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:06.953208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.797378Z digest=sha256:13353653b08baf1a706489be925cc68d139dfb3a5275beafd143a4af96df50ee

Observation 3837b475-2097-41fa-ae2e-c81018b4ab17 · outbound

This paper cites Measuring Agents in Production.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Measuring Agents in Production

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-14T05:50:05.802315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T05:50:05.802315Z digest=sha256:75591db52fc725eda70ef7cd84b28d0725a39cab8896de391cceffd6d4a57030

Observation 8b9ac719-4b3d-4f6a-b323-812a45d49b10 · outbound

This paper cites Splitwise: Efficient generative LLM inference using phase splitting.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Splitwise: Efficient generative LLM inference using phase splitting

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:06.933862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.807259Z digest=sha256:c84b4413d9c58d4438026438d4be788e3f2c2517f047587aca5a6cffc38b66ca

Observation 2fbffbea-b1bd-4ee0-9093-4e8370d75c5d · outbound

This paper cites Hierarchical Autoscaling for Large Language Model Serving with Chiron.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Hierarchical Autoscaling for Large Language Model Serving with Chiron

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-14T05:50:05.811741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T05:50:05.811741Z digest=sha256:6c351acff9b6a6ec421df698da60afa1c118a42128b6af49739aaa830e90b40e

Observation 3341333c-74d4-42ea-a23c-6251c0123749 · outbound

This paper cites Gonzalez, Ion Stoica, and Harry Xu.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Gonzalez, Ion Stoica, and Harry Xu

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:06.913621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.816995Z digest=sha256:dca30b3d756406cff80a65e7044ab9239b1d63d0a1722e77f349f46b7ede21a2

Observation d3859c9b-3093-4662-b8c0-68de115c3106 · outbound

This paper cites Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-14T05:50:05.822059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T05:50:05.822059Z digest=sha256:2ff2367f2dc37f11b900d009b71a54b75d8c07b8bc05dc1b1b67f1dbb32e99d3

Observation 6bbe73cc-c469-4f6d-908c-af50bd5574b0 · outbound

This paper cites FIRM: An intelligent fine-grained resource management framework for SLO-oriented microservices.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving FIRM: An intelligent fine-grained resource management framework for SLO-oriented microservices

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:06.896442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.827601Z digest=sha256:570bac6876636d01b64d97b6400ded0142fb14336bf6654bb08c55de77a4f9b3

Observation c5fb8028-6a9e-41e4-89fd-e945f150bfc9 · outbound

This paper cites ModServe: Scalable and resource-efficient large multimodal model serving.arXiv preprint arXiv:2502.00937, 2025.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving ModServe: Scalable and resource-efficient large multimodal model serving.arXiv preprint arXiv:2502.00937, 2025

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-14T05:50:05.832604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T05:50:05.832604Z digest=sha256:a5532a2fc6ed74995ebadd1e8c2da45d5098520b8f52e3ebb616125a358cef38

Observation e409514e-9414-42e2-a094-11fa45ee5b8b · outbound

This paper cites Power-aware deep learning model serving with µ-Serve.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Power-aware deep learning model serving with µ-Serve

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:06.881336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.837628Z digest=sha256:53db21f168fea744958150adfdba53fb7ceea4b6ef30bbd674c6999f93ad5d12

Observation dc9df5f0-ac90-482d-aef0-696a9fcb3403 · outbound

This paper cites USHER: Holistic interference avoidance for resource optimized ML inference.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving USHER: Holistic interference avoidance for resource optimized ML inference

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:06.865672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.842612Z digest=sha256:79f4de60aef42be928d3f14d681b0d0a911edabf9fca51a945f936ba54649cf0

Observation c9ad68a2-064b-453d-9a35-568e673923f5 · outbound

This paper cites Efficiently Serving Large Multimodal Models Using EPD Disaggregation.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Efficiently Serving Large Multimodal Models Using EPD Disaggregation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-14T05:50:05.847560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T05:50:05.847560Z digest=sha256:d12288b0a1bd9a2a7ce9928cc26e151b2f1abb63210f70944c9ab12be8d7652d

Observation 6c846edf-bf77-4054-80c6-90ea870afdf4 · outbound

This paper cites DynamoLLM: Designing LLM inference clusters for performance and energy efficiency.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving DynamoLLM: Designing LLM inference clusters for performance and energy efficiency

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:06.849384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.852309Z digest=sha256:781e2f7ec3a3bd7c213c3190e7feb306e059788eaa0e6d82e0fe12823598605d

Observation 434b7763-0b9d-44ff-bb37-b8b3c97e6e35 · outbound

This paper cites Orion: Interference-aware, fine-grained GPU sharing for ML applications.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Orion: Interference-aware, fine-grained GPU sharing for ML applications

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:06.831719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.857081Z digest=sha256:c10b6c8057d0a64ea09d6cf46944dd042aea1be8df34cd3fe7d522ff7da86cbb

Observation 005c611f-6144-43ca-b66f-cb6d1b47f03d · outbound

This paper cites AIBrix: Towards Scalable, Cost-Effective Large Language Model Inference Infrastructure.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving AIBrix: Towards Scalable, Cost-Effective Large Language Model Inference Infrastructure

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-14T05:50:05.862038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T05:50:05.862038Z digest=sha256:e1e5a5c07d96b91ff71cd77afb143421f5997ad4f461f21e2892e3f299e6bfef

Observation 31a91211-922b-4548-aba9-dd83a56be30a · outbound

This paper cites Distributed Inference and Serving.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Distributed Inference and Serving

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:06.814067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.866937Z digest=sha256:bf38cde199b63d50075e70b69a9eae1c5add051fbc6a6843f54f1cc17829740d

Observation e08d7657-36b4-4a3c-8e1a-2631e020b52e · outbound

This paper cites vLLM Profiler.https://docs.vllm.ai/en/ stable/contributing/profiling/, 2025.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving vLLM Profiler.https://docs.vllm.ai/en/ stable/contributing/profiling/, 2025

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:06.797238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.873942Z digest=sha256:c46b2689ad8afb859bda877bc8eb9293942101a98d10c3d067438330036ddf71

Observation 3f41d75b-28c2-43c8-920c-4d854305b720 · outbound

This paper cites Step-3 is large yet affordable: Model-system co-design for cost-effective decoding.arXiv preprint arXiv:2507.19427, 2025.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Step-3 is large yet affordable: Model-system co-design for cost-effective decoding.arXiv preprint arXiv:2507.19427, 2025

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-14T05:50:05.879083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T05:50:05.879083Z digest=sha256:22748c283a7b48158d59e4ffbd39047a1910861e39afe1a550860769babc6dd7

Observation 7cf61f99-b599-4973-bf5f-447bf464003c · outbound

This paper cites Autothrottle: A practical bi-level approach to resource management for SLO-targeted microservices.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Autothrottle: A practical bi-level approach to resource management for SLO-targeted microservices

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:06.778567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.884034Z digest=sha256:761dca00e895ce1a0c8a70d389e43aadd1f80f46f12ced29c717427a6bd23f32

Observation c200162e-ce8a-440c-86a6-587522659815 · outbound

This paper cites DeepScaling: microservices autoscaling for stable cpu utilization in large scale cloud systems.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving DeepScaling: microservices autoscaling for stable cpu utilization in large scale cloud systems

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:06.761726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.889694Z digest=sha256:0b7cbd75dc6d2670f66d0d612450b784c9e63075f9e2169e069f28fe83cfa3c7

Observation 623bbe61-556d-4498-be9c-999e8dba87dc · outbound

This paper cites Aegaeon: Effective GPU pooling for concurrent LLM serving on the market.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Aegaeon: Effective GPU pooling for concurrent LLM serving on the market

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:06.745531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.894643Z digest=sha256:db81823dbba81cfd5b560a706d5b11b82fabeac122e099a0601d7200f15f5114

Observation b6bf6a03-1957-45cf-af65-9246a48b9bd8 · outbound

This paper cites Towards Efficient and Practical GPU Multitasking in the Era of LLM.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Towards Efficient and Practical GPU Multitasking in the Era of LLM

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-14T05:50:05.899552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T05:50:05.899552Z digest=sha256:40176612e0ec561dc41ac46229f28fec39c30682d36f5656411555b0b5b2d8ae

Observation d7174262-3b1c-40b6-9d86-119ebe8cdb6e · outbound

This paper cites Prism: Cost-Efficient Multi-LLM Serving via GPU Memory Ballooning.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Prism: Cost-Efficient Multi-LLM Serving via GPU Memory Ballooning

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-14T05:50:05.905030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T05:50:05.905030Z digest=sha256:547b912a50cc49200c98ec080cff36675ed77eea4ebe0213e1232063dfdff675

Observation c956474c-8c6b-4762-b6d5-434fa09e7650 · outbound

This paper cites DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:06.730071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.910551Z digest=sha256:b225a8bf4dc2ad7e1dc8bc41efd6735a136e7ff7d41c96a9a46ac77d1bd476d7

Observation e268cb21-855f-4aa2-b214-de1acc1dcedf · outbound

This paper cites NanoFlow: Towards optimal large language model serving throughput.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving NanoFlow: Towards optimal large language model serving throughput

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T05:50:06.713390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T05:50:05.915430Z digest=sha256:67ce5938a26bafb4d892efdad6f58694ec1f46125fac5b53633d134170bad4bf

Observation 6fd1a60d-da8f-4b10-9a6f-d793d5a53838 · outbound

This paper cites MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-14T05:50:05.920154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T05:50:05.920154Z digest=sha256:8ecdbbac324c24acd34177d95fddc4256f924e62c9ecb91ecaf0ff1655e9d111

Observation 3e3af8a0-484b-4851-98c2-04187c6ab46c · outbound

This paper cites Serving Large Language Models on Huawei CloudMatrix384.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Serving Large Language Models on Huawei CloudMatrix384

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-14T05:50:05.925477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T05:50:05.925477Z digest=sha256:43a1bf75fbddb5c1a8489a49c339ca1c169b3cb8a941bfb1a9f037518b132dfb

Pith citing papers

No inbound Pith citation observations are available.