Pith. sign in

Paper Citation Record · LEDGER

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures

As of 7 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2506.05871.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05871 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:18:24.943745Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T20:35:48.828322Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9a84f048-fc01-47c7-bdd5-71fb2a72c78c · outbound

This paper cites Gqa: Training generalized multi-query transformer models from multi-head checkpoints, 2023.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Gqa: Training generalized multi-query transformer models from multi-head checkpoints, 2023

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:21.975462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:21.975462Z digest=sha256:95f2f110896475353a89595275c953628f4125f355550eb2c318f9944291ba32

Observation b76c699d-04ab-4a89-9df4-bb88f9d7a139 · outbound

This paper cites How continuous batching enables 23x throughput in LLM inference while reducing p50 latency.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures How continuous batching enables 23x throughput in LLM inference while reducing p50 latency

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:32.117041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:22.047032Z digest=sha256:2c3410ce93b78457b680a47ecc0b14d7309078fb5f599c0b647ded15ed4b236a

Observation f116b98f-bafb-43fe-a337-bb60bf1d77e5 · outbound

This paper cites an unresolved cited work.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:18:31.863421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:22.111097Z digest=sha256:a7193cf0c7385ce9e1fc9d8358fa05cdc7729fabb8e911e0cc7f5b2fc2089cd4

Observation d3569a65-42ef-4074-a853-8884f11aae61 · outbound

This paper cites an unresolved cited work.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:22.166883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:22.166883Z digest=sha256:bcda7051ca06622c124389a7e26b888729e5ee3bb069722bf735f6b8edf7962f

Observation 4ffec0d6-76ce-4efd-934c-30698fb7ab4c · outbound

This paper cites Throughput is not all you need: Maximizing goodput in llm serving using prefill-decode disaggregation.https://hao-ai- lab.github.io/blogs/distserve/, 2024.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Throughput is not all you need: Maximizing goodput in llm serving using prefill-decode disaggregation.https://hao-ai- lab.github.io/blogs/distserve/, 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:31.602003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:22.238130Z digest=sha256:bdf5a4e2c406f3e6be746b95db55123b110c02c37aa1273f0479219dbc0b1c8d

Observation 3a7fe4fb-4914-4440-b118-d8163dfad57a · outbound

This paper cites an unresolved cited work.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:18:31.356513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:22.310379Z digest=sha256:a79530fd7bde3ad6bae6ad8170306b4d93a958f19e6a3b83d26cd39907f9f934

Observation f8c79ad7-811b-472b-ab2e-bb55a4b7f564 · outbound

This paper cites an unresolved cited work.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:18:31.052695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:22.376684Z digest=sha256:3c551b9537feb5f32895a129d2c97d03a3c9ccfc822f5b5f7498bc79168ee9fc

Observation 32076b1c-5f97-41ab-ab8d-614d6ea9393c · outbound

This paper cites an unresolved cited work.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:18:30.817164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:22.458150Z digest=sha256:16daa646cbfe956b869d5291c32cfbc4c2bba3a9b6579aeb66115421f33502c2

Observation 24b96557-fb56-44a3-95ff-7b54f2e5471a · outbound

This paper cites Sigmoid-weighted linear units for neural network function approximation in reinforcement learning, 2017.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Sigmoid-weighted linear units for neural network function approximation in reinforcement learning, 2017

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:22.540447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:22.540447Z digest=sha256:03def2814156679e442feec6e1d16ea422d8426095dfe1a63e7e1670b61fa22d

Observation c16b42c3-e6aa-402a-918d-71d2ccaafcb1 · outbound

This paper cites Text generation inference.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Text generation inference

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:30.599741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:22.605821Z digest=sha256:382d0972bb268010eadec240765578fe55a3dfc7d80e7e268fe77939c4713838

Observation 6e795329-54e9-4350-9bd8-073f3c64a412 · outbound

This paper cites Low latency rnn inference with cellular batching.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Low latency rnn inference with cellular batching

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:30.273233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:22.666922Z digest=sha256:6a1c84c8a6451b11f6f3b2c57062c2c6afd88747e5428b13ea16c097f4619e6d

Observation d676dfcc-085d-457a-8234-0659cbf90478 · outbound

This paper cites Getting started with CUDA graphs.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Getting started with CUDA graphs

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:30.035124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:22.721275Z digest=sha256:450f7c808bfe02b682daa03cc021b36049414e8071bc692c6fc0f1b60f08e0d2

Observation 86bb2b9d-0fd9-496c-b9fb-290bcd0678d7 · outbound

This paper cites Shortle, James M.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Shortle, James M

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:29.734840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:22.782207Z digest=sha256:7c0842323e441785460aa9d4e464b650e2d49939e410cc8bcf3bf456f537bf56

Observation 0aee096f-153b-446d-b9b7-33c76b19bbb1 · outbound

This paper cites Pipedream: Fast and efficient pipeline parallel dnn training, 2018.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Pipedream: Fast and efficient pipeline parallel dnn training, 2018

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:29.478952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:22.842703Z digest=sha256:19fde837b0f35c78f97a4f8fcf6585d712250546abb3cc31f92c6c410a371a4a

Observation f32f8b3b-45e4-4fcc-8ad3-536d46c7d48a · outbound

This paper cites Inference without interference: Disaggregate llm inference for mixed downstream workloads, 2024.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Inference without interference: Disaggregate llm inference for mixed downstream workloads, 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:22.918596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:22.918596Z digest=sha256:34addb050f2782aa347d116af7db07285c9c00ccae836475f52df327cae51df3

Observation fc0b7dfd-3473-4204-aaa8-8d8ad4ed900b · outbound

This paper cites Le, Yonghui Wu, and Zhifeng Chen.GPipe: efficient training of giant neural networks using pipeline parallelism.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Le, Yonghui Wu, and Zhifeng Chen.GPipe: efficient training of giant neural networks using pipeline parallelism

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:29.186315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:23.004203Z digest=sha256:9dc17c5a060321484d598a0ee628c610ab1c5d75229dc4370f41067a50d4c423

Observation 6a9ad0c6-582b-4c8c-bf15-b812ea35e28e · outbound

This paper cites an unresolved cited work.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:23.064579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:23.064579Z digest=sha256:c2af41cbaeecc2862dd892bd14ac9b506c5243977fcd7fa8723de5dbe0a2702b

Observation 145f5443-f828-4631-87b0-8ea6a3de16bb · outbound

This paper cites Efficient memory management for large language model serving with PagedAttention.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Efficient memory management for large language model serving with PagedAttention

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:29.105821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:23.122706Z digest=sha256:9ab17dcbfeb9ea3c457eacd179f52dc91a5e70f9ed2cc3d4e5828197776626dd

Observation 0d23d8a8-73ad-40d6-91da-9b8aedc2cdec · outbound

This paper cites Transformers KV caching explained.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Transformers KV caching explained

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:29.006858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:23.182303Z digest=sha256:10c783368f65b3fba24d19fac7090dc3e1ffc0ee36c68302c9296ff20c998e20

Observation 45caf8d9-9714-4566-b9fa-0db1dcdb3cea · outbound

This paper cites Sequence parallelism: Long sequence training from system perspective, 2022.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Sequence parallelism: Long sequence training from system perspective, 2022

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:28.761856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:23.303096Z digest=sha256:3407df8087ad299ba55fbe0da15f85c4ea03ac4db61f6273c6a6b691a5da2abc

Observation 64e64901-3b20-4764-b6a6-651291e110b3 · outbound

This paper cites Llama 3.2: Revolutionizing edge AI and vision with open, customizable models.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Llama 3.2: Revolutionizing edge AI and vision with open, customizable models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:28.549453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:23.353733Z digest=sha256:fdd265c8f05fb805b45d8c11a89647ea0530984031b18e49f9364c2807c555ae

Observation e8618f5a-0dd0-46b8-a1b4-6c6368e7b5cd · outbound

This paper cites NVIDIA TensorRT-LLM.https: //docs.nvidia.com/tensorrt-llm/index.html.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures NVIDIA TensorRT-LLM.https: //docs.nvidia.com/tensorrt-llm/index.html

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:28.335252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:23.468448Z digest=sha256:8b8ddee7fef175fb12a174e2bb2e471ae9b89b27d4196794524dea7669ccca79

Observation bc6c61ce-9f85-444a-8f49-441f6dd4c07c · outbound

This paper cites OpenAI o3-mini.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures OpenAI o3-mini

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:28.084574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:23.581221Z digest=sha256:76bbd2be932220a62284f9aa1606f0d823974d57c31582ca847ab0f43e9df6e5

Observation 294d1086-4bb9-4ba4-8247-257ff3a68a55 · outbound

This paper cites Splitwise: Efficient generative llm inference using phase splitting, 2024.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Splitwise: Efficient generative llm inference using phase splitting, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:27.810127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:23.701320Z digest=sha256:ca141b2a51562c9d87d13d99a787034e0e4db6d6abf6633300bb6c3b7cc35e75

Observation dd690d7d-cf99-41df-8990-1f63a9f18590 · outbound

This paper cites Mooncake: A KVCache-centric disaggregated architecture for LLM serving, 2024.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Mooncake: A KVCache-centric disaggregated architecture for LLM serving, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:27.600512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:23.849100Z digest=sha256:93d67d83bce41250201a9b10ab902ed978ca02cfde6cb641943d1cbd1c40364f

Observation fa45246a-90e3-447d-98ee-280bf8e2edb9 · outbound

This paper cites Focus: For tech giants, AI like Bing and Bard poses billion-dollar search problem.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Focus: For tech giants, AI like Bing and Bard poses billion-dollar search problem

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:27.412217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:23.901389Z digest=sha256:5f7c652b7f2650a9772f9a7936db05af79f90f9b3d71dfb0def275282eba8d2f

Observation a06f4a7f-9185-4cf2-9fd8-08725cf7b59b · outbound

This paper cites an unresolved cited work.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:18:27.145596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:23.942676Z digest=sha256:5818bfb68b81bd442e56af5b625d7820d6a7bfbd0feb61ce51df7c91b4e8ef00

Observation 23909fa7-09e8-4b89-96af-7ccdb9dbc366 · outbound

This paper cites Megatron-lm: Training multi-billion parameter language models using model parallelism, 2020.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Megatron-lm: Training multi-billion parameter language models using model parallelism, 2020

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:23.999552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:23.999552Z digest=sha256:5265388c63e3d65c0a5f599275796065457d33ca3313344ee69bf0127dc8175d

Observation c7662fe8-a85d-4e52-a49c-f766c9634da4 · outbound

This paper cites Gomez, Łukasz Kaiser, and Illia Polosukhin.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Gomez, Łukasz Kaiser, and Illia Polosukhin

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:26.742702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:24.118959Z digest=sha256:a3f73ae3e05aaca38af277778300efa109eef2870862d38ec2fb20efc502b721

Observation ed9c0fd2-6578-4842-a611-4a965f03cead · outbound

This paper cites https://docs.vllm.ai/en/v0.4.2/index.html.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures https://docs.vllm.ai/en/v0.4.2/index.html

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:26.408769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:24.229841Z digest=sha256:6e5213e89b771f93ab5e266f52d61cd0722e29fa3e3bdf3299861e7b8919b0b1

Observation 2f0c0e4d-6004-450f-8805-292f75ab23b8 · outbound

This paper cites https://github.com/vllm-project/vllm-ascend.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures https://github.com/vllm-project/vllm-ascend

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:26.114753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:24.346122Z digest=sha256:6b3f375ccb1c3905e3346e89eac3bb70e5fc6ff51a8a3a1d22a964c0b570bb39

Observation 248364bd-e2c4-427e-ac45-4f940f6b129f · outbound

This paper cites Simai: Unifying architecture design and performance tuning for large-scale large language model training with scalability and precision.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Simai: Unifying architecture design and performance tuning for large-scale large language model training with scalability and precision

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:25.835717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:24.449900Z digest=sha256:b1db4a6579fb6b2120593635b9bc146869e83cd6c4c6b202f7f54f2c4909fd71

Observation 86364566-a7b8-4d13-b770-c3fc07d52163 · outbound

This paper cites Roofline: an insightful visual performance model for multicore architectures.Commun.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Roofline: an insightful visual performance model for multicore architectures.Commun

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:25.636873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:24.579266Z digest=sha256:e8bfed9c157d8b6fb431947c08fc8906622acc1c67d72bc732e925d72428e2a5

Observation 9ac34d5f-de27-4130-a262-bbfaf2fe842d · outbound

This paper cites Orca: A distributed serving system for Transformer-Based generative models.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Orca: A distributed serving system for Transformer-Based generative models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:25.376580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:24.712230Z digest=sha256:ab7e6c5797a85299eacbf558cf88914b9111a7d798fcc2fdc13c608360d75272

Observation 9e1804fc-072c-44a9-b496-b7a853b79dec · outbound

This paper cites Root mean square layer normalization, 2019.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures Root mean square layer normalization, 2019

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:24.846042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:24.846042Z digest=sha256:669d187e6f358ce4ec6884d215343baf18d2bd07024cb09f6ec5875bd7a2f13f

Observation a0d35056-e58c-43c2-81eb-367868a0e7fd · outbound

This paper cites DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving.

BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:25.146940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:24.943745Z digest=sha256:e2f1101e4922987f6174b6794d701175cd55571e911db508115cb7f876e5add7

Pith citing papers

Observation 23e4ae32-5d5a-4207-aa2e-5810f04e42cb · inbound

Simple is Better: Multiplication May Be All You Need for LLM Request Scheduling cites this paper.

Simple is Better: Multiplication May Be All You Need for LLM Request Scheduling BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T20:35:48.828322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T20:35:48.828322Z digest=sha256:d55517f11591f9437bd1fba2a6494990e8acb8ffae9c4352bf7b631502fc6aab