Pith. sign in

Paper Citation Record · LEDGER

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling

As of 7 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 0 inbound Pith citation observations for arXiv:2507.18006.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.18006 v1

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:45:20.821661Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

64 of 64 outbound references displayed

  • verified exact0
  • verified fuzzy43
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c2e070f6-fd83-452e-81fd-310b7c0dd7c0 · outbound

This paper cites GPT-4 Technical Report.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.562559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.562559Z digest=sha256:1113648a95c519cc0a6bfc3bab6827ded3aef8a23f683a60aa4259008584a6e8

Observation f5dd5571-5204-4e44-b7d8-0eccd1211acf · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.567366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.567366Z digest=sha256:139867320109200b6e69921e3ec7533904566e055d78efdf212469107d2f6e18

Observation 603b8f94-dccd-40a7-8181-7e40f6adac5b · outbound

This paper cites DeepSeek-V3 Technical Report.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling DeepSeek-V3 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.571852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.571852Z digest=sha256:7d3f33fe321cef8b2ed1690614ca2b10f4902fb79cd6bfb24b2021ae03da0d6b

Observation 738dbaeb-a711-43db-bfa1-62f073dc36c5 · outbound

This paper cites Black-Box Tuning for Language-Model-as-a-Service.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Black-Box Tuning for Language-Model-as-a-Service

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.576394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.576394Z digest=sha256:b2f9dc2aaa6c58e8aef2225851bcce5f6e37fb2eb8d2f22eb7daa1a8883773fa

Observation db8eb68e-5004-45fc-b8fa-c030f11af1f7 · outbound

This paper cites Xing, Hao Zhang, Joseph E.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Xing, Hao Zhang, Joseph E

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:28.556477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.581123Z digest=sha256:1f47a2d80dda27c8a9ce9e493e08542cb5fc6dcc000f028e6e2cc2aac5a14e05

Observation 55b91946-2e24-42ce-8a21-b0f42ff41d2e · outbound

This paper cites A Comprehensive Survey of AI-Generated Content (AIGC): A History of Generative AI from GAN to ChatGPT.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling A Comprehensive Survey of AI-Generated Content (AIGC): A History of Generative AI from GAN to ChatGPT

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.585569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.585569Z digest=sha256:d88717d8a3a28ab468b6f5214b514911af6b1fd9a993e8c8e3b6d8618605b5c7

Observation 71922fdf-ee1e-4418-99a7-d8fa77b1303b · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Evaluating Large Language Models Trained on Code

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.591384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.591384Z digest=sha256:392ee68dbcf34307a0a9bf5109eb929042b5fb3201b85b7b7db9238b272e69c5

Observation bdf5a873-497e-407e-862f-77e370601bbc · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Code Llama: Open Foundation Models for Code

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.595686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.595686Z digest=sha256:07b691ccb43f54f41163bf2b0ab8260299508036830ce8e900cd916a3e3283a8

Observation 51fe50b3-27fa-4952-8078-3a7759ba61d3 · outbound

This paper cites Accessed: Apr.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Accessed: Apr

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:28.337931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.600056Z digest=sha256:cda94eaf769b1f2abfb00e1ec7c5f1670c161e572705e39e473565cafa25cf47

Observation b7a24f66-cd92-485f-a4af-09779ae31ac7 · outbound

This paper cites Accessed: Apr.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Accessed: Apr

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:28.135926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.603833Z digest=sha256:d3affc005fb4c60c912cef776622a54d79fc178a439a38337739c93eec1a38c4

Observation d5d67e8e-aa31-40b0-85f3-1718fcff4046 · outbound

This paper cites Smoothquant: Accurate and efficient post-training quantization for large language models, 2024.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Smoothquant: Accurate and efficient post-training quantization for large language models, 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.607908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.607908Z digest=sha256:d7d84c2ee169268f43651c8698554d3ad92d83c15bd2b8db505a82698bbe0fec

Observation 8ed255ca-b436-478e-be84-cba993f9765f · outbound

This paper cites Plug-and-play: An efficient post-training pruningmethodforlargelanguagemodels.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Plug-and-play: An efficient post-training pruningmethodforlargelanguagemodels

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:27.983499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.612417Z digest=sha256:a8471ae13d48b843e8ec48236ee84e2351bbe5d2d85f2872a5889b751db48520

Observation 226107fc-4b94-4798-8697-ebada942d07b · outbound

This paper cites Awq: Activation-aware weight quantization for llm compression and acceleration, 2024.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Awq: Activation-aware weight quantization for llm compression and acceleration, 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:27.805979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.616704Z digest=sha256:0fe39efa2d1346103abb47b79d487f2fce57f5ad7c855d73fba40cf2f122b13e

Observation bbac1862-4e2a-4f20-9b8f-9b4d6315ab69 · outbound

This paper cites Zico Kolter.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Zico Kolter

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:27.665893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.620795Z digest=sha256:4b7e03c97cc81214135cfa4248f5b17da7e8562293729f5abecccf3c2c440631

Observation 7f6173b9-8bbc-4803-a0ec-d755b9c17f1b · outbound

This paper cites Multiplexing dynamic deep learning workloads with slo-awareness in gpu clusters.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Multiplexing dynamic deep learning workloads with slo-awareness in gpu clusters

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:27.515390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.624755Z digest=sha256:cadbe6537ab000ce4b0baf35f6a7a7409c8bd3be743820899806cb18702ca083

Observation 43514fe2-a55b-48e3-862c-9f96811117c8 · outbound

This paper cites Cloudnativesim: A toolkit for modeling and simulation of cloud-native applications.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Cloudnativesim: A toolkit for modeling and simulation of cloud-native applications

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:27.346301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.628509Z digest=sha256:46d1dabeec1d20b3cc77ec1aeb4aa791c34e00a3a54a87e403794dd1a89c3674

Observation ab4ad36b-d4ea-4784-8c2b-96f02774f583 · outbound

This paper cites Llminaflash:Efficientlargelanguagemodelinferencewith limited memory, 2024.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Llminaflash:Efficientlargelanguagemodelinferencewith limited memory, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:27.162975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.632443Z digest=sha256:6e369be49d3751017d9e33b9fbee8468f890217cb914c21738b55a0d951c07bd

Observation 59bf1481-178e-4fc5-b5d8-ae5b96ba2e8a · outbound

This paper cites Spotserve: Serving generative large language models on preemptible instances.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Spotserve: Serving generative large language models on preemptible instances

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.636501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.636501Z digest=sha256:a54197e48136e9035eae8c3cacc848fab429f983cb480b65cdbf3481073bdc96

Observation 0736df2e-d30d-49e9-8f47-458134c404ed · outbound

This paper cites Llumnix:Dynamicschedulingforlargelanguage model serving.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Llumnix:Dynamicschedulingforlargelanguage model serving

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:26.979547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.640277Z digest=sha256:e2340457b996a398aa36d99912683001f55e94408a3a904fc2f41cb60e7ff15f

Observation 559ffcaf-15c7-4c90-8d29-7712367d1853 · outbound

This paper cites Serving heterogeneous machine learning models on multi-gpu servers with spatio-temporal sharing.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Serving heterogeneous machine learning models on multi-gpu servers with spatio-temporal sharing

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:26.787021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.643981Z digest=sha256:d43314d1e168577cc851650a70c48699ec35edf03a0d89ab2dcdbe5b90ac26a4

Observation 9c86c373-b4bc-422f-b8be-c2e46e1f2c0b · outbound

This paper cites Inferline:latency-aware provisioningandscalingforpredictionservingpipelines.In Proceedings of the 11th ACM Symposium on Cloud Computing, pages 477–491, 2020.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Inferline:latency-aware provisioningandscalingforpredictionservingpipelines.In Proceedings of the 11th ACM Symposium on Cloud Computing, pages 477–491, 2020

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:26.604327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.647609Z digest=sha256:f0b4b1fe2576b921ea16bed9acfb5b443fc549fca2f193cb78d820761b5279a1

Observation c8937796-4580-4d32-a58c-fee8349597d3 · outbound

This paper cites Optimizing llm inference throughput via memory-aware and sla-constrained dynamic batching, 2025.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Optimizing llm inference throughput via memory-aware and sla-constrained dynamic batching, 2025

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:26.431926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.651829Z digest=sha256:c9bc733046d9a604a68bcef5246deb681745b95bce3c88a9e0b3af046c3c3b74

Observation ae8538a9-5dcf-4c02-88d2-141892950ff8 · outbound

This paper cites Alloystack: A library operating system for serverless workflow applications.Pro- ceedings of the Twentieth European Conference on Computer Systems, 2025.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Alloystack: A library operating system for serverless workflow applications.Pro- ceedings of the Twentieth European Conference on Computer Systems, 2025

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:26.270323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.655748Z digest=sha256:26470d51fbafab8473e9127279d46a0cecc4d5ef8e71548075f6922613ba6087

Observation 0777050f-f079-44e5-8789-ffd80508e5ea · outbound

This paper cites Lora-flow: Dynamic lora fusion for large lan- guage models in generative tasks.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Lora-flow: Dynamic lora fusion for large lan- guage models in generative tasks

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:26.052390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.660020Z digest=sha256:ee407ba5795bee888f7e149aace8e8c2e04e3939a8cf7aa5c3ab58d7773dcdea

Observation f1344fe9-90df-4ba6-8d48-ab40f74d133f · outbound

This paper cites Accessed: Apr.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Accessed: Apr

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:25.830080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.663827Z digest=sha256:96d956de132913a7548009bc7c8b492370edbf4ce06c4a52b25fad38cf40c014

Observation 01a69b05-5d9b-463a-b295-4217399f77da · outbound

This paper cites an unresolved cited work.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:45:25.642539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.667723Z digest=sha256:163952cd8b5e30082e668b086981d942af41b9abad2726a6f5b73c4f45024532

Observation 2213fac0-1631-4eec-bdc7-3978d02e6221 · outbound

This paper cites an unresolved cited work.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:45:25.486001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.675885Z digest=sha256:0c3d1debaa83d164084d468236564e3b6ac95e9ff4d977371b87d60bfc3a3772

Observation 1421a709-5b0d-4168-99c6-cff475463e4b · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.681562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.681562Z digest=sha256:20839fd752f400030d9f3667ec47e036620492d1785af8c7df93e4e398ce926b

Observation 25fa635e-8597-4a2c-a108-282095a2f577 · outbound

This paper cites Llm inference serving: Survey of recent advances and opportunities, 2024.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Llm inference serving: Survey of recent advances and opportunities, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:25.265670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.686206Z digest=sha256:b57c5b885e09eeb16538ba727b22f4dcfde5c453f29d528b183310a23e873a4a

Observation 85001b70-bf8c-41bd-9932-9291d297961a · outbound

This paper cites Optimizing mixture-of-experts inference time combining model deployment and communication scheduling, 2024.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Optimizing mixture-of-experts inference time combining model deployment and communication scheduling, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:25.101053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.690334Z digest=sha256:44eeb8e75775990a825f878d3d5c13f6f90a9d42e45679159873b649e94ef6af

Observation 3df35fa0-3e6a-4d50-95ce-866af36fc334 · outbound

This paper cites Mooncake: A kvcache-centric disag- gregated architecture for llm serving, 2024.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Mooncake: A kvcache-centric disag- gregated architecture for llm serving, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:24.909654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.694330Z digest=sha256:e0be0e09670609087e1c5acbb992fe1c42f24c301ac1a25d392795f4b0da9fba

Observation ec14dcb2-cd43-4691-8eb5-08930a6dfee8 · outbound

This paper cites Spinfer: Leveraging low-level sparsity for efficient large language model inference on gpus.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Spinfer: Leveraging low-level sparsity for efficient large language model inference on gpus

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:24.699198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.698281Z digest=sha256:57bf5bf17ce23d018e9186a9f07416ca80b551152b0d8ec72917cd8d11f15584

Observation 1ecc3806-31e9-491d-ae80-e0550faa2b0b · outbound

This paper cites H2o:Heavy-hitteroracleforefficientgenerativeinference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling H2o:Heavy-hitteroracleforefficientgenerativeinference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:24.511967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.702463Z digest=sha256:5cca1683c635f9de34bb0b526fa73d2ff352fe5b57ec27f09d51e79e28530327

Observation 76faaf30-a734-4153-8442-4a54223cfa77 · outbound

This paper cites Splitwise: Efficient gen- erative llm inference using phase splitting.2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA), pages 118–132, 2023.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Splitwise: Efficient gen- erative llm inference using phase splitting.2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA), pages 118–132, 2023

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:24.286745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.706607Z digest=sha256:5283edc4a9088f6ca2eecfcac1f06f70cc5fa2ca312272351d35343101621ac3

Observation c4c725c7-5a59-4d32-8bf4-3e95d7168ab7 · outbound

This paper cites an unresolved cited work.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:45:24.086241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.710774Z digest=sha256:ae664d772590338ba93ac4b43c6593dfb5b723dc31655ed0b19227b0d7ccf5ea

Observation 43ce76af-0cca-4fae-bae4-1514640d0afa · outbound

This paper cites Inference without interference: Disaggregate llm inference for mixed downstream workloads, 2024.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Inference without interference: Disaggregate llm inference for mixed downstream workloads, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:23.843614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.714736Z digest=sha256:a847a0159000fc604bcc42f055a1d0dbf450feff161ee46cf1dfbdb63255c39d

Observation 904791a1-6169-4a39-a9f4-18cb5f46ec30 · outbound

This paper cites Dynamollm:Designingllminferenceclustersforperformance and energy efficiency, 2024.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Dynamollm:Designingllminferenceclustersforperformance and energy efficiency, 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:23.669168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.719014Z digest=sha256:3fa93f9285034fbabaa4b3b49fe0ee847466af176e71a3cea3c00016bd891a38

Observation d9ebb7f2-f091-404d-8285-19b0415de7b8 · outbound

This paper cites Skyserve:Servingaimodels across regions and clouds with spot instances.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Skyserve:Servingaimodels across regions and clouds with spot instances

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:23.437689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.722857Z digest=sha256:b87d6c5f29f22a835a1cf7924e920cbdd9dc6243178139c48c040d7bfdb35281

Observation d650dde7-70e7-4800-9354-556790f414ee · outbound

This paper cites Towardsefficientandreliablellmserving: A real-world workload study.arXiv preprint arXiv:2402.XXXXX, 2024.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Towardsefficientandreliablellmserving: A real-world workload study.arXiv preprint arXiv:2402.XXXXX, 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:23.227733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.727165Z digest=sha256:bce9b56e936cf2880e73ca4c2f720327962006afa4a6262f2d8d420aebe8fcf4

Observation 50e2209c-e0bc-4164-972b-9d186e868a2e · outbound

This paper cites Usher: Holistic interference avoidance for resource optimized ML inference.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Usher: Holistic interference avoidance for resource optimized ML inference

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:23.012903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.731400Z digest=sha256:30e3671850d561a4033206430f470f27822f7bec0df2397fd23970aa5034eaa5

Observation 71fe8728-f510-4dde-8c3a-93b4cea88868 · outbound

This paper cites an unresolved cited work.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:45:22.779244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.735534Z digest=sha256:236be70d5085f7e8a0a1cd889a0abc30ffdd11ee71ac45cf11d9d97e6c104d8e

Observation 515dbeab-5865-4a1e-89d1-4b2dbc2a8595 · outbound

This paper cites Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:22.599134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.739514Z digest=sha256:bf57adb24407e42277b080b1e26d1b4c1c5594ab63e1beea2deb076a7b5302f7

Observation b17279ee-733a-4137-87dc-7c2105163f89 · outbound

This paper cites Alpaserve: Statistical multiplexing with model parallelism for deep learning serving.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Alpaserve: Statistical multiplexing with model parallelism for deep learning serving

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:22.387183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.743488Z digest=sha256:dca9b76169665c98f15daa631e4e39ee5f3870489005c4907d99b648b0f43a4e

Observation 3db551ec-4c4a-449c-8a13-15ea022f6fdf · outbound

This paper cites an unresolved cited work.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:45:22.205088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.747191Z digest=sha256:c6f59725f6b8ccf8ee26973a05ec64aaf264d8344b9399e61e73c183a20131fc

Observation 8e952f21-82e8-4dfb-aeab-26bb3c1c3e81 · outbound

This paper cites Yadwadkar, and Christos Kozyrakis.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Yadwadkar, and Christos Kozyrakis

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.999977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.750754Z digest=sha256:6191a2ab8cf4bdc243bbd6fdfb4fbecfc3410f6b7ceb1cee4e641df8eb0a638a

Observation 29feca7a-4ba7-4d1c-870a-f553e4fdb36f · outbound

This paper cites Reducing Activation Recomputation in Large Transformer Models.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Reducing Activation Recomputation in Large Transformer Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.758677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.758677Z digest=sha256:0bf8331d9af97d934e8fe85986f4c2aa45348f64b60006a1f4e4f884747a73ab

Observation 6192e0ab-e576-407f-938f-4a21570f9f37 · outbound

This paper cites On parallel processing systems: Amdahl’s law generalized and some results on optimal design.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling On parallel processing systems: Amdahl’s law generalized and some results on optimal design

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.610721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.762745Z digest=sha256:5da443aa8ad5fbdadd1acdc58d895fb694604bfee16183f0577fcfd436ab32bd

Observation 3607326f-b68b-41b0-bfc0-29a9cb5085ec · outbound

This paper cites xformers: A modular and hackable transformer modelling library.https://github.com/facebookresearch/ xformers, 2022.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling xformers: A modular and hackable transformer modelling library.https://github.com/facebookresearch/ xformers, 2022

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.412262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.766581Z digest=sha256:8e2cd0a9fad569dc0833374d71906f4da2a192e46c228bbf4a58be2e6b640d41

Observation b995e9ff-346e-43c5-a422-8101a97dc13e · outbound

This paper cites https://developer.nvidia.com/management-library-nvml, 2025.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling https://developer.nvidia.com/management-library-nvml, 2025

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.313583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.770284Z digest=sha256:62daa3b1a5b31aaa2a7ff88d3094a388ca2f7da44f70211bcaf3b37ede4806c9

Observation 33dca7d2-896f-4afa-a787-f6a83496e92f · outbound

This paper cites Orca: A distributed serving system for transformer- based generative models.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Orca: A distributed serving system for transformer- based generative models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.283738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.774358Z digest=sha256:5565eef0af62a3ebd65de2b0b9060e2c640897eb03c15454179aa7ed39878e00

Observation 080e1029-b4cb-43b7-9733-66fb8c5041d6 · outbound

This paper cites Uellm: A unified and efficient approach for large language model inference serving.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Uellm: A unified and efficient approach for large language model inference serving

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.245867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.778185Z digest=sha256:37b4a909e9cce78587a02de9dfd88d3e32ad670020a6a84aef2cb43d96ead534

Observation 85ae5a38-9a43-400b-9e93-4dcf3748b8cf · outbound

This paper cites Stanford alpaca: An instruction- following llama model.https://github.com/tatsu-lab/stanford_alpaca,.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Stanford alpaca: An instruction- following llama model.https://github.com/tatsu-lab/stanford_alpaca,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.202598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.781963Z digest=sha256:88250d322c90dd3a44852ea882e39d23dea41aab0957b2c60811a6aa82764a8d

Observation f8ce4997-8690-4e79-92bb-6d436cbc5305 · outbound

This paper cites Mepipe: Democratizing llm training with memory- efficient slice-level pipeline scheduling on cost-effective accelerators.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Mepipe: Democratizing llm training with memory- efficient slice-level pipeline scheduling on cost-effective accelerators

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.113047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.789659Z digest=sha256:71bfb634824d715078dc138f08b0e282363839c082d034eb47ddebe33241d56a

Observation d44c4132-fcb8-4e9b-b5eb-fe849f84db51 · outbound

This paper cites Le, and Z.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Le, and Z

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.078885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.793446Z digest=sha256:77afb7939a252e67ee0d38cf8a2ca2d491d9be0a5ef46309c95d0b48620464f0

Observation ab48421c-b07c-4781-905c-e160553b1ac6 · outbound

This paper cites Mist:Efficientdistributedtrainingoflargelanguagemodelsviamemory- parallelism co-optimization.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Mist:Efficientdistributedtrainingoflargelanguagemodelsviamemory- parallelism co-optimization

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.055530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.797367Z digest=sha256:8b4dad90790e6401c5d61e34a8e7baa3be3d5ac80cef26d57bc64abf887a71cd

Observation b2b93895-5df9-4f87-91b4-4b5a78aab060 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.801131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.801131Z digest=sha256:06a4bf37194b5d631e3bc1e6f692392510749411d3a5c4be5d1cee42230df54a

Observation 848a515f-29f3-418e-8124-111d46d06335 · outbound

This paper cites Alpa: Automating inter and intra- operator parallelism for distributed deep learning.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Alpa: Automating inter and intra- operator parallelism for distributed deep learning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.041953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.805316Z digest=sha256:5d184801bda70f2e79578c8db8c30ac537697f4e4d6df2c42e25e90e08736596

Observation a7567b5a-d5ea-4aff-aee2-d1c5c27566fc · outbound

This paper cites Fast state restoration in llm serving with hcache.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Fast state restoration in llm serving with hcache

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.028193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.809631Z digest=sha256:3c7c8e9e8878eb7f97bc1911ea2038753a918b1465fe3ffea06eee4e1263cf81

Observation 5b2ca156-3b79-4afc-8b43-a3107e2d25aa · outbound

This paper cites Fast and live model auto scaling with o(1) host caching, 2024.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Fast and live model auto scaling with o(1) host caching, 2024

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:21.013948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.813723Z digest=sha256:d71f4875bead0a5e9510fea5aee7bbfcd75558b8cefa9f4f35fccbc975501a20

Observation 106593b2-91f9-4854-87af-93f5faf2cc46 · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.817827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.817827Z digest=sha256:48805fdce199c37958064cb8f0676e41c59d436ea72de19f253d9f3db6859ab7

Observation d07250cf-5f05-4c41-8fdb-e4ecfcde3f15 · outbound

This paper cites Infinigen: Efficientgenerativeinferenceoflargelanguagemodelswithdynamickv cache management.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Infinigen: Efficientgenerativeinferenceoflargelanguagemodelswithdynamickv cache management

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:45:20.991136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.821661Z digest=sha256:cb831385fb995f1bb8e4f7bebb6fa67ae39e658768d8c25315327491baf0897d

Observation bb71e73a-2187-431a-8428-e8c3931b8909 · outbound

This paper cites an unresolved cited work.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Unresolved cited work

Reference 411

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:45:21.824444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.754810Z digest=sha256:f87f5d0b283b07df83fb0a489a0647620928b930275343dc57914b992b8952dd

Observation a18aea7a-6abb-454a-a786-988f05df980b · outbound

This paper cites an unresolved cited work.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Unresolved cited work

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.671778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.671778Z digest=sha256:27104edf8af77ef4ff5e46a3cb3ae5542abc0a619efc6e39cd0caee412b5d186

Observation 2ef497c9-742a-44aa-9990-dbfd9f1c122c · outbound

This paper cites an unresolved cited work.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Unresolved cited work

Reference 2023

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:45:21.165178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T14:45:20.785956Z digest=sha256:e55ec930f3dada2662f8075f22b8eab137fb49a99c183886c7ce22d3df5b7b1f

Pith citing papers

No inbound Pith citation observations are available.