Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T14:45:20.821661Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 0 inbound Pith citation observations for arXiv:2507.18006.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T14:45:20.821661Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
64 of 64 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c2e070f6-fd83-452e-81fd-310b7c0dd7c0 · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5dd5571-5204-4e44-b7d8-0eccd1211acf · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 603b8f94-dccd-40a7-8181-7e40f6adac5b · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling DeepSeek-V3 Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 738dbaeb-a711-43db-bfa1-62f073dc36c5 · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Black-Box Tuning for Language-Model-as-a-Service
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db8eb68e-5004-45fc-b8fa-c030f11af1f7 · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Xing, Hao Zhang, Joseph E
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 55b91946-2e24-42ce-8a21-b0f42ff41d2e · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling A Comprehensive Survey of AI-Generated Content (AIGC): A History of Generative AI from GAN to ChatGPT
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71922fdf-ee1e-4418-99a7-d8fa77b1303b · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Evaluating Large Language Models Trained on Code
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdf5a873-497e-407e-862f-77e370601bbc · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Code Llama: Open Foundation Models for Code
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51fe50b3-27fa-4952-8078-3a7759ba61d3 · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Accessed: Apr
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b7a24f66-cd92-485f-a4af-09779ae31ac7 · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Accessed: Apr
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d5d67e8e-aa31-40b0-85f3-1718fcff4046 · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Smoothquant: Accurate and efficient post-training quantization for large language models, 2024
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ed255ca-b436-478e-be84-cba993f9765f · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Plug-and-play: An efficient post-training pruningmethodforlargelanguagemodels
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 226107fc-4b94-4798-8697-ebada942d07b · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Awq: Activation-aware weight quantization for llm compression and acceleration, 2024
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bbac1862-4e2a-4f20-9b8f-9b4d6315ab69 · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Zico Kolter
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7f6173b9-8bbc-4803-a0ec-d755b9c17f1b · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Multiplexing dynamic deep learning workloads with slo-awareness in gpu clusters
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 43514fe2-a55b-48e3-862c-9f96811117c8 · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Cloudnativesim: A toolkit for modeling and simulation of cloud-native applications
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ab4ad36b-d4ea-4784-8c2b-96f02774f583 · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Llminaflash:Efficientlargelanguagemodelinferencewith limited memory, 2024
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 59bf1481-178e-4fc5-b5d8-ae5b96ba2e8a · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Spotserve: Serving generative large language models on preemptible instances
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0736df2e-d30d-49e9-8f47-458134c404ed · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Llumnix:Dynamicschedulingforlargelanguage model serving
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 559ffcaf-15c7-4c90-8d29-7712367d1853 · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Serving heterogeneous machine learning models on multi-gpu servers with spatio-temporal sharing
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9c86c373-b4bc-422f-b8be-c2e46e1f2c0b · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Inferline:latency-aware provisioningandscalingforpredictionservingpipelines.In Proceedings of the 11th ACM Symposium on Cloud Computing, pages 477–491, 2020
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c8937796-4580-4d32-a58c-fee8349597d3 · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Optimizing llm inference throughput via memory-aware and sla-constrained dynamic batching, 2025
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ae8538a9-5dcf-4c02-88d2-141892950ff8 · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Alloystack: A library operating system for serverless workflow applications.Pro- ceedings of the Twentieth European Conference on Computer Systems, 2025
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0777050f-f079-44e5-8789-ffd80508e5ea · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Lora-flow: Dynamic lora fusion for large lan- guage models in generative tasks
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f1344fe9-90df-4ba6-8d48-ab40f74d133f · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Accessed: Apr
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 01a69b05-5d9b-463a-b295-4217399f77da · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2213fac0-1631-4eec-bdc7-3978d02e6221 · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1421a709-5b0d-4168-99c6-cff475463e4b · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Attention is all you need.Advances in neural information processing systems, 30, 2017
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25fa635e-8597-4a2c-a108-282095a2f577 · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Llm inference serving: Survey of recent advances and opportunities, 2024
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 85001b70-bf8c-41bd-9932-9291d297961a · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Optimizing mixture-of-experts inference time combining model deployment and communication scheduling, 2024
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3df35fa0-3e6a-4d50-95ce-866af36fc334 · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Mooncake: A kvcache-centric disag- gregated architecture for llm serving, 2024
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ec14dcb2-cd43-4691-8eb5-08930a6dfee8 · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Spinfer: Leveraging low-level sparsity for efficient large language model inference on gpus
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1ecc3806-31e9-491d-ae80-e0550faa2b0b · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling H2o:Heavy-hitteroracleforefficientgenerativeinference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 76faaf30-a734-4153-8442-4a54223cfa77 · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Splitwise: Efficient gen- erative llm inference using phase splitting.2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA), pages 118–132, 2023
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c4c725c7-5a59-4d32-8bf4-3e95d7168ab7 · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 43ce76af-0cca-4fae-bae4-1514640d0afa · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Inference without interference: Disaggregate llm inference for mixed downstream workloads, 2024
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 904791a1-6169-4a39-a9f4-18cb5f46ec30 · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Dynamollm:Designingllminferenceclustersforperformance and energy efficiency, 2024
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d9ebb7f2-f091-404d-8285-19b0415de7b8 · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Skyserve:Servingaimodels across regions and clouds with spot instances
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d650dde7-70e7-4800-9354-556790f414ee · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Towardsefficientandreliablellmserving: A real-world workload study.arXiv preprint arXiv:2402.XXXXX, 2024
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 50e2209c-e0bc-4164-972b-9d186e868a2e · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Usher: Holistic interference avoidance for resource optimized ML inference
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 71fe8728-f510-4dde-8c3a-93b4cea88868 · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 515dbeab-5865-4a1e-89d1-4b2dbc2a8595 · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b17279ee-733a-4137-87dc-7c2105163f89 · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Alpaserve: Statistical multiplexing with model parallelism for deep learning serving
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3db551ec-4c4a-449c-8a13-15ea022f6fdf · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8e952f21-82e8-4dfb-aeab-26bb3c1c3e81 · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Yadwadkar, and Christos Kozyrakis
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 29feca7a-4ba7-4d1c-870a-f553e4fdb36f · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Reducing Activation Recomputation in Large Transformer Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6192e0ab-e576-407f-938f-4a21570f9f37 · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling On parallel processing systems: Amdahl’s law generalized and some results on optimal design
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3607326f-b68b-41b0-bfc0-29a9cb5085ec · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling xformers: A modular and hackable transformer modelling library.https://github.com/facebookresearch/ xformers, 2022
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b995e9ff-346e-43c5-a422-8101a97dc13e · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling https://developer.nvidia.com/management-library-nvml, 2025
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 33dca7d2-896f-4afa-a787-f6a83496e92f · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Orca: A distributed serving system for transformer- based generative models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 080e1029-b4cb-43b7-9733-66fb8c5041d6 · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Uellm: A unified and efficient approach for large language model inference serving
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 85ae5a38-9a43-400b-9e93-4dcf3748b8cf · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Stanford alpaca: An instruction- following llama model.https://github.com/tatsu-lab/stanford_alpaca,
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f8ce4997-8690-4e79-92bb-6d436cbc5305 · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Mepipe: Democratizing llm training with memory- efficient slice-level pipeline scheduling on cost-effective accelerators
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d44c4132-fcb8-4e9b-b5eb-fe849f84db51 · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Le, and Z
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ab48421c-b07c-4781-905c-e160553b1ac6 · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Mist:Efficientdistributedtrainingoflargelanguagemodelsviamemory- parallelism co-optimization
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b2b93895-5df9-4f87-91b4-4b5a78aab060 · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 848a515f-29f3-418e-8124-111d46d06335 · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Alpa: Automating inter and intra- operator parallelism for distributed deep learning
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a7567b5a-d5ea-4aff-aee2-d1c5c27566fc · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Fast state restoration in llm serving with hcache
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5b2ca156-3b79-4afc-8b43-a3107e2d25aa · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Fast and live model auto scaling with o(1) host caching, 2024
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 106593b2-91f9-4854-87af-93f5faf2cc46 · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d07250cf-5f05-4c41-8fdb-e4ecfcde3f15 · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Infinigen: Efficientgenerativeinferenceoflargelanguagemodelswithdynamickv cache management
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bb71e73a-2187-431a-8428-e8c3931b8909 · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Unresolved cited work
Reference 411
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a18aea7a-6abb-454a-a786-988f05df980b · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Unresolved cited work
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ef497c9-742a-44aa-9990-dbfd9f1c122c · outbound
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Unresolved cited work
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
No inbound Pith citation observations are available.