Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T18:28:01.641782Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2411.11560.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T18:28:01.641782Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
42 of 42 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8f64c548-f65d-4eba-b4dd-645c130c0fbc · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads A survey on large language models: Applications, challenges, limitations, and practical usage
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3303931-8436-4e1c-b87e-27fcafcf2e7a · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads Efficient Training of Large Language Models on Distributed Infrastructures: A Survey
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation decb8d3c-e512-4ab7-a628-a57eb2b26359 · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1838389c-2088-4752-ae94-a643ce2ca9f1 · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads Gödel: Unified large-scale resource management and scheduling at bytedance
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 686d5a79-898d-46f3-9724-e2a7a3495913 · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads Topology-aware gpu scheduling for learning workloads in cloud environments
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cb7af90d-5efa-4522-b89b-27aadcf7a300 · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads Numa (non-uniform memory access): An overview: Numa becomes more common because memory controllers get close to execution units on microprocessors
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c2a87714-bfe2-4b13-bb56-e6dce973d287 · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b274c18f-8aa3-4049-8c3b-ef85184042f6 · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads Fastertransformer
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d72f1f46-7198-4ed8-b7e6-929dac152349 · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads Efficient memory management for large language model serving with pagedattention
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8059a009-b02e-4a5e-90ab-5b5462e55ffc · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 55bbb4f4-33c7-4400-9a72-8dedbf17fc3d · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads Huggingface text generation inference
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8cf4ed1a-5c32-4f90-8fc4-6828229232e4 · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads Deepspeed inference
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 24929857-98da-480d-bb69-5d8d84af3978 · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads Tensorrt-llm
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c33301c8-4278-4565-8476-81bae1b4adad · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35e1fa7c-29a6-48ae-8bf0-d7501de9fdf6 · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads Kubernetes topology manager moves to beta
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4a4273a1-70b4-4d0e-a9fb-d136a6eafbe2 · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads Pod priority and preemption
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 56641ff5-55ad-4d69-a0ed-47f142ee90b4 · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads Godel scheduler: a unified scheduler for online and offline tasks
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 65743995-cd8e-4356-bc95-b0d38e2e1028 · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads Daemonset
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 38571f2a-df86-4258-95d3-9677b53151bd · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads Kubernetes without kubelet
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 60eb0db1-4c97-4a32-ab61-2449af02ad9f · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads Control topology management policies on a node
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 88a88c86-aed8-4e04-8cf0-847c994707ee · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads Towards {GPU} utilization prediction for cloud deep learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b1a8bceb-c2a4-40f6-8530-0d45e9a5ee3f · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads Horus: Interference-aware and prediction-based scheduling in deep learning systems
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 925f7790-266f-4f84-b56e-a683690cf4c0 · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads Beware of fragmentation: Scheduling {GPU-Sharing} workloads with fragmentation gradient descent
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b9e132b9-40af-44c1-bf05-587c08e6e83b · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads {HiveD}: Sharing a {GPU} cluster for deep learning with guarantees
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a71a4c0e-220e-4f0c-b4f4-ff8a56479a6f · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads Supporting gpu sharing in cloud environments with a transparent runtime consolidation framework
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 748ddaea-6be4-49c5-ac6c-c15e16774fd0 · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads Fine-grained gpu sharing primitives for deep learning applications
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8223a453-87fd-43c1-93e8-110117107960 · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads Gpushare: Fair-sharing middleware for gpu clouds
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4506a771-c34a-4ba0-9258-99d951452b22 · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads Nvidia cloud native technologies: Gpu sharing
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4c09c855-110c-49b5-b9a1-3475a4e1f3c7 · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads Advanced features in ibm power8 systems
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1fb6f2b0-c3d3-45af-a511-191befec4d7f · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads Performance evaluation of the nvidia tesla p100: Our directive-based partitioning and pipelining vs
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation deeb195f-5236-4900-8171-18ad8d71598f · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads Topology-aware scheduling framework for microservice applications in cloud
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation dbf852d4-d6ba-44a7-a2ef-b31305b0d091 · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads Katalyst core
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 93263d25-13c5-437c-9a69-5016325fcf88 · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads Topology-aware resource allocation for data-intensive workloads
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6f4c2c72-2123-4979-afa7-c94df6b035d9 · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads Towards topology aware pre-emptive job scheduling with deep reinforcement learning
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 792f1e0a-0f86-4e97-a7ec-91a733a6150d · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads Microsecond-scale preemption for concurrent {GPU- accelerated}{DNN} inferences
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cc0d2b9d-a5d2-457a-9492-ffaf42566c25 · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads Efficiently programming large language models using sglang
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d548ec38-5922-4dd3-826e-3af1d996ebbd · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads Orca: A distributed serving system for {Transformer-Based} generative models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ff1a8a3-098a-48de-9bd6-1c601afba6c4 · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads Fast Distributed Inference Serving for Large Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 799a9f9b-e866-443b-a8f8-1e266cf3c4c9 · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads Bert loses patience: Fast and robust inference with early exit
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2ef5b5fb-8604-47f4-9f99-e25ed9042c7d · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads Flashattention: Fast and memory-efficient exact attention with io-awareness
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7a7f4e6-21d2-45c6-bc1d-78087adeda08 · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads Sparsegpt: Massive language models can be accurately pruned in one-shot
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 563b8ecf-5b3a-47ba-af67-92051c370136 · outbound
Topology-aware Preemptive Scheduling for Co-located LLM Workloads Awq: Activation-aware weight quantization for on-device llm compression and acceleration
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
No inbound Pith citation observations are available.