Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T17:12:42.195329Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2509.11076.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T17:12:42.195329Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
62 of 62 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 797489cb-11f9-413d-98b5-cfc49ddc171f · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Murray, Benoit Steiner, Paul Tucker, Vijay Va- sudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f6eb688-5040-446e-94c0-2d7e4ac829ce · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Tensorflow eager: A multi-stage, python-embedded dsl for machine learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a26902c-15f2-4f2b-9a93-0107a2327dde · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Accessed: 2024-12
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc988602-2645-4db1-9e0f-18c3d9c30f3e · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00ec6fc1-a5c3-4e06-ab35-60000b2a2eae · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Accessed: 2025-06
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36fd0431-fc20-45f3-8832-33b7e06c7949 · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Accessed: 2025-06
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc0438bf-8698-4d22-8476-b9707d6b137d · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Accessed: 2025-08
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92711af2-aa84-4f57-9e5b-f9c68b8f4ffc · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences B, Anshuj Garg, and Purushottam Kulkarni
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac6b6635-7635-43d9-a051-917d6e76d559 · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Qwen2.5-vl technical report, 2025
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31ce71fd-f98c-4098-b640-211ffe57b7e9 · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences CSWAP: A self-tuning compression framework for accelerating tensor swapping in gpus
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b97bfda7-41b9-4de2-aa49-ad4d51ebbf61 · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccc4a09a-8b73-4d0e-a888-bec7640f8965 · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Training Deep Nets with Sublinear Memory Cost
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5bf22d7-432e-45f2-8786-874d6d5015f2 · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 991759fb-2d42-4678-8ba3-50592d5dfc2d · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dc67573-48e4-4a10-8800-166503e54001 · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Parallel training of pre-trained models via chunk-based dynamic memory management.IEEE Transactions on Parallel and Distributed Systems, 34(1):304–315, 2023
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d488a174-db82-432b-a6e4-cbf9d82e8be2 · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 577d206e-533d-4e19-8552-98d9b477e10d · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences En- abling parallelism hot switching for efficient training of large language models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e734dde-fc62-49af-ac54-bb191f94b410 · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences The Llama 3 Herd of Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1664a8c0-4114-42f1-b40d-291df46e3d21 · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Gmlake: Efficient and transparent gpu memory defragmentation for large-scale dnn training with virtual memory stitching
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b24d8825-5e2f-40c8-ae01-4e0cb9ffa612 · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f4bbcf7-afd8-4938-b592-9972c2fae70b · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Transcending runtime-memory trade- offs in checkpointing by being fusion aware
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e70ee0bb-0f6a-43c6-9bc1-d01d03381512 · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Gpu memory usage optimization for backward propagation in deep network training.Journal of Parallel and Distributed Computing, 199:105053, 2025
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec6e2bc7-2f9e-4af7-acdd-b6a30267b4ad · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Meg- taichi: dynamic tensor-based memory management optimization for DNN training
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17c1bc48-47e6-410a-a885-9293eda36044 · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Swapadvisor: Pushing deep learning beyond the gpu memory limit via smart swapping
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14854789-2eaa-4efe-a6d4-a27db1a2deb8 · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Gpipe: Efficient training of giant neural networks using pipeline parallelism
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66d0deac-cbf2-48f4-a6f0-546be34848af · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Oobleck: Resilient distributed training of large models using pipeline templates
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ff79736-8a72-40af-ae55-8c2e55ce3678 · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Caffe: Convolutional architecture for fast feature embedding
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df975625-0761-4e0f-8c74-8680373070e1 · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences MegaScale: Scaling large language model training to more than 10,000 GPUs
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 652e6a8d-12a5-4cb6-883d-3c3082d39e6f · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Scaling Laws for Neural Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65947a89-090d-4701-9c62-18e187245652 · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences On model parallelization and scheduling strategies for distributed machine learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fca8bb7-8ace-43bf-875b-b563b2d02924 · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Cognitive Intelligence and Robotics
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d421a1d-4d1e-42bd-8ef1-5a95c2d6bdf6 · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Pytorch distributed: Experiences on accelerating data parallel training.Proc
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0c567be-f043-4bf9-887a-a4268ca5874a · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Parameter-efficient sparsity for large language models fine-tuning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f0adc43-116f-483a-9fe9-aae4ad391a05 · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Model compression for deep neural networks: A survey.Computers, 12(3), 2023
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86ea1121-37cf-41cc-ad4e-653d37d0273c · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Ascend: a scalable and unified architecture for ubiquitous deep neural network computing : Industry track paper
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 803c171f-6c55-4eb2-9368-4dd7c474f4cf · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Diamos, Erich Elsen, David García, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0af72ea8-2d94-41af-8d53-0b3ce91bd13a · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Devanur, Gregory R
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 530e2014-7b46-441a-a75d-097e2356a29b · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Accessed: 2025-06
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d355a5f-310c-4c2c-a5c6-020146ad02d6 · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Accessed: 2025-08
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5af6af36-3c22-4e58-8ea6-47c062f99fa0 · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Unresolved cited work
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c057b31-3913-4ae6-a3bb-de4910fa96fc · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Pytorch: An imperative style, high-performance deep learning library
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b97db3e3-59a7-4e66-9ae7-b05b95b51096 · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Capuchin: Tensor-based gpu memory management for deep learning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cdb95ca-0b29-46ad-b612-db34d0449ba7 · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Accessed: 2024-12
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0127d64-5468-4a1f-aab3-8549765671f7 · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Accessed: 2024-12
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d504001-4fe5-46bf-90d8-71309c1a96d6 · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Zero: Memory optimizations toward training trillion parameter mod- els
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24671d30-c663-4185-902b-cf4e76e265dc · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Zero-infinity: breaking the gpu memory wall for extreme scale deep learning
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53cef652-1a4c-4b02-9fa2-3421141fba23 · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Deepspeed: System optimizations enable training deep learning mod- els with over 100 billion parameters
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 419cae02-0a1a-4773-94b8-b87654d7d60f · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Sentinel: Efficient tensor migration and allocation on heteroge- neous memory systems for deep learning
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b689f6b-2dca-4230-9d7c-9976f561fe33 · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences ZeRO-Offload: Democratizing Billion-Scale model training
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 093ce6c5-fe1c-4f2b-a3c8-896e179120dd · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Unresolved cited work
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba47f7de-9267-43e5-a731-1ddbffb5a420 · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Unresolved cited work
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd02899f-e2d0-4484-92e4-04d008c6d704 · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Cntk: Microsoft’s open-source deep- learning toolkit
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3efdc510-6f06-4b0c-b126-ec89b2d1038c · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Neural machine translation of rare words with subword units, 2016
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 328f9522-c21b-4049-937f-8db6f5e0747e · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63e9c211-9b63-49fa-b2d5-cf73a98b25f1 · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Unresolved cited work
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d197b74b-c442-4dca-9805-c1495afc1e26 · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Bamboo: Making preemptible instances resilient for affordable training of large DNNs
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4155000a-7a13-4a9d-8bcf-a9f4bbcfafd3 · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Superneurons: dynamic gpu memory management for training deep neural networks
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 152ffe27-0bca-4fe3-902c-ff49c1742f05 · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69cc361f-fde4-42be-bfaf-7a27b74d9724 · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Unresolved cited work
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 933fbcee-59db-412d-829e-dd2240db6f7d · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Accelerating the training of large language models using efficient activation rematerialization and optimal hybrid parallelism
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 788b3f30-19de-46ae-bb15-efcf0b91e973 · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Pytorch fsdp: Experiences on scaling fully sharded data parallel.Proc
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eae79eea-5b24-4664-adc3-0160a0de8350 · outbound
SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences Unresolved cited work
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.