Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T20:39:32.123038Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 1 inbound Pith citation observation for arXiv:2508.21613.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T20:39:32.123038Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-15T13:17:47.603967Z
A source-named dated measurement, never combined with another source.
Source: cited_works
37 of 37 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a07ee32b-c396-40ac-94e6-593d3e531650 · outbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection The Llama 3 Herd of Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a669348c-f4e7-4781-aa37-1e9f897d1663 · outbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 69a3850e-4219-44da-9acd-4ea8f9816459 · outbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection URLhttps://doi.org/10.1145/3600006.3613145
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f09d9a6c-27fa-4ba6-8249-1960f72f5a28 · outbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Check- N-Run: a checkpointing system for training deep learning recommendation models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a0ac690c-53b7-4fe9-a033-64c61ee2c191 · outbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 794654ba-8e13-4695-8cb7-ba2f66719e9e · outbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Elan: Towards generic and efficient elastic training for deep learning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d4297cb2-d11b-4e9f-9cef-3bd8fc5418ae · outbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Bamboo: Making preemptible instances resilient for affordable training of large{DNNs}
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 67f52b03-5938-475c-ace7-616fbdcd2d7f · outbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Oobleck: Resilient distributed training of large models using pipeline templates
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 29c021fc-a986-49b2-aaa3-2a54e214225a · outbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Recycle: Re- silient training of large dnns using pipeline adaptation
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6f353f01-72b2-41cf-9dcf-189c503483a4 · outbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Tenplex: Dynamic parallelism for deep learning using parallelizable tensor col- lections
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d9b38c0b-7c6c-4324-81aa-9e6315f0da30 · outbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Ascend: a scalable and unified architecture for ubiquitous deep neural network computing: Industry track paper
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f5908091-346d-4114-b213-3a3c3abb219f · outbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Parallel scan on ascend ai accelerators
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 710d63b9-f237-49f7-aa67-04ff0cbfa936 · outbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 23a3272d-7a27-41ee-b508-837b62afecac · outbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Horovod: fast and easy distributed deep learning in TensorFlow
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 31891851-6ccf-4e31-85e2-c3352e6269df · outbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection ImageNet Training in Minutes
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3356bce0-fe8f-4732-ac13-4893bee81912 · outbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2c2149b2-0772-49b8-8097-ef3a26bf5d0c · outbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e77b0450-8f64-4cc5-b1fd-2ff4c85a717e · outbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection BPipe: Memory-balanced pipeline parallelism for training large language models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d77e79dd-3b5b-4704-b993-0431996b3dbc · outbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 42bc76e4-b3cf-475d-bf2b-430ed40f5930 · outbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 863dbd24-f1b2-4d47-8830-6895e417b805 · outbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2944bfda-781b-43ad-b12b-2891759dd942 · outbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Moe parallel folding: Heterogeneous parallelism mappings for efficient large-scale moe model training with megatron core
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c0b33fbc-ced1-4f00-a407-01b95eaefe16 · outbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Switch transformers: scaling to trillion parameter models with simple and efficient sparsity
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ea35a451-8a4a-4010-99dc-292664997f1d · outbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Tutel: Adaptive Mixture-of-Experts at Scale
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8c68d54a-5f6f-4055-8ceb-25edab022570 · outbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Megascale-moe: Large-scale communication-efficient training of mixture-of-experts models in production
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2e840f52-915e-4557-b1ca-4001fd5428a6 · outbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Understanding communication characteristics of distributed training
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8e0a46fb-5b31-474c-bba2-271646555743 · outbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Amped: An analytical model for performance in distributed training of transformers
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2310f253-5f7a-4c4f-b35c-c2b3c75cacbb · outbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Reducing Activation Recomputation in Large Transformer Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 19f3c7a3-a5bf-47ea-b116-53ad995d2a1d · outbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , articleno =
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f6dbc7bb-df01-4242-a69f-09c725584504 · outbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Pytorch distributed: experiences on accelerating data parallel training
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b240e950-ad48-452e-9dcb-e6e8d4cee2eb · outbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Varuna: scalable, low-cost training of massive deep learning models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7fea385a-ea6b-4e33-bc6e-a204e02792b8 · outbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Failures in large scale systems: Long-term measurement, analysis, and implications
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ff22b037-5091-495f-a53b-7b2cc18184a2 · outbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection MLaaS in the wild: Workload analysis and scheduling in Large-Scale heterogeneous GPU clusters
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 93b186d6-d0e8-4f84-97e9-d482ff7898e2 · outbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Minder: Faulty machine detection for large-scale distributed model training
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation babf272d-7115-4c67-8df3-a0bcf6e60010 · outbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Alpa: Automating inter- and Intra-Operator parallelism for distributed deep learning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 492049b6-3fee-44ae-8f88-c66af5985182 · outbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection The hungarian method for the assignment problem
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6b00ff77-1b97-4770-ae2a-8655b4f723cd · outbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9d0e2fe2-aee7-4187-91f1-908239989285 · inbound
Enhancing OLAP Resilience at LinkedIn Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.