Pith. sign in

Paper Citation Record · LEDGER

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection

As of 7 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 1 inbound Pith citation observation for arXiv:2508.21613.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.21613 v4

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T20:39:32.123038Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-15T13:17:47.603967Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

37 of 37 outbound references displayed

  • verified exact18
  • verified fuzzy16
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a07ee32b-c396-40ac-94e6-593d3e531650 · outbound

This paper cites The Llama 3 Herd of Models.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection The Llama 3 Herd of Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:41:50.688302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:7442158f08323bc0c21d666f4c94968df7e5e2f0b99c8d5809731246f8c5a3e2

Observation a669348c-f4e7-4781-aa37-1e9f897d1663 · outbound

This paper cites GPT-4 Technical Report.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection GPT-4 Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:41:50.693665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:6dc0b377b7e4bc514f884451d212ffa5da17409f7d5de54241896a7c9564eed6

Observation 69a3850e-4219-44da-9acd-4ea8f9816459 · outbound

This paper cites URLhttps://doi.org/10.1145/3600006.3613145.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection URLhttps://doi.org/10.1145/3600006.3613145

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T20:41:50.315045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:213e6cf73c545d5f70998f29f8fafac8b55332eabc2013096292fa5c99d37a40

Observation f09d9a6c-27fa-4ba6-8249-1960f72f5a28 · outbound

This paper cites Check- N-Run: a checkpointing system for training deep learning recommendation models.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Check- N-Run: a checkpointing system for training deep learning recommendation models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.905063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:89fd325edb4b5c249328c7a4b569fdb4510e62323204c70f20fc60f280ad7e2b

Observation a0ac690c-53b7-4fe9-a033-64c61ee2c191 · outbound

This paper cites MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:41:50.682623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:c17597498f64702a8aed1c21b8db316b5c06b8666b65881356474331cde92efc

Observation 794654ba-8e13-4695-8cb7-ba2f66719e9e · outbound

This paper cites Elan: Towards generic and efficient elastic training for deep learning.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Elan: Towards generic and efficient elastic training for deep learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.901286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:cf6d30a0b4cf5ac4f5581443cb712ccc7d8a48d6035309d9c68764e47c67f81e

Observation d4297cb2-d11b-4e9f-9cef-3bd8fc5418ae · outbound

This paper cites Bamboo: Making preemptible instances resilient for affordable training of large{DNNs}.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Bamboo: Making preemptible instances resilient for affordable training of large{DNNs}

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.897817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:86cee22a0f2108431aef2fde0acf92db5415a33d9dddf842788b759e7493eff0

Observation 67f52b03-5938-475c-ace7-616fbdcd2d7f · outbound

This paper cites Oobleck: Resilient distributed training of large models using pipeline templates.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Oobleck: Resilient distributed training of large models using pipeline templates

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.894134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:9641943772e6e49764880ea4ac83e5174e41256d58dbbd3ed6b09cebbe01ff01

Observation 29c021fc-a986-49b2-aaa3-2a54e214225a · outbound

This paper cites Recycle: Re- silient training of large dnns using pipeline adaptation.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Recycle: Re- silient training of large dnns using pipeline adaptation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.890778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:8f0ec743d09f4433e5fa578a5011743fd12c7a4ef0f7c34217454e407a79d795

Observation 6f353f01-72b2-41cf-9dcf-189c503483a4 · outbound

This paper cites Tenplex: Dynamic parallelism for deep learning using parallelizable tensor col- lections.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Tenplex: Dynamic parallelism for deep learning using parallelizable tensor col- lections

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.887146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:66099cdf2521e6c5cada53f5619796c45b804dc9a5906da75665f60660340919

Observation d9b38c0b-7c6c-4324-81aa-9e6315f0da30 · outbound

This paper cites Ascend: a scalable and unified architecture for ubiquitous deep neural network computing: Industry track paper.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Ascend: a scalable and unified architecture for ubiquitous deep neural network computing: Industry track paper

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.883628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:9a1034d2b136c2605e25f549e86dbf252edd249e9df9df039f1c41a83f50808c

Observation f5908091-346d-4114-b213-3a3c3abb219f · outbound

This paper cites Parallel scan on ascend ai accelerators.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Parallel scan on ascend ai accelerators

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:41:50.666848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:6d89816483e488fa7dfcc3070b6cee1e6859de80f1f140fb5539bf90d44dee74

Observation 710d63b9-f237-49f7-aa67-04ff0cbfa936 · outbound

This paper cites Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:41:50.651257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:6f2bac7638f0f6b4c87cb26ef6cb692ba5526938a7ab01933577d93ac83c0582

Observation 23a3272d-7a27-41ee-b508-837b62afecac · outbound

This paper cites Horovod: fast and easy distributed deep learning in TensorFlow.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Horovod: fast and easy distributed deep learning in TensorFlow

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:41:50.641241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:13b5e4b4182f8ed10300c96b5079dabd2643c3a96eac61284909603ba66d8deb

Observation 31891851-6ccf-4e31-85e2-c3352e6269df · outbound

This paper cites ImageNet Training in Minutes.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection ImageNet Training in Minutes

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:41:50.630641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:3e29a599b589f7195f06796f097ef63c40a7dff45797f3351e82701ef858efd3

Observation 3356bce0-fe8f-4732-ac13-4893bee81912 · outbound

This paper cites Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:41:50.646394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:bb261bebc7df59a03c7925fc7891a3d47531851db0d53ac6f882d6cc540f40e9

Observation 2c2149b2-0772-49b8-8097-ef3a26bf5d0c · outbound

This paper cites GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:41:50.677249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:c4dd5a28dad78f9ea578d73b0f3e13b832a101e0daa9ecf55ef6ec3948e1654c

Observation e77b0450-8f64-4cc5-b1fd-2ff4c85a717e · outbound

This paper cites BPipe: Memory-balanced pipeline parallelism for training large language models.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection BPipe: Memory-balanced pipeline parallelism for training large language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.879787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:4d9c79d81132d08bd25a28034af689aa113c07bc9a032340752a8fcfecd1d566

Observation d77e79dd-3b5b-4704-b993-0431996b3dbc · outbound

This paper cites ZeRO: Memory Optimizations Toward Training Trillion Parameter Models.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection ZeRO: Memory Optimizations Toward Training Trillion Parameter Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:41:50.656090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:e1a1bc1ef569a0ecaa2cd4c1ca8dadf493fee94e7d5d6c1129f7d976d6000520

Observation 42bc76e4-b3cf-475d-bf2b-430ed40f5930 · outbound

This paper cites Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:41:50.619868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:a9ecec2e271624ae76a481d4af051f19e3f20c22b02a732ff92a2fdd1804cbab

Observation 863dbd24-f1b2-4d47-8830-6895e417b805 · outbound

This paper cites EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:41:50.608867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:fb08fb7c5269ce688ba2b8910ce6b78a994e78c456e2db1ec254bdefa014202a

Observation 2944bfda-781b-43ad-b12b-2891759dd942 · outbound

This paper cites Moe parallel folding: Heterogeneous parallelism mappings for efficient large-scale moe model training with megatron core.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Moe parallel folding: Heterogeneous parallelism mappings for efficient large-scale moe model training with megatron core

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:41:50.614182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:bc326faf6d9902eed21597de1103a5f705ad7f9c40ec4378220b04e7d4cc696d

Observation c0b33fbc-ced1-4f00-a407-01b95eaefe16 · outbound

This paper cites Switch transformers: scaling to trillion parameter models with simple and efficient sparsity.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Switch transformers: scaling to trillion parameter models with simple and efficient sparsity

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.876089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:4afad092defba2f68a14304ce9ef28369e702101af1685d24cfc9facf2e3053b

Observation ea35a451-8a4a-4010-99dc-292664997f1d · outbound

This paper cites Tutel: Adaptive Mixture-of-Experts at Scale.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Tutel: Adaptive Mixture-of-Experts at Scale

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:41:50.636452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:f703f567d6d713b8298231109d0edcef632c27a25a3a9669e6f504796ef31f70

Observation 8c68d54a-5f6f-4055-8ceb-25edab022570 · outbound

This paper cites Megascale-moe: Large-scale communication-efficient training of mixture-of-experts models in production.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Megascale-moe: Large-scale communication-efficient training of mixture-of-experts models in production

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:41:50.625516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:654039c5ca14f94a53ee9801117617bde0b97b630fa5c878c01b8add2cec4299

Observation 2e840f52-915e-4557-b1ca-4001fd5428a6 · outbound

This paper cites Understanding communication characteristics of distributed training.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Understanding communication characteristics of distributed training

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:41:50.308004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:fa65251c906e8a310cc9e47bfed42f9e73264e88a4f65d2de7152e043d227b40

Observation 8e0a46fb-5b31-474c-bba2-271646555743 · outbound

This paper cites Amped: An analytical model for performance in distributed training of transformers.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Amped: An analytical model for performance in distributed training of transformers

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.872686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:ad5d50f313ecae97b44032b576ea38c63b389e1b345fa2a1e354d976827e6305

Observation 2310f253-5f7a-4c4f-b35c-c2b3c75cacbb · outbound

This paper cites Reducing Activation Recomputation in Large Transformer Models.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Reducing Activation Recomputation in Large Transformer Models

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T20:41:50.661629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:4fc4987997afdbf145922f47391c63fe1b9d2e14063fac05bbc9e473583bd7cd

Observation 19f3c7a3-a5bf-47ea-b116-53ad995d2a1d · outbound

This paper cites Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , articleno =.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , articleno =

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T20:41:50.295300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:7d24313030f1f60278216008c0b0ab4a730e3bdb0344ab13abbc152320dab4d8

Observation f6dbc7bb-df01-4242-a69f-09c725584504 · outbound

This paper cites Pytorch distributed: experiences on accelerating data parallel training.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Pytorch distributed: experiences on accelerating data parallel training

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:41:50.301514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:a4e3b84d9de871047dde708f9780fadaaaa2d37eaaa8e5bb4b3162ef92d75bd3

Observation b240e950-ad48-452e-9dcb-e6e8d4cee2eb · outbound

This paper cites Varuna: scalable, low-cost training of massive deep learning models.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Varuna: scalable, low-cost training of massive deep learning models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.869630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:9463fa8567c04b3c621fceb0ad8077d070e900319223ce05c6c6478fc61ff364

Observation 7fea385a-ea6b-4e33-bc6e-a204e02792b8 · outbound

This paper cites Failures in large scale systems: Long-term measurement, analysis, and implications.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Failures in large scale systems: Long-term measurement, analysis, and implications

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.866624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:ddc5b891512895607dc59460113491e9399864f410d1ba5401e3e6bfad6df79b

Observation ff22b037-5091-495f-a53b-7b2cc18184a2 · outbound

This paper cites MLaaS in the wild: Workload analysis and scheduling in Large-Scale heterogeneous GPU clusters.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection MLaaS in the wild: Workload analysis and scheduling in Large-Scale heterogeneous GPU clusters

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.863489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:615b0f7264cf1dac521ccc0210fdb5329481b15dcd7a78fc320d02e82bd4152c

Observation 93b186d6-d0e8-4f84-97e9-d482ff7898e2 · outbound

This paper cites Minder: Faulty machine detection for large-scale distributed model training.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Minder: Faulty machine detection for large-scale distributed model training

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.860249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:fcfd98a8419174530c699d9ccc9db2798f0987ae6eaa08d8a7f96c9bd6dc4eff

Observation babf272d-7115-4c67-8df3-a0bcf6e60010 · outbound

This paper cites Alpa: Automating inter- and Intra-Operator parallelism for distributed deep learning.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Alpa: Automating inter- and Intra-Operator parallelism for distributed deep learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.856876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:14f5615ee11433b5e3f61046a5e8ebb89b85b84e7155733c86b11660e6c241de

Observation 492049b6-3fee-44ae-8f88-c66af5985182 · outbound

This paper cites The hungarian method for the assignment problem.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection The hungarian method for the assignment problem

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.853505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:ea8b2a7737fd13bc62b1da3be4ab7cca7b7dd199bb0b98b00e198c542f81ca5b

Observation 6b00ff77-1b97-4770-ae2a-8655b4f723cd · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:41:50.671674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:5167ea5a2e23a0b00bf0e12e2f1b36695a2c9ae6ef4f639569490a02bfb0e705

Pith citing papers

Observation 9d0e2fe2-aee7-4187-91f1-908239989285 · inbound

Enhancing OLAP Resilience at LinkedIn cites this paper.

Enhancing OLAP Resilience at LinkedIn Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection

Reference 71

Resolution
unresolved
no resolver link, observed 2026-07-15T13:17:47.603967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:17:47.603967Z digest=sha256:f84df061ff7d9afa9f71128dfd46d7aae5c31fe4aadac74df731fb854082e099