Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:15:28.228455Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2506.22175.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:15:28.228455Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
35 of 35 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 80e302e1-70b3-48e4-9fb2-efa3d16c2011 · outbound
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism On the optimization of deep networks: Implicit acceleration by overparameterization,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation da7e2b1f-e19f-4471-8b74-013bcc6e3fac · outbound
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Exploring the limits of weakly supervised pretraining,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c0832550-adb5-4c73-999e-1b1107862a00 · outbound
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Antman: Dynamic scaling on gpu clusters for deep learning,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b4e4428f-2f87-4a99-a327-3399924419b9 · outbound
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Whale: Efficient giant model training over heterogeneous gpus,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4a45c5aa-fed0-4ad2-903a-aa8ea74b4ac7 · outbound
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Axonn: An asynchronous, message-driven parallel framework for extreme-scale deep learning,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 202fd7a0-346f-4d78-b43d-6d7a2e0321a8 · outbound
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism An efficient and non-intrusive gpu schedul- ing framework for deep learning training systems,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f4fe10b9-0f1b-4450-9e60-b23b2571da6f · outbound
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Bert: Pre-training of deep bidirectional transformers for language understanding,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ace6e3a0-ef11-45f4-a95e-78859d2aefd2 · outbound
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism RoBERTa: A Robustly Optimized BERT Pretraining Approach
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ea1faec-ca1e-4b18-8d4d-900f39b7ebea · outbound
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Exploring the limits of transfer learning with a unified text-to-text transformer
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c0368665-aa1a-4d71-a103-10d3923d409e · outbound
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Language mod- els are few-shot learners,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e9724b0-cf24-495c-a834-667601713bd7 · outbound
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism BLISS: Robust Sequence-to-Sequence Learning via Self-Supervised Input Representation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9129e399-e0bb-4626-8d95-36dbff9f54fe · outbound
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism E2S2: Encoding-Enhanced Sequence-to-Sequence Pretraining for Language Understanding and Generation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a588efc2-b4fd-4dd4-8bdc-8dc75e06d433 · outbound
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Unsu- pervised cross-lingual representation learning at scale,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c51cf98b-e571-425f-bbf9-cb073dfb0325 · outbound
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Toward Efficient Language Model Pretraining and Downstream Adaptation via Self-Evolution: A Case Study on SuperGLUE
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45b1a33d-616a-487a-bccb-fb8ccd773284 · outbound
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e365e47-038f-4936-851c-975894fcde0d · outbound
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25121131-c2c5-4c52-98d2-7cfd717c1894 · outbound
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 52d7527b-358b-4360-af2c-cbe2b11d13a6 · outbound
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism PAD-Net: An Efficient Framework for Dynamic Networks
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d1ab2883-4be9-4ebd-8bcc-18f2bfd914b5 · outbound
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Base layers: Simplifying training of large, sparse models,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 81121fdc-711e-4918-8173-af010a61e654 · outbound
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Gating dropout: Communication-efficient regularization for sparsely activated transform- ers,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f5762c01-f99e-42bd-903b-cb895bac91df · outbound
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Deepspeed-moe: Advancing mixture-of-experts inference and training to power next-generation AI scale,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 10ad3246-ce10-4f75-8cf0-b201c7eae5b4 · outbound
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Fastermoe: modeling and optimizing training of large-scale dynamic pre-trained models,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation da501c63-0f1f-4977-a420-cf8acc594f93 · outbound
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Scalable distributed dl training: Batching communication and computation,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 77bc7860-4971-447f-867d-c7965f20e622 · outbound
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Zero: Memory optimizations toward training trillion parameter models,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 89d2c24a-c7ce-4c6e-b494-0c6e08c1f0db · outbound
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Scalable and Efficient MoE Training for Multitask Multilingual Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3157efc-d1e2-4f18-928f-81e7d59f0558 · outbound
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Training Deep Nets with Sublinear Memory Cost
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61131dc7-3ae8-4b15-b543-4815bf7e6cda · outbound
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism vdnn: Virtualized deep neural networks for scalable, memory-efficient neural network design,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 954749e8-cf15-4840-8bcd-07b8fddcad2b · outbound
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Buddy compression: Enabling larger memory for deep learning and hpc workloads on gpus,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0637980d-29df-4b1c-b823-c91db9e0624f · outbound
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Efficient large-scale language model training on gpu clusters using megatron-lm,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5f77f9e4-68da-46a4-9a5b-fe15d0db598c · outbound
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Adam: A method for stochastic optimization,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1df6a646-9d37-49d9-94b6-61dbb4a70cfb · outbound
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Gpipe: Efficient training of giant neu- ral networks using pipeline parallelism,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc6a1c04-8cc7-41c1-9dfd-bd462064a2e0 · outbound
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Tutel: Adaptive Mixture-of-Experts at Scale
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74773f93-26a2-4e1a-bccc-250a63346df8 · outbound
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Mesh-tensorflow: Deep learning for supercomputers,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0cc0be77-5179-45f1-b922-db552ce61723 · outbound
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcc388e8-10f4-48cc-a89d-78c841b56b78 · outbound
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism PipeDream: Fast and Efficient Pipeline Parallel DNN Training
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.