Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:55:15.447206Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 3 inbound Pith citation observations for arXiv:2506.01260.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:55:15.447206Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-28T18:53:47.187437Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-28T19:42:36.073867Z
63 of 63 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 592c6e93-48c5-42ed-a329-6175452b8bf2 · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Transformers learn through gradual rank increase
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c64f9b57-fbda-4ae6-9d11-f765a78e9e18 · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Qsgd: Communication-efficient sgd via gradient quantization and encoding
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d9f4a23d-1c97-4aff-8476-1549f87bd169 · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Dissecting adam: The sign, magnitude and variance of stochastic gradients
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b0970cf0-a55e-4c8e-bd80-4d94511bc6d9 · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism signsgd: Compressed optimisation for non-convex problems
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f0b24b53-ce89-4eeb-8b57-0b237e83940b · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1add04da-defb-470c-90cd-02b6b5ca622c · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Does compressing activations help model parallel training? Proceedings of Machine Learning and Systems, 6: 0 239--252, 2024
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c8e9404a-916d-4c17-804c-81fdebbac5f7 · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Low-rank gradient descent
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 68b31344-16d3-455d-bc0d-ca08bcbba0c4 · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism DeepSeek LLMs , 2023
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4b22ed7a-9bb6-4124-bbc1-0c5aaec5abd2 · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism A Simple Convergence Proof of Adam and Adagrad
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f320d63e-c91c-46e4-9a45-535ee198c85f · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism 8-bit Optimizers via Block-wise Quantization
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ca8007e-5add-4001-a83b-f107719004ae · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Distributed deep learning in open collaborations
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f036de15-21bb-48ec-810e-d27ca7cb503e · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Attention is not all you need: Pure attention loses rank doubly exponentially with depth
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6e906f3c-ff5a-495d-ab76-86472ce8777e · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism DiLoCo: Distributed Low-Communication Training of Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 657f8464-4b43-4ef3-a757-c4b350906cce · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism The Llama 3 Herd of Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1517c73-b813-4383-9d2b-c86d6c3edb31 · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Openwebtext corpus
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4dd99cc9-a68d-4bba-b2a7-7be5f5266f9b · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Gradient Descent Happens in a Tiny Subspace
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af20aac7-bd93-48cd-8526-2d1f2499c4b3 · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Training compute-optimal large language models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 683d281f-56b2-437a-8a1e-24729c94ee26 · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Gpipe: Efficient training of giant neural networks using pipeline parallelism
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4c27129-cdab-49cd-99c9-b8aa62ec37ca · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Error feedback fixes signsgd and other gradient compression schemes
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64d9c80c-1fe9-42e7-bb18-22149f1a0fc3 · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Adam: A Method for Stochastic Optimization
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9e3c9a4-ab2e-456e-aeaa-ce89aaa02670 · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Big transfer (bit): General visual representation learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 28c35edf-fce5-4f31-9f2e-739c15297eb7 · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Decentralized stochastic optimization and gossip algorithms with compressed communication
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4e2ef206-c2d0-4d33-bc94-9ab270d35c10 · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism A unified theory of decentralized sgd with changing topology and local updates
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 886cc201-60f9-4848-93b9-49a731d1a1f7 · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Imagenet classification with deep convolutional neural networks
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d731396-90cc-43ae-bf02-63c8765e4443 · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Convergence of adam under relaxed assumptions
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4cc5152f-e7a0-4997-a5f0-5e9f547c6b78 · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Learning on transformers is provable low-rank and sparse: A one-layer analysis
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 19dd66f6-69f6-4bb8-a90d-8468618ce0ff · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism PyTorch Distributed: Experiences on Accelerating Data Parallel Training
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 732b3de9-bc91-4a0d-a6ca-47b1d0f59ec7 · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0bd52271-10b8-4d50-a78c-22c4a00bb911 · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8cf2029-b266-4302-bfb0-99f8904247e0 · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Deep Gradient Compression: Reducing the Communication Bandwidth for Distributed Training
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f8f90bb-c2ea-4452-9de8-6966bae1c399 · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Adam$^+$: A Stochastic Method with Adaptive Variance Reduction
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9d2acdea-e1c8-4755-8b44-50e70a49009d · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Decoupled Weight Decay Regularization
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dc712ca-9b3e-4e6c-99e1-0292bccb2473 · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Pointer sentinel mixture models, 2016
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c77da94-f270-4761-bd07-7264e30cae9f · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Efficient large-scale language model training on gpu clusters using megatron-lm
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc818491-3547-47da-94c9-deac8862c8f2 · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Decoupled momentum optimization
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23dbdad6-0ba3-4c34-afb7-f26c30eb8305 · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Ai and compute
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 06aeac6e-0e51-469d-9d5e-ff5d00d31a4f · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Unresolved cited work
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 084b05a6-4f5a-47db-801b-fb9d6b952adf · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism PanGu-{\Sigma}: Towards Trillion Parameter Language Model with Sparse Heterogeneous Computing
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95d26b16-c41d-4c86-b777-8211317f30dd · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Activations and gradients compression for model-parallel training
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 847e1442-37ef-42f9-b5bd-b2febdfc801f · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Towards crowdsourced training of large neural networks using decentralized mixture-of-experts
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ea379800-0903-4da6-9da1-a2b0399d6d36 · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Moshpit sgd: Communication-efficient decentralized training on heterogeneous unreliable devices
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ea093586-11fc-473b-a71d-e4a9a660c855 · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Swarm parallelism: Training large models can be surprisingly communication-efficient
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 01da2141-3467-4517-86e1-04d69921597a · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Inheritune: Training smaller yet more attentive language models, 2024
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28adb1f0-5240-469f-a239-f90147d83e63 · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c5ed812-b146-4583-acd0-2f391285cda7 · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism 1-bit adam: Communication efficient large-scale training with adam’s convergence speed
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d3f47848-25af-49fb-8b0c-f25666683362 · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Scaling the summit: deploying the world’s fastest supercomputer
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3670813a-87e0-4f51-ae8f-e6a7fff0a382 · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Powersgd: Practical low-rank gradient compression for distributed optimization
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcaaf8b2-61dc-427c-bae4-f0c9bf3f473a · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Kalman Gradient Descent: Adaptive Variance Reduction in Stochastic Optimization
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6874d0b4-192d-415b-96a5-3531593fc3fa · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Variance reduction for stochastic gradient optimization
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7f84e35e-cba5-42b3-8109-1a64f96e43ab · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Atomo: Communication-efficient learning via atomic sparsification
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cc829027-62a7-422f-915c-8e12a8b7961c · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Pufferfish: Communication-efficient models at no extra cost
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2193dd3a-4213-465e-b126-d8717960a4fe · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Efficient distributed learning with sparsity
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d4b889cc-8acb-4882-82c0-a01bebc72cff · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Cocktailsgd: Fine-tuning foundation models over 500mbps networks
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bc257cec-5b28-4b92-8d2c-e30af183a949 · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Gradient sparsification for communication-efficient distributed optimization
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 74a6b4c1-4926-4340-bc05-6b093c11153f · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Error compensated quantized sgd and its applications to large-scale distributed optimization
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d054a76c-f635-49b6-b8ac-0faffc2dc425 · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism A Spectral Condition for Feature Learning
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc8965be-0e92-426b-80f8-55b176dddbf8 · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Stochastic Gradient Variance Reduction by Solving a Filtering Problem
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e20f2efd-df6b-429e-8765-f24f65a631e5 · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Decentralized training of foundation models in heterogeneous environments
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 62d9db7f-7614-4cbd-9382-b7706f10294e · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Convergence Guarantees for RMSProp and Adam in Generalized-smooth Non-convex Optimization with Affine Noise Variance
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54ddbd28-fa93-40ef-a94b-65c092a6f745 · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism ZerO Initialization: Initializing Neural Networks with only Zeros and Ones
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 529c2762-7a6b-4c2d-b08d-0186b91e35ff · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03303144-9887-4265-b371-af3dcdde92c0 · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84ae9fa8-ad4c-46f6-8881-7f811796074d · outbound
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9a3fc158-0571-435a-9652-263a421bd0b8 · inbound
On the Surprising Effectiveness of a Single Global Merging in Decentralized Learning Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5ab30c17-9f1e-48dc-b734-49a5b9295181 · inbound
ResBM: Residual Bottleneck Models for Low-Bandwidth Pipeline Parallelism Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d3d7d422-f931-43f2-85ba-457aed58aacd · inbound
GNMR: Runtime Stability Control for Low-Precision Large Language Model Training Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.