Pith. sign in

Paper Citation Record · LEDGER

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism

As of 7 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 3 inbound Pith citation observations for arXiv:2506.01260.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01260 v3

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:55:15.447206Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T18:53:47.187437Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T19:42:36.073867Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact1
  • verified fuzzy32
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 592c6e93-48c5-42ed-a329-6175452b8bf2 · outbound

This paper cites Transformers learn through gradual rank increase.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Transformers learn through gradual rank increase

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:19.678630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:55:12.762172Z digest=sha256:510cd6c87bcda291e1b6fbbfda287a3ee3bd4d98678b068377c72052220b513a

Observation c64f9b57-fbda-4ae6-9d11-f765a78e9e18 · outbound

This paper cites Qsgd: Communication-efficient sgd via gradient quantization and encoding.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Qsgd: Communication-efficient sgd via gradient quantization and encoding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:19.504881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:55:12.792678Z digest=sha256:d44a5ee933b3a7409309e20ba231fa354d7181521434a7edf9591aa4a1f3fc55

Observation d9f4a23d-1c97-4aff-8476-1549f87bd169 · outbound

This paper cites Dissecting adam: The sign, magnitude and variance of stochastic gradients.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Dissecting adam: The sign, magnitude and variance of stochastic gradients

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:19.330916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:55:12.887755Z digest=sha256:89a42e689592e14aa9b96700a75805cdd6d05d92218b8bdb0d99a187f5204ff5

Observation b0970cf0-a55e-4c8e-bd80-4d94511bc6d9 · outbound

This paper cites signsgd: Compressed optimisation for non-convex problems.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism signsgd: Compressed optimisation for non-convex problems

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:19.199270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:55:12.938754Z digest=sha256:ded865e21bd2470259cdabf79c13cfd9bb5bda670328f1e3a488fcfe87340248

Observation f0b24b53-ce89-4eeb-8b57-0b237e83940b · outbound

This paper cites Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.978236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:12.978236Z digest=sha256:d5e9b8a989cca373c95598f39fd9d5dde18efddd882c3d8477f343c984d9c0b9

Observation 1add04da-defb-470c-90cd-02b6b5ca622c · outbound

This paper cites Does compressing activations help model parallel training? Proceedings of Machine Learning and Systems, 6: 0 239--252, 2024.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Does compressing activations help model parallel training? Proceedings of Machine Learning and Systems, 6: 0 239--252, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:19.045514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:55:13.031459Z digest=sha256:7f14b242b84fc8715578627f09ea1671f594025fa82c5881ba288473c836cc16

Observation c8e9404a-916d-4c17-804c-81fdebbac5f7 · outbound

This paper cites Low-rank gradient descent.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Low-rank gradient descent

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:18.889782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:55:13.092064Z digest=sha256:996387102d74d94b1a533f8f97040fb5cd039bdaae412427e0dccedcfb27ec93

Observation 68b31344-16d3-455d-bc0d-ca08bcbba0c4 · outbound

This paper cites DeepSeek LLMs , 2023.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism DeepSeek LLMs , 2023

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:18.694587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:55:13.159428Z digest=sha256:472cd9890716cb6a2791218048b02c290ff0d39cd6bc0ef52e656c17e64a0d74

Observation 4b22ed7a-9bb6-4124-bbc1-0c5aaec5abd2 · outbound

This paper cites A Simple Convergence Proof of Adam and Adagrad.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism A Simple Convergence Proof of Adam and Adagrad

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.202772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:13.202772Z digest=sha256:4b7ba85ce43c09efca9d4ff556e3754f198e9be89e372ead948ec0c291a83d5b

Observation f320d63e-c91c-46e4-9a45-535ee198c85f · outbound

This paper cites 8-bit Optimizers via Block-wise Quantization.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism 8-bit Optimizers via Block-wise Quantization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.250462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:13.250462Z digest=sha256:5071c1c8dac413fcf59942a0d76dcd190cee2295da0a1ddea6109eacf3b281fa

Observation 0ca8007e-5add-4001-a83b-f107719004ae · outbound

This paper cites Distributed deep learning in open collaborations.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Distributed deep learning in open collaborations

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:18.526941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:55:13.281551Z digest=sha256:dcf2cec9e3e0e83aecd3a957a46ece1c180845eda2e0f20cc49f6be2639a94f0

Observation f036de15-21bb-48ec-810e-d27ca7cb503e · outbound

This paper cites Attention is not all you need: Pure attention loses rank doubly exponentially with depth.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Attention is not all you need: Pure attention loses rank doubly exponentially with depth

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:18.456406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:55:13.326951Z digest=sha256:2228a09457ff4a5ffd48e785bffad81e0700b11d41754b873cdd2f486921f662

Observation 6e906f3c-ff5a-495d-ab76-86472ce8777e · outbound

This paper cites DiLoCo: Distributed Low-Communication Training of Language Models.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism DiLoCo: Distributed Low-Communication Training of Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.360801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:13.360801Z digest=sha256:8da918ea6f70455e69730d9be7b30ada91168fd74a4bdbcd66256ae719917547

Observation 657f8464-4b43-4ef3-a757-c4b350906cce · outbound

This paper cites The Llama 3 Herd of Models.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.407145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:13.407145Z digest=sha256:a63593fb583e0d437438fa47fad2c68445e4deb2e1cf48ed4aa637407aa8f3c8

Observation e1517c73-b813-4383-9d2b-c86d6c3edb31 · outbound

This paper cites Openwebtext corpus.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Openwebtext corpus

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.449970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:13.449970Z digest=sha256:cb68fb7d84aba29e36b23ed261690720ffe30dacf20cfaa5ffdfed99748ae73f

Observation 4dd99cc9-a68d-4bba-b2a7-7be5f5266f9b · outbound

This paper cites Gradient Descent Happens in a Tiny Subspace.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Gradient Descent Happens in a Tiny Subspace

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.490228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:13.490228Z digest=sha256:aed5e3f6a039f45f9ac02073236252f8249b84add2c3b9cf9f34aec9fba1f9ac

Observation af20aac7-bd93-48cd-8526-2d1f2499c4b3 · outbound

This paper cites Training compute-optimal large language models.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Training compute-optimal large language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:18.325287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:55:13.527372Z digest=sha256:34594f2536b6a2ddc0cb2925ce0806fe23f34e8389ab332e5cde5ae5d71f8ccd

Observation 683d281f-56b2-437a-8a1e-24729c94ee26 · outbound

This paper cites Gpipe: Efficient training of giant neural networks using pipeline parallelism.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Gpipe: Efficient training of giant neural networks using pipeline parallelism

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.561151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:13.561151Z digest=sha256:0ed6b3ce2f2be471a6f4780c33586c1fb699e48cca4ca499a607f3d031e17e86

Observation b4c27129-cdab-49cd-99c9-b8aa62ec37ca · outbound

This paper cites Error feedback fixes signsgd and other gradient compression schemes.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Error feedback fixes signsgd and other gradient compression schemes

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.601856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:13.601856Z digest=sha256:1f8e4f081b91338dbfd29a30921aa5960872e02389f9fb3cc8d757e53c04aa53

Observation 64d9c80c-1fe9-42e7-bb18-22149f1a0fc3 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Adam: A Method for Stochastic Optimization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.644921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:13.644921Z digest=sha256:3d6ef40714d59dbc0448afa269c133e6b2e647faae08f950557ad8abd0ab2827

Observation c9e3c9a4-ab2e-456e-aeaa-ce89aaa02670 · outbound

This paper cites Big transfer (bit): General visual representation learning.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Big transfer (bit): General visual representation learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:18.191509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:55:13.685383Z digest=sha256:119362ee9dbcefe6949382be79a177efce2a8a39962358bc2c860405b9d16391

Observation 28c35edf-fce5-4f31-9f2e-739c15297eb7 · outbound

This paper cites Decentralized stochastic optimization and gossip algorithms with compressed communication.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Decentralized stochastic optimization and gossip algorithms with compressed communication

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:18.111648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:55:13.722281Z digest=sha256:6f2b68a31065a93814b40271d55dcaa9d85aaca67cb7a1e3e58b9378cecadf36

Observation 4e2ef206-c2d0-4d33-bc94-9ab270d35c10 · outbound

This paper cites A unified theory of decentralized sgd with changing topology and local updates.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism A unified theory of decentralized sgd with changing topology and local updates

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:18.022141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:55:13.774841Z digest=sha256:96653268798921cf8b24b3f6cc4b60c1faf9ae97244b59ccfd47cf424c7d993e

Observation 886cc201-60f9-4848-93b9-49a731d1a1f7 · outbound

This paper cites Imagenet classification with deep convolutional neural networks.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Imagenet classification with deep convolutional neural networks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.802110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:13.802110Z digest=sha256:aa6e6ed302fe372b25a331876f69f627ebfe6d5d8fba9c18fbf50990c5967abe

Observation 0d731396-90cc-43ae-bf02-63c8765e4443 · outbound

This paper cites Convergence of adam under relaxed assumptions.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Convergence of adam under relaxed assumptions

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.906167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:55:13.825919Z digest=sha256:5b4401c092de61bfd7cb18aa2124135975520c52933e5d6b9e8e1e29d0392ea4

Observation 4cc5152f-e7a0-4997-a5f0-5e9f547c6b78 · outbound

This paper cites Learning on transformers is provable low-rank and sparse: A one-layer analysis.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Learning on transformers is provable low-rank and sparse: A one-layer analysis

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.727061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:55:13.868539Z digest=sha256:50b17a51b59752149ce903ecbdf217cd9f5e8eac77c5fa9b671ab25a34219d10

Observation 19dd66f6-69f6-4bb8-a90d-8468618ce0ff · outbound

This paper cites PyTorch Distributed: Experiences on Accelerating Data Parallel Training.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism PyTorch Distributed: Experiences on Accelerating Data Parallel Training

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.906478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:13.906478Z digest=sha256:517a1ef3ab57e95650a97299dd0622576ac107e511453515a8151b0a07e73caf

Observation 732b3de9-bc91-4a0d-a6ca-47b1d0f59ec7 · outbound

This paper cites Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.641992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:55:13.949599Z digest=sha256:6d843f80bb4759a5b95bdd25faa35b530cb96933b54e95ba7a393827f6b75ce5

Observation 0bd52271-10b8-4d50-a78c-22c4a00bb911 · outbound

This paper cites TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.990201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:13.990201Z digest=sha256:f93fb0514a7b93f7d1022fc0c1d869e7ed2a7328c9ba26303ab7ff5fd378db5a

Observation e8cf2029-b266-4302-bfb0-99f8904247e0 · outbound

This paper cites Deep Gradient Compression: Reducing the Communication Bandwidth for Distributed Training.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Deep Gradient Compression: Reducing the Communication Bandwidth for Distributed Training

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.033583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:14.033583Z digest=sha256:51e6e91282a3f84c4a32318eca03ba2064fcd25c98a3ccee9e26a61bc2f0667a

Observation 6f8f90bb-c2ea-4452-9de8-6966bae1c399 · outbound

This paper cites Adam$^+$: A Stochastic Method with Adaptive Variance Reduction.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Adam$^+$: A Stochastic Method with Adaptive Variance Reduction

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:55:15.848572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:55:14.067024Z digest=sha256:fec809bb0b8fe68b21407f79e75b234b610569b2ab67edaa39e890e35d0a7e4e

Observation 9d2acdea-e1c8-4755-8b44-50e70a49009d · outbound

This paper cites Decoupled Weight Decay Regularization.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Decoupled Weight Decay Regularization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.108295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:14.108295Z digest=sha256:44f93bfcbee026ffba86edf3b463b73c30c5b43ea8175be7f0546def6a9f4f3d

Observation 0dc712ca-9b3e-4e6c-99e1-0292bccb2473 · outbound

This paper cites Pointer sentinel mixture models, 2016.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Pointer sentinel mixture models, 2016

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.137281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:14.137281Z digest=sha256:310ea2ed4cb166a0ccfc9adb272e7c831af6ba5a50e3b7cac9545a99f4807d84

Observation 8c77da94-f270-4761-bd07-7264e30cae9f · outbound

This paper cites Efficient large-scale language model training on gpu clusters using megatron-lm.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Efficient large-scale language model training on gpu clusters using megatron-lm

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.176931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:14.176931Z digest=sha256:e6cfca848e74bc6eebe1cf4f97e313c6ad1145b5c9d8726cea17053bd513c075

Observation fc818491-3547-47da-94c9-deac8862c8f2 · outbound

This paper cites Decoupled momentum optimization.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Decoupled momentum optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.217550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:14.217550Z digest=sha256:733628aa8256982a07e180684ab46d11fe852b6bcc41be52e9a3912446e79158

Observation 23dbdad6-0ba3-4c34-afb7-f26c30eb8305 · outbound

This paper cites Ai and compute.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Ai and compute

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.532560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:55:14.258302Z digest=sha256:54cd6d464bf63960f8764c4aeac965578c9100d454697d4be92071abc8d15e05

Observation 06aeac6e-0e51-469d-9d5e-ff5d00d31a4f · outbound

This paper cites an unresolved cited work.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.298139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:14.298139Z digest=sha256:4ac1e2043fc497c777dc359b088b455163d6066e3707ea28f6ba9a86be296e70

Observation 084b05a6-4f5a-47db-801b-fb9d6b952adf · outbound

This paper cites PanGu-{\Sigma}: Towards Trillion Parameter Language Model with Sparse Heterogeneous Computing.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism PanGu-{\Sigma}: Towards Trillion Parameter Language Model with Sparse Heterogeneous Computing

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.338207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:14.338207Z digest=sha256:123a70358f15cc6420608e5b0072c5b557f1f1fbf95c0948b1b20874487a8b6f

Observation 95d26b16-c41d-4c86-b777-8211317f30dd · outbound

This paper cites Activations and gradients compression for model-parallel training.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Activations and gradients compression for model-parallel training

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.445422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:55:14.379343Z digest=sha256:6a8926e73508ca894082f39619763046bf13815964df1096d59d410d535887fa

Observation 847e1442-37ef-42f9-b5bd-b2febdfc801f · outbound

This paper cites Towards crowdsourced training of large neural networks using decentralized mixture-of-experts.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Towards crowdsourced training of large neural networks using decentralized mixture-of-experts

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.356432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:55:14.444090Z digest=sha256:3c3d734ed4af3302a03cab88d22c832d0eabd58042c81928ba8d36837e67c9df

Observation ea379800-0903-4da6-9da1-a2b0399d6d36 · outbound

This paper cites Moshpit sgd: Communication-efficient decentralized training on heterogeneous unreliable devices.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Moshpit sgd: Communication-efficient decentralized training on heterogeneous unreliable devices

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.271867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:55:14.490832Z digest=sha256:41accd36d7af7e20cc34f158b0c104b6571149e34438541e4a29c8851715170a

Observation ea093586-11fc-473b-a71d-e4a9a660c855 · outbound

This paper cites Swarm parallelism: Training large models can be surprisingly communication-efficient.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Swarm parallelism: Training large models can be surprisingly communication-efficient

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.187559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:55:14.539044Z digest=sha256:ab3d6603260246656b75276e5be3e549edd9947441c4ed5b1a3841c207aaa169

Observation 01da2141-3467-4517-86e1-04d69921597a · outbound

This paper cites Inheritune: Training smaller yet more attentive language models, 2024.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Inheritune: Training smaller yet more attentive language models, 2024

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.585859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:14.585859Z digest=sha256:6b1086e0901038dc77a2462fce20cd886919b1d47b88506a0586f569bf182f0d

Observation 28adb1f0-5240-469f-a239-f90147d83e63 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.633452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:14.633452Z digest=sha256:21388c24050949e5146fd05f03493135434ca0bf6688b2460eee7eabfd8e27da

Observation 0c5ed812-b146-4583-acd0-2f391285cda7 · outbound

This paper cites 1-bit adam: Communication efficient large-scale training with adam’s convergence speed.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism 1-bit adam: Communication efficient large-scale training with adam’s convergence speed

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.120045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:55:14.675721Z digest=sha256:70f8313dac25049cf94d627020cf485cf164ae92b6bd11fd72b8384ee3691ad7

Observation d3f47848-25af-49fb-8b0c-f25666683362 · outbound

This paper cites Scaling the summit: deploying the world’s fastest supercomputer.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Scaling the summit: deploying the world’s fastest supercomputer

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.055192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:55:14.729222Z digest=sha256:cc467d9532493bbda42b80eea9c6b5f9ad23530500362dce4915b4aa17d0d05d

Observation 3670813a-87e0-4f51-ae8f-e6a7fff0a382 · outbound

This paper cites Powersgd: Practical low-rank gradient compression for distributed optimization.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Powersgd: Practical low-rank gradient compression for distributed optimization

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.772989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:14.772989Z digest=sha256:550bde7ae31de4cddd2c72acc56ff1193c93945a070a26a4baf40a9a4d482747

Observation dcaaf8b2-61dc-427c-bae4-f0c9bf3f473a · outbound

This paper cites Kalman Gradient Descent: Adaptive Variance Reduction in Stochastic Optimization.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Kalman Gradient Descent: Adaptive Variance Reduction in Stochastic Optimization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.838055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:14.838055Z digest=sha256:e6b75408072c7e61c9c8d20b61009dff6723843bc28f652d70dffaceb920eccf

Observation 6874d0b4-192d-415b-96a5-3531593fc3fa · outbound

This paper cites Variance reduction for stochastic gradient optimization.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Variance reduction for stochastic gradient optimization

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:16.949401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:55:14.864069Z digest=sha256:c09c154efecec2edf0e74803aa8df1ded3f2e5bbfe49fca835b3d87c4e582d59

Observation 7f84e35e-cba5-42b3-8109-1a64f96e43ab · outbound

This paper cites Atomo: Communication-efficient learning via atomic sparsification.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Atomo: Communication-efficient learning via atomic sparsification

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:16.890495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:55:14.902786Z digest=sha256:984830a57054d3e7d22383be78321a8f3094888e7f2f540a49202723e436408a

Observation cc829027-62a7-422f-915c-8e12a8b7961c · outbound

This paper cites Pufferfish: Communication-efficient models at no extra cost.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Pufferfish: Communication-efficient models at no extra cost

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:16.804411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:55:14.942915Z digest=sha256:2f8877fb15d5837038f1b622a86c83b915acf426033e6c588ba7220737527c6d

Observation 2193dd3a-4213-465e-b126-d8717960a4fe · outbound

This paper cites Efficient distributed learning with sparsity.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Efficient distributed learning with sparsity

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:16.716469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:55:15.004690Z digest=sha256:ad68ef5955e3feee2d24d5b892b31b3cd327d8d1f8d3b28fa90847057690e01a

Observation d4b889cc-8acb-4882-82c0-a01bebc72cff · outbound

This paper cites Cocktailsgd: Fine-tuning foundation models over 500mbps networks.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Cocktailsgd: Fine-tuning foundation models over 500mbps networks

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:16.633501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:55:15.041815Z digest=sha256:79cd8d19b3c2175db33d071d201bc8a78ee1912035565152f4eb03531cb4d252

Observation bc257cec-5b28-4b92-8d2c-e30af183a949 · outbound

This paper cites Gradient sparsification for communication-efficient distributed optimization.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Gradient sparsification for communication-efficient distributed optimization

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:16.537448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:55:15.092180Z digest=sha256:793d6fec68c3e6483d6f1c45fdd634a7809c1b5ba3f6aff3a76d69ff00efee0b

Observation 74a6b4c1-4926-4340-bc05-6b093c11153f · outbound

This paper cites Error compensated quantized sgd and its applications to large-scale distributed optimization.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Error compensated quantized sgd and its applications to large-scale distributed optimization

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:16.430575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:55:15.139453Z digest=sha256:8879be9fea8019a32280c3362858264c0e9db1511199cff16c4802700b52bfe4

Observation d054a76c-f635-49b6-b8ac-0faffc2dc425 · outbound

This paper cites A Spectral Condition for Feature Learning.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism A Spectral Condition for Feature Learning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:15.171300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:15.171300Z digest=sha256:e55f3f9da1a6d519277e529425965f6d19194a7108e3204cf05aba6bb6b583f1

Observation fc8965be-0e92-426b-80f8-55b176dddbf8 · outbound

This paper cites Stochastic Gradient Variance Reduction by Solving a Filtering Problem.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Stochastic Gradient Variance Reduction by Solving a Filtering Problem

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:15.205516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:15.205516Z digest=sha256:eb145111fad12985217bf2b72c9862a7b2c9f3dfc6aab7bb28e7b195135f2407

Observation e20f2efd-df6b-429e-8765-f24f65a631e5 · outbound

This paper cites Decentralized training of foundation models in heterogeneous environments.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Decentralized training of foundation models in heterogeneous environments

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:16.241381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:55:15.244018Z digest=sha256:9ef82a38c7b03069dbfe3c610ed3a84f2cf42aace6b8f225cfad157945755398

Observation 62d9db7f-7614-4cbd-9382-b7706f10294e · outbound

This paper cites Convergence Guarantees for RMSProp and Adam in Generalized-smooth Non-convex Optimization with Affine Noise Variance.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Convergence Guarantees for RMSProp and Adam in Generalized-smooth Non-convex Optimization with Affine Noise Variance

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:15.284168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:15.284168Z digest=sha256:596f660803fc6c18deab0c510ec03d28c490e9009014861f54da42cc90892278

Observation 54ddbd28-fa93-40ef-a94b-65c092a6f745 · outbound

This paper cites ZerO Initialization: Initializing Neural Networks with only Zeros and Ones.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism ZerO Initialization: Initializing Neural Networks with only Zeros and Ones

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:15.314868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:15.314868Z digest=sha256:2fe40b5b8d21d0d210bc75929d3ed4d025f04f574b9c824ec3ab7b03a47c9568

Observation 529c2762-7a6b-4c2d-b08d-0186b91e35ff · outbound

This paper cites GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:15.366875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:15.366875Z digest=sha256:c1d8498ce9e821b881b86dad77db5f79d308ac0b01448db4869ca311e039a5da

Observation 03303144-9887-4265-b371-af3dcdde92c0 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:15.412077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:15.412077Z digest=sha256:1a646719aeac54ff63bbca70088f77a689d356c3c0f228bdd32982ecc91c8bca

Observation 84ae9fa8-ad4c-46f6-8881-7f811796074d · outbound

This paper cites Aligning books and movies: Towards story-like visual explanations by watching movies and reading books.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism Aligning books and movies: Towards story-like visual explanations by watching movies and reading books

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:16.124732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:55:15.447206Z digest=sha256:4fc8a3bf252f8ab52a76e50a22ab94280205943a4a79653eec4015ba4de92ed5

Pith citing papers

Observation 9a3fc158-0571-435a-9652-263a421bd0b8 · inbound

On the Surprising Effectiveness of a Single Global Merging in Decentralized Learning cites this paper.

On the Surprising Effectiveness of a Single Global Merging in Decentralized Learning Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-07-09T00:19:35.965961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T05:39:53.088948Z digest=sha256:61b004cf1aaca55075fee4e2d0af9241bc53a39a4bb30349c9545763f8533a9a

Observation 5ab30c17-9f1e-48dc-b734-49a5b9295181 · inbound

ResBM: Residual Bottleneck Models for Low-Bandwidth Pipeline Parallelism cites this paper.

ResBM: Residual Bottleneck Models for Low-Bandwidth Pipeline Parallelism Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-09T00:19:35.965961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:36:15.760164Z digest=sha256:fca8f2afed3d935ddb11783f53e802b393ba22e9816e15dcb2ed88274bc7f6d0

Observation d3d7d422-f931-43f2-85ba-457aed58aacd · inbound

GNMR: Runtime Stability Control for Low-Precision Large Language Model Training cites this paper.

GNMR: Runtime Stability Control for Low-Precision Large Language Model Training Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T00:19:35.965961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T18:53:47.187437Z digest=sha256:8b816f7806b7f24622e1bdc1662204fdd33c7927c3d5ccd93aecc3b413cd3b6e