Pith. sign in

Paper Citation Record · LEDGER

DiLoCo: Distributed Low-Communication Training of Language Models

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 46 inbound Pith citation observations for arXiv:2311.08105.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.08105 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 46 of 46 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:39:02.026235Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

7
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0a63f94a-d0b3-4149-871b-8cff7a9af167 · inbound

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training cites this paper.

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training DiLoCo: Distributed Low-Communication Training of Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:08.191511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:57:08.191511Z digest=sha256:fdfc6694c52f1231d62a454d20305460c958414fc1b0108bbec7727aa8ba9e07

Observation 1cf412a7-5472-452d-a3cd-5238a62f440d · inbound

INTELLECT-1 Technical Report cites this paper.

INTELLECT-1 Technical Report DiLoCo: Distributed Low-Communication Training of Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:26.378365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:44:26.378365Z digest=sha256:70f1d6d904d4d5dc7ef233ed12b01d5722ccdb153a98e70d6019b196e6197055

Observation aa79819e-02f6-440b-a216-0172499fa719 · inbound

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models cites this paper.

EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models DiLoCo: Distributed Low-Communication Training of Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:50.746824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:01:50.746824Z digest=sha256:af9811228822e233256086a7dd117f680bf02b7c71e3a3b32cef6929ee9696b4

Observation 1e12dd8c-6f6b-409e-a0db-f0d1cc01789f · inbound

Decentralized Diffusion Models cites this paper.

Decentralized Diffusion Models DiLoCo: Distributed Low-Communication Training of Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T21:17:58.310968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:17:58.310968Z digest=sha256:73546a735547f2fb74a2cadd2e99be4d7068b039675f005dbfa337901f66eb28

Observation 71a652a7-53d8-4c60-b46b-7335b1675d41 · inbound

Cross-region Model Training with Communication-Computation Overlapping and Delay Compensation cites this paper.

Cross-region Model Training with Communication-Computation Overlapping and Delay Compensation DiLoCo: Distributed Low-Communication Training of Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T10:39:02.026235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:39:02.026235Z digest=sha256:ff47a891d0e832bfcb53d887477a312fe1a67abcba2ef1e44fbe967a03e6bbf1

Observation 15ad925d-b52d-460d-877d-99fbcae9f9fa · inbound

Prime Collective Communications Library -- Technical Report cites this paper.

Prime Collective Communications Library -- Technical Report DiLoCo: Distributed Low-Communication Training of Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:45:00.694739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:45:00.694739Z digest=sha256:5a50f447ca1bc44f024d0e6c2f891ea9edc579968d9140b4abba26f4c4336dfd

Observation 071f1e30-5c6f-4492-bf64-3c50eca7c604 · inbound

Incentivizing Permissionless Distributed Learning of LLMs cites this paper.

Incentivizing Permissionless Distributed Learning of LLMs DiLoCo: Distributed Low-Communication Training of Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:19.164843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:19.164843Z digest=sha256:170de6539cd4a31f018f93c70ee67603a9b9a864cc4fe206a398070a6e956c4a

Observation ec5ff9dd-d323-429c-b9ae-5d73bd6acdd0 · inbound

DES-LOC: Desynced Low Communication Adaptive Optimizers for Training Foundation Models cites this paper.

DES-LOC: Desynced Low Communication Adaptive Optimizers for Training Foundation Models DiLoCo: Distributed Low-Communication Training of Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:21.194758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:21.194758Z digest=sha256:4499401d67b906003dafc419b69679ecb8c8493068c5b0f74651a27accf2c4bd

Observation 4449a919-5659-43e7-bf23-9d0dc44eb9a2 · inbound

MuLoCo: Muon is a practical inner optimizer for DiLoCo cites this paper.

MuLoCo: Muon is a practical inner optimizer for DiLoCo DiLoCo: Distributed Low-Communication Training of Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:26.610592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:26.610592Z digest=sha256:0499622d7305d0467043313ef622177a629129efbe662de317e634740b77e961

Observation 6e906f3c-ff5a-495d-ab76-86472ce8777e · inbound

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism cites this paper.

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism DiLoCo: Distributed Low-Communication Training of Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.360801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:13.360801Z digest=sha256:601607fea7f79c872f235ebf07678115e810be990954399bc3af71650d13cd6b

Observation 76f073f6-21b1-4e85-aabd-74f486939cb4 · inbound

PC-MoE: Memory-Efficient and Privacy-Preserving Collaborative Training for Mixture-of-Experts LLMs cites this paper.

PC-MoE: Memory-Efficient and Privacy-Preserving Collaborative Training for Mixture-of-Experts LLMs DiLoCo: Distributed Low-Communication Training of Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:17:08.727540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:17:08.727540Z digest=sha256:dc50926459bae548507fa93e6a0b2813adb35c9d2f9989d0a040c23874dc64d0

Observation 416f7470-76d8-4077-8710-3ab49683b68f · inbound

HALoS: Hierarchical Asynchronous Local SGD over Slow Networks for Geo-Distributed Large Language Model Training cites this paper.

HALoS: Hierarchical Asynchronous Local SGD over Slow Networks for Geo-Distributed Large Language Model Training DiLoCo: Distributed Low-Communication Training of Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:46:06.468296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:46:06.468296Z digest=sha256:4cca47be1f0fc387737d0a684de6d4c6db5841c5cf68251cd37789d56f8171bb

Observation 407c224f-5d82-43e0-9a22-a0ba775b36cf · inbound

NoLoCo: No-all-reduce Low Communication Training Method for Large Models cites this paper.

NoLoCo: No-all-reduce Low Communication Training Method for Large Models DiLoCo: Distributed Low-Communication Training of Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:01.338312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:01.338312Z digest=sha256:0e1e843727321443681d1c2485bfc12954afe229dca1260e9db16789beaa73fa

Observation 8549cf5b-dd06-4a11-98b6-9f82744c914d · inbound

DiLoCoX: A Low-Communication Large-Scale Training Framework for Decentralized Cluster cites this paper.

DiLoCoX: A Low-Communication Large-Scale Training Framework for Decentralized Cluster DiLoCo: Distributed Low-Communication Training of Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:40:13.436255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:40:13.436255Z digest=sha256:2781e2971b9ca124f99bd344004a41b18f3a637c9b696d94f5778f7b74427890

Observation 547be796-941e-40ce-aa18-496d59b245c6 · inbound

On the Surprising Effectiveness of a Single Global Merging in Decentralized Learning cites this paper.

On the Surprising Effectiveness of a Single Global Merging in Decentralized Learning DiLoCo: Distributed Low-Communication Training of Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:42:06.087985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-19T05:39:53.088948Z digest=sha256:1cf21927db8ffb4d114be8db8e2a83d8d1b6051d13b8ca36128b7ceb3b972991

Observation 4581f219-a55e-46eb-8f6f-dac9a32201d5 · inbound

DICE: Data Influence Cascade in Decentralized Learning cites this paper.

DICE: Data Influence Cascade in Decentralized Learning DiLoCo: Distributed Low-Communication Training of Language Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T19:00:00.918889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:00:00.918889Z digest=sha256:91e4da07184cf41bf6dd2346c9efaecb410a2bf115a0f3bc7bd6ba6101523df5

Observation ef29e758-7d1f-40ed-898a-be261a4fc9a9 · inbound

Distributed and Decentralised Training: Technical Governance Challenges in a Shifting AI Landscape cites this paper.

Distributed and Decentralised Training: Technical Governance Challenges in a Shifting AI Landscape DiLoCo: Distributed Low-Communication Training of Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:45.531277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:45.531277Z digest=sha256:23899d41eaa7e25a87eb5b2b19c00c0a884f8db77966a925bcefc85558c1df10

Observation ee585c18-0e62-4758-9dbe-f2062d9917bb · inbound

Model Parallelism With Subnetwork Data Parallelism cites this paper.

Model Parallelism With Subnetwork Data Parallelism DiLoCo: Distributed Low-Communication Training of Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:59.333518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:59.333518Z digest=sha256:cd9feb9dc459e3becd3e4244397a94f9edb5f74d0f94d662a271c51c598c4524

Observation 719527bf-0822-43b7-ac90-8e208eac7646 · inbound

Compute Requirements for Algorithmic Innovation in Frontier AI Models cites this paper.

Compute Requirements for Algorithmic Innovation in Frontier AI Models DiLoCo: Distributed Low-Communication Training of Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:50.265538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:50.265538Z digest=sha256:da75c912ef047419370c6e05f0d7ff9385b29a13893c7485bc9bbf37b7bc415e

Observation 56cd2afe-9f0a-41fd-a231-e5696fb4f400 · inbound

Overcoming the Communication-Performance Tradeoff in LLM Pretraining cites this paper.

Overcoming the Communication-Performance Tradeoff in LLM Pretraining DiLoCo: Distributed Low-Communication Training of Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T17:50:42.327928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:50:42.327928Z digest=sha256:4e74e1553278326d8bab1c87e89e07c3a1dd60b277f844dfa6593d1247b95988

Observation d20fa830-0265-4a0b-8cef-3700bff4f7c1 · inbound

AdLoCo: adaptive batching significantly improves communications efficiency and convergence for Large Language Models cites this paper.

AdLoCo: adaptive batching significantly improves communications efficiency and convergence for Large Language Models DiLoCo: Distributed Low-Communication Training of Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T16:36:50.502069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:36:50.502069Z digest=sha256:36a8e88f498eb5005fdc2897cb3b028da69c7bdbadc8d58de1d9bb1ea9c29b20

Observation 735fd172-a4bc-4a6b-a7b2-3b5eb52e84cb · inbound

Paris: A Decentralized Trained Open-Weight Diffusion Model cites this paper.

Paris: A Decentralized Trained Open-Weight Diffusion Model DiLoCo: Distributed Low-Communication Training of Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T12:37:33.139636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:37:33.139636Z digest=sha256:00a4f1552fc099db48799326693037791ac21a6249a0a1b3f43a9d2dd818a828

Observation ee686517-8e08-47fb-bea2-93f0aa6fde3e · inbound

Towards a future space-based, highly scalable AI infrastructure system design cites this paper.

Towards a future space-based, highly scalable AI infrastructure system design DiLoCo: Distributed Low-Communication Training of Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T21:00:25.350044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:00:25.350044Z digest=sha256:c3191befcc954e24269999d76ef59331879af498455f739572698a1efd7a308d

Observation 8affa5a0-6c2d-45ed-b5c8-488d36d803c2 · inbound

Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic cites this paper.

Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic DiLoCo: Distributed Low-Communication Training of Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:50:49.408158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T09:47:47.051969Z digest=sha256:f8bbfe26d7aa7dcccd79b8ad011542b9a69b82b0c16ea95ca337ce3ae828e247

Observation 14ec8947-395e-4b2e-8ba7-f4de7bc932a9 · inbound

Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic cites this paper.

Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic DiLoCo: Distributed Low-Communication Training of Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:38.886381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:38.886381Z digest=sha256:0b8f71f7dbbb7ff3d7c839aced65ee5d8ee0c612c9ca99aff93d17037faee065

Observation 60c3ecbb-c895-41d5-9e0d-221fa60f8ccb · inbound

LoRDO: Distributed Low-Rank Optimization with Infrequent Communication cites this paper.

LoRDO: Distributed Low-Rank Optimization with Infrequent Communication DiLoCo: Distributed Low-Communication Training of Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T04:44:55.319460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:44:55.319460Z digest=sha256:151d56e43284e8a91937c7a32a9e6b5b12755b200238bbc0a37fc059e87afb34

Observation d1c9f7f1-8bbe-4e3c-baf0-9adfa6831473 · inbound

Distributed Generative Inference of LLM at Internet Scales with Multi-Dimensional Communication Optimization cites this paper.

Distributed Generative Inference of LLM at Internet Scales with Multi-Dimensional Communication Optimization DiLoCo: Distributed Low-Communication Training of Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-09T23:04:17.846136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-09T23:01:56.712670Z digest=sha256:4bf1316a2f96c3e17edb2f90782e854786899c244d9882ae7760e9fdb555275c

Observation 669f8d3c-b7c6-48da-a9fa-389612f0a24d · inbound

Scalable Hyperparameter-Divergent Ensemble Training with Automatic Learning Rate Exploration for Large Models cites this paper.

Scalable Hyperparameter-Divergent Ensemble Training with Automatic Learning Rate Exploration for Large Models DiLoCo: Distributed Low-Communication Training of Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:51:20.305569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T04:08:25.530778Z digest=sha256:2a7f8c041e58a73a1b149e3a7a8ecc00aee622d584c99ebc7dd51c149854ec36

Observation 7d01d556-d04a-45ed-b874-a4ebec0d0f21 · inbound

Rennala MVR: Improved Time Complexity for Parallel Stochastic Optimization via Momentum-Based Variance Reduction cites this paper.

Rennala MVR: Improved Time Complexity for Parallel Stochastic Optimization via Momentum-Based Variance Reduction DiLoCo: Distributed Low-Communication Training of Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:51:31.493255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-12T01:51:20.003552Z digest=sha256:15cb5b4915603c333a2ff4869752be258be9d885453f5e8ce908789696547ce4

Observation da34be27-988f-409b-83a0-7ec12094c728 · inbound

Cosine-Gated Adam-Decay: Drop-In Staleness-Aware Outer Optimization for Decoupled DiLoCo cites this paper.

Cosine-Gated Adam-Decay: Drop-In Staleness-Aware Outer Optimization for Decoupled DiLoCo DiLoCo: Distributed Low-Communication Training of Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:31:16.903062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T02:30:11.470788Z digest=sha256:ff56536bdabc99525ce781db416690bd56488fd494a155e744c4c02be39e5350

Observation 4e08d48d-d067-4fd7-8043-523dfc129b2f · inbound

Optimistic Dual Averaging Unifies Modern Optimizers cites this paper.

Optimistic Dual Averaging Unifies Modern Optimizers DiLoCo: Distributed Low-Communication Training of Language Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:52:21.837939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T05:52:16.805180Z digest=sha256:d3074e75731691e796ab0e5ed8c0866dd29bc1dda8335d263b1f3838326688fc

Observation 0dbb234a-1f2d-48ae-8b63-af31aca8ebeb · inbound

ScheduleFree+: Scaling Learning-Rate-Free & Schedule-Free Learning to Large Language Models cites this paper.

ScheduleFree+: Scaling Learning-Rate-Free & Schedule-Free Learning to Large Language Models DiLoCo: Distributed Low-Communication Training of Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:23:16.790219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-20T12:22:30.263086Z digest=sha256:dd74ff41fc3a88e1126747f2e3d6293b8413c86d6b3e2aebebfd4e8627acfb14

Observation f2347315-cd60-43e8-ad4c-7105119d1cd7 · inbound

ScaleAcross Explorer: Exploring Communication Optimization for Scale-Across AI Model Training cites this paper.

ScaleAcross Explorer: Exploring Communication Optimization for Scale-Across AI Model Training DiLoCo: Distributed Low-Communication Training of Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:04:40.496903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T12:57:37.344369Z digest=sha256:6acc61a4e7a440ba94678289afaeb578b24997e7fd1c20a5da8d7fae7ad8e5c1

Observation a9082207-48d1-4bc4-b2f5-e4ba2ba2779c · inbound

Outer-Momentum Restarting in High-Dimensional Two-Phase Optimization cites this paper.

Outer-Momentum Restarting in High-Dimensional Two-Phase Optimization DiLoCo: Distributed Low-Communication Training of Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:33:27.499368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T13:31:23.664332Z digest=sha256:2d5f2658b2bb263836769ce24561f27c594203d53409c1771805f9d031d9660f

Observation 89e02c70-d3cc-40a7-a260-fd4ab45ed7d0 · inbound

Local MixVR: Breaking the Communication-Sample Dependence in Distributed Learning cites this paper.

Local MixVR: Breaking the Communication-Sample Dependence in Distributed Learning DiLoCo: Distributed Low-Communication Training of Language Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T21:16:14.072138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T17:22:36.503465Z digest=sha256:7a84f4d3f59618b47e64553457e506dcf6b4f212af03f00f1f1329db8ee232b6

Observation e18b6555-34f8-4f6c-96e7-27c2a768b362 · inbound

Echelon: Auditable Aggregate-Only Language-Model Adaptation Across Privacy Boundaries cites this paper.

Echelon: Auditable Aggregate-Only Language-Model Adaptation Across Privacy Boundaries DiLoCo: Distributed Low-Communication Training of Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:06:23.646181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T13:38:48.333498Z digest=sha256:709295c9d98f923c8cd85ddf4cac3ca5461603b7cf48e68e1144f03b5c0ddcd4

Observation cbdf6fdf-bd18-4ff0-8067-26d1d8d21a0d · inbound

Learned Subspace Compression for Communication-Efficient Pipeline Parallelism cites this paper.

Learned Subspace Compression for Communication-Efficient Pipeline Parallelism DiLoCo: Distributed Low-Communication Training of Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:46:46.170228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T06:41:58.301482Z digest=sha256:c589979f6aa563a50c8a64b3135fe38b038a919f3857fe2647e8a303bc18bb30

Observation 6c2cffcd-d7b8-49f4-a770-06a1de08f54f · inbound

Unifying Local Communications and Local Updates for LLM Pretraining cites this paper.

Unifying Local Communications and Local Updates for LLM Pretraining DiLoCo: Distributed Low-Communication Training of Language Models

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:07:37.145086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T14:07:42.062805Z digest=sha256:3c7077a8cc3bc86b5deabb91f86dce67e51fef8f36c6238c2df46a155788ec99

Observation a21a84fc-055c-42ac-af50-6b1ca4ac5690 · inbound

FoMoE: Breaking the Full-Replica Barrier with a Federation of MoEs cites this paper.

FoMoE: Breaking the Full-Replica Barrier with a Federation of MoEs DiLoCo: Distributed Low-Communication Training of Language Models

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-06-26T21:30:02.824990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-26T21:25:15.709652Z digest=sha256:33e75e984a1cc751a3cfb13bca33e270c8994e0c903c05b2cc7ee67dfcb9090a

Observation da0efcae-8cec-4d96-94fd-204f83bcd174 · inbound

Quantum-Resilient Decentralized AI Economies: Proof-of-Useful-Work and Post-Quantum Security cites this paper.

Quantum-Resilient Decentralized AI Economies: Proof-of-Useful-Work and Post-Quantum Security DiLoCo: Distributed Low-Communication Training of Language Models

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T11:49:50.890394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T07:34:47.006198Z digest=sha256:a09b308b8213bdc4587674b598c5631ee524286b3df16c721e5ad407de621f83

Observation 49c6774a-1edd-4d97-bfe8-5f87a536ef4a · inbound

Not Every Sync Is Safe: Calibrated DiLoCo Scheduling for Shared AI Infrastructure cites this paper.

Not Every Sync Is Safe: Calibrated DiLoCo Scheduling for Shared AI Infrastructure DiLoCo: Distributed Low-Communication Training of Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T12:16:10.904456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T12:16:10.904456Z digest=sha256:7bf269444487b94569613060931d8a66f7d73f445f79b9e2dd3921835b83c3cc

Observation f05a57ea-6304-4234-a165-48f9392853ae · inbound

Byzantine Accountability Without Consensus: Strong Eventual Consistency for Non-Associative, Stochastic, Robust Aggregation cites this paper.

Byzantine Accountability Without Consensus: Strong Eventual Consistency for Non-Associative, Stochastic, Robust Aggregation DiLoCo: Distributed Low-Communication Training of Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T12:46:51.615514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:46:51.615514Z digest=sha256:6572909d1e20c35270192aa887c6746e370dc0537137d18a7ee52389d0f818e5

Observation 2148f948-01ab-44ea-bc42-d3748978e21f · inbound

What's in a Smoothness Constant? Tighter Rates for Local SGD with Bounded Second-order Heterogeneity cites this paper.

What's in a Smoothness Constant? Tighter Rates for Local SGD with Bounded Second-order Heterogeneity DiLoCo: Distributed Low-Communication Training of Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T01:23:30.172502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:23:30.172502Z digest=sha256:84e79ba4f4b48181dec8bbfc792b55004464e9e6f80bd5eb41e0ed8b4f7ecec9

Observation 19bc8e28-174c-4e81-9325-4f63e772b91f · inbound

Federated Lightweight Fine-Tuning cites this paper.

Federated Lightweight Fine-Tuning DiLoCo: Distributed Low-Communication Training of Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T17:36:03.863586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:36:03.863586Z digest=sha256:701f73bad8706708ab639bd088083c695cd1c83c0d09f7677ef7dc26d630b386

Observation 6bfc9d83-eeff-4b6c-96eb-2155984c6f95 · inbound

Controlled Periodic Synchronization for Efficient Data-Parallel Training cites this paper.

Controlled Periodic Synchronization for Efficient Data-Parallel Training DiLoCo: Distributed Low-Communication Training of Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T08:10:51.408706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:10:51.408706Z digest=sha256:231573389dc7d57f7a0c95023f8f28f4e3fb58f8f1c6d55cf78228e49369dffb

Observation 4244ccb7-abc3-4f0a-aa46-e3d037045424 · inbound

DistMoE: Private-data Rehearsal-free Routing in Mixture-of-Experts for Distributed Instruction Tuning cites this paper.

DistMoE: Private-data Rehearsal-free Routing in Mixture-of-Experts for Distributed Instruction Tuning DiLoCo: Distributed Low-Communication Training of Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T14:28:03.357493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:28:03.357493Z digest=sha256:af9624ac2f6d688cdebe8183d0b8e1c94b33530ea20ef4e272091bbab19fb7e3