Pith. sign in

REVIEW 5 minor 39 references

Isotropic merging of worker updates turns DiLoCo into a low-communication trainer that degrades far less as worker count grows.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 05:27 UTC pith:FOCFNAR2

load-bearing objection Clean, useful bridge from model merging to DiLoCo outer steps; IsoLoCo is a practical win that scales better with workers, and the evidence holds inside the stated regime.

arxiv 2607.03011 v1 pith:FOCFNAR2 submitted 2026-07-03 cs.LG cs.AIcs.CL

Can Model Merging Improve Aggregation in DiLoCo?

classification cs.LG cs.AIcs.CL
keywords model mergingDiLoCoIsoLoColow-communication trainingtask arithmeticisotropic aggregationlanguage model pre-trainingdistributed optimization
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Training large language models across many weakly connected machines is attractive because it cuts communication, but standard DiLoCo-style aggregation of local updates loses quality as the number of workers or local steps grows. This paper treats each worker's parameter change as a task vector from the model-merging literature and shows that the same aggregation step can be replaced by a merging rule designed to reduce interference. After testing several modern merging methods, the authors find that isotropic correction of the averaged update (Iso-C) already beats plain DiLoCo, even without momentum. They then add Nesterov momentum to produce IsoLoCo. On language-model pre-training from 178 M to 1 B parameters, IsoLoCo consistently reaches lower validation loss than DiLoCo; the gap widens with more workers and remains when local steps are lengthened. A fast Newton-Schulz version keeps nearly the same accuracy while cutting orthogonalization cost dramatically.

Core claim

Replacing DiLoCo's pseudo-gradient average with Iso-C-style isotropic spectral correction, then equipping the corrected update with Nesterov momentum, yields IsoLoCo: a low-communication outer step that substantially narrows the performance gap to data-parallel training as worker count and inner-step count increase.

What carries the argument

IsoLoCo: after averaging worker pseudo-gradients, replace every matrix's singular values by their mean (or RMS proxy) so the update becomes isotropic, then feed the corrected gradient into a Nesterov momentum outer optimizer.

Load-bearing premise

The ranking of methods rests on single-run validation loss under a FLOP-matched protocol that shrinks each worker's batch as workers increase; if multi-seed variance, different data, or larger models reverse that ranking, the claim collapses.

What would settle it

Train matched 178 M or 512 M models with DiLoCo versus IsoLoCo at R=64 or R=128 under identical FLOP and global-batch budgets; if IsoLoCo no longer shows a clear validation-loss advantage that grows with R, the central claim fails.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 5 minor

Summary. The paper draws a structural analogy between DiLoCo’s outer pseudo-gradient aggregation and task-arithmetic model merging (Eqs. 1–2, §2.3), then systematically substitutes several merging methods for DiLoCo’s outer step. Iso-C is identified as the strongest drop-in aggregator (Table 1). The authors equip it with Nesterov momentum to obtain IsoLoCo (Alg. 2) and evaluate it on Llama-style pre-training (178 M–1 B, DCLM, Chinchilla 20× tokens) under a FLOP-matched, fixed-global-batch protocol. IsoLoCo consistently lowers validation loss relative to DiLoCo, with the gap widening as the number of workers grows to R=128 (Tables 2–3, Fig. 1); the advantage also holds across inner-step counts H and composes with a Muon inner optimizer. A Newton–Schulz + RMS approximation yields a practical speed-up with negligible loss degradation.

Significance. If the empirical ranking holds, the work supplies a concrete, low-overhead outer-step improvement for the increasingly important DiLoCo family of low-communication optimizers. The explicit merging–distributed-training bridge is novel and immediately actionable; the public code, extensive hyper-parameter tables (Appendix A), and ablations (momentum type, orthogonalization target, CLIP_HIGH/LOW, Muon composition) raise the bar for reproducibility in this area. The result is of practical interest to groups training across heterogeneous or geo-distributed clusters and of conceptual interest to both the model-merging and federated-optimization communities.

minor comments (5)
  1. All main tables report single-run validation losses. Even a brief multi-seed check at one (R,H) setting would strengthen confidence that the ranking is not seed-dependent.
  2. Figure 1 and Tables 2–4 would benefit from a short note clarifying that the AdamW DP baseline uses the same global batch and token budget; the FLOP-matched claim is stated in §3.1 but is easy to miss when reading the figures alone.
  3. In §4.3 the RMS proxy is introduced without an explicit inequality relating it to the arithmetic mean; a one-line remark that RMS ≥ mean (with equality only for flat spectra) would make the approximation transparent.
  4. Appendix B’s comparison of IsoLoCo versus Muon-as-outer-optimizer is useful; a pointer to it from the main-text Muon discussion (§4.4) would help readers who skip the appendix.
  5. A few typographical inconsistencies remain (e.g., “dimen-sions”, “matriecs” in the appendix). A light copy-edit pass would polish the camera-ready version.

Circularity Check

0 steps flagged

No significant circularity: purely empirical ranking of aggregation rules under FLOP-matched training; no prediction is forced by construction.

full rationale

The paper's central claim is that IsoLoCo (Iso-C isotropic correction of averaged pseudo-gradients plus Nesterov outer momentum) yields lower held-out validation loss than DiLoCo across worker counts, model sizes, and inner-step counts. That claim is established by training runs whose final losses are measured on a held-out DCLM subset (Tables 1–4, Figure 1). The only formal step is the structural analogy in Section 2.3 (Eqs. 1–2) showing that DiLoCo's SGD outer step is task-arithmetic merging; this is an identity of update forms, not a derivation that forces IsoLoCo's superiority. Iso-C and DiLoCo are taken from independent prior literature and re-implemented; hyper-parameters are swept (Appendix A); ablations (Section 4.5) and the Newton–Schulz variant are also measured, not fitted-then-predicted. No equation, uniqueness theorem, or self-citation chain makes the reported ranking true by construction. Score 0 is therefore the correct outcome.

Axiom & Free-Parameter Ledger

3 free parameters · 3 axioms · 1 invented entities

The central claim is empirical and rests on standard optimizer and data-parallel assumptions plus the experimental protocol choices. No new physical entities are postulated; free parameters are the usual learning-rate and momentum knobs that are swept rather than fitted to force the result.

free parameters (3)
  • outer learning rate η_out = setting-dependent (e.g. 1.8 for IsoLoCo R=8)
    Swept extensively (0.5–25+); optimal values differ markedly between DiLoCo and IsoLoCo and are reported per setting. The ranking of methods is sensitive to this choice.
  • outer momentum μ = 0.7–0.9
    Swept in {0.6,0.7,0.8,0.9}; Nesterov form preferred. Affects absolute numbers but not the qualitative ranking.
  • inner learning rate η_in = setting-dependent
    Swept on a √2 grid; transferred across model scales after 178 M tuning. Standard practice but still a free choice that influences final loss.
axioms (3)
  • domain assumption FLOP-matched comparison with fixed global batch size (inner batch = global / R) is the correct fairness criterion for low-communication methods.
    Stated in Section 3.1 and used for all tables; if wall-clock or communication-volume matching were used instead, rankings could change.
  • domain assumption Pseudo-gradients after H local AdamW steps behave sufficiently like task vectors for merging methods designed for multi-task fine-tuning to transfer.
    Core analogy of Section 2.3; empirically supported for Iso-C but not proven for arbitrary merging methods.
  • domain assumption Single-run validation loss on a 100 M-token DCLM subset is a reliable ranking signal.
    All reported numbers are single-run; no error bars or multi-seed statistics are given.
invented entities (1)
  • IsoLoCo outer optimizer no independent evidence
    purpose: Combines Iso-C isotropic spectral correction of the averaged pseudo-gradient with Nesterov momentum for DiLoCo-style training.
    Defined in Algorithm 2; the paper's main algorithmic contribution. Independent evidence is the empirical tables themselves; no external prediction is made.

pith-pipeline@v1.1.0-grok45 · 22582 in / 2713 out tokens · 27089 ms · 2026-07-12T05:27:45.086662+00:00 · methodology

0 comments
read the original abstract

Model merging techniques, which aggregate independently finetuned models into one to combine their capabilities, have become a topic of significant interest in recent years, with a broad array of methods having been proposed to tackle this problem. Simultaneously, an emerging trend in distributed learning has been the use of methods such as local SGD and DiLoCo, which greatly reduce communication costs by periodically aggregating the independently trained local models. However, these communication-efficient methods have been shown to degrade in performance relative to the FLOP-matched data-parallel gold standard as the number of independent local models grows and as the number of local training steps before global communication is increased. In this work, we draw an explicit analogy between the pseudo-gradient aggregation step in local SGD/DiLoCo and task arithmetic-based model merging, establishing a straightforward way to utilize merging methods in the context of distributed optimization. We then evaluate multiple state-of-the-art model merging methods in this setting and identify one method in particular, Iso-C, as a promising approach for improving DiLoCo. We find that DiLoCo SGD with Iso-C aggregation outperforms not only simple pseudo-gradient averaging but even the momentum-based DiLoCo, despite lacking a momentum mechanism itself. Building on this finding, we propose IsoLoCo, which adapts Iso-C for distributed training by equipping it with Nesterov momentum. Our empirical evaluations on language model pre-training across varying numbers of local workers show that IsoLoCo significantly outperforms DiLoCo, with the gap between them widening as the number of workers increases. This advantage remains present across model sizes and inner step counts, confirming that merging-inspired aggregation is an effective strategy for low-communication distributed training.

Figures

Figures reproduced from arXiv: 2607.03011 by Benjamin Th\'erien, Eugene Belilovsky, Guy Wolf, Stefan Horoi.

Figure 1
Figure 1. Figure 1: IsoLoCo scales better with work￾ers than DiLoCo. Final validation loss is shown versus worker count for a 178M LLaMa-2 model in a FLOP-matched setting. Each cluster, or worker, trains its own model replica through multiple local optimization steps before shar￾ing the accumulated parameter changes, i.e. pseudo￾gradients, with the remaining clusters for a global optimization step. Performing many inner optim… view at source ↗
Figure 2
Figure 2. Figure 2: DiLoCo as iterative model merging. Dots represent model checkpoints, with blue dots denoting locally trained models: fine-tuned experts in model merging and worker replicas in DiLoCo respectively. Blue arrows are local updates—task vectors in merging and pseudo-gradients in DiLoCo—measuring displacement from the shared initialization. Red arrows denote aggregation: a one-shot merge on the left, and an oute… view at source ↗
Figure 3
Figure 3. Figure 3: IsoLoCo further improves MuLoCo’s worker-count scaling. Final validation loss is shown versus worker count for a 178M LLaMa-2 model in a FLOP-matched setting [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

39 extracted references · 3 linked inside Pith

  1. [1]

    K. Ahn, B. Xu, N. Abreu, Y . Fan, G. Magakyan, P. Sharma, Z. Zhan, and J. Langford. Dion: Distributed orthonormalized updates, 2025

  2. [2]

    AI, :, I

    E. AI, :, I. Shah, A. M. Polloreno, K. Stratos, P. Monk, A. Chaluvaraju, A. Hojel, A. Ma, A. Thomas, A. Tanwer, D. J. Shah, K. Nguyen, K. Smith, M. Callahan, M. Pust, M. Parmar, P. Rushton, P. Mazarakis, R. Kapila, S. Srivastava, S. Singla, T. Romanski, Y . Vanjani, and A. Vaswani. Practical efficiency of muon for pretraining, 2025

  3. [3]

    Björck and C

    Å. Björck and C. Bowie. An iterative algorithm for computing the best estimate of an orthogonal matrix.SIAM Journal on Numerical Analysis, 8(2):358–364, 1971

  4. [4]

    Charles, G

    Z. Charles, G. Teston, L. M. Dery, J. K. Rush, N. Fallen, Z. Garrett, A. Szlam, and A. Douillard. Communication-efficient language model training scales reliably and robustly: Scaling laws for diloco. InThe Thirty-ninth Annual Conference on Neural Information Processing Systems, 2026

  5. [5]

    Davari and E

    M. Davari and E. Belilovsky. Model breadcrumbs: Scaling multi-task model merging with sparse masks. InComputer Vision – ECCV 2024: 18th European Conference, Milan, Italy, September 29–October 4, 2024, Proceedings, Part LXXV, page 270–287, Berlin, Heidelberg,

  6. [6]

    Douillard, Q

    A. Douillard, Q. Feng, A. A. Rusu, R. Chhaparia, Y . Donchev, A. Kuncoro, M. Ranzato, A. Szlam, and J. Shen. Diloco: Distributed low-communication training of language models, 2024

  7. [7]

    Draxler, K

    F. Draxler, K. Veschgini, M. Salmhofer, and F. Hamprecht. Essentially no barriers in neural network energy landscape. In J. Dy and A. Krause, editors,International Conference on Machine Learning (ICML), volume 80 ofProceedings of Machine Learning Research, pages 1309–1318. PMLR, 10–15 Jul 2018

  8. [8]

    Frankle, G

    J. Frankle, G. K. Dziugaite, D. Roy, and M. Carbin. Linear mode connectivity and the lottery ticket hypothesis. In H. D. III and A. Singh, editors,International Conference on Machine Learning (ICML), volume 119 ofProceedings of Machine Learning Research, pages 3259–3269. PMLR, 13–18 Jul 2020

  9. [9]

    A. A. Gargiulo, D. Crisostomi, M. S. Bucarelli, S. Scardapane, F. Silvestri, and E. Rodolà. Task singular vectors: Reducing task interference in model merging. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 18695–18705, June 2025

  10. [10]

    Garipov, P

    T. Garipov, P. Izmailov, D. Podoprikhin, D. P. Vetrov, and A. G. Wilson. Loss surfaces, mode connectivity, and fast ensembling of DNNs. In S. Bengio, H. Wallach, H. Larochelle, K. Grau- man, N. Cesa-Bianchi, and R. Garnett, editors,Advances in Neural Information Processing Systems (NeurIPS), volume 31. Curran Associates, Inc., 2018

  11. [11]

    Hoffmann, S

    J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, J. W. Rae, O. Vinyals, and L. Sifre. Training compute-optimal large language models, 2022

  12. [12]

    Horoi, G

    S. Horoi, G. Wolf, E. Belilovsky, and G. K. Dziugaite. Less is more: Undertraining experts improves model upcycling.arXiv preprint arXiv:2506.14126, 2025

  13. [13]

    Ilharco, M

    G. Ilharco, M. T. Ribeiro, M. Wortsman, L. Schmidt, H. Hajishirzi, and A. Farhadi. Editing models with task arithmetic. InInternational Conference on Learning Representations (ICLR), 2023

  14. [14]

    Jaghouar, J

    S. Jaghouar, J. M. Ong, M. Basra, F. Obeid, J. Straube, M. Keiblinger, E. Bakouch, L. Atkins, M. Panahi, C. Goddard, M. Ryabinin, and J. Hagemann. INTELLECT-1 technical report.CoRR, abs/2412.01152, 2024. 12

  15. [15]

    Jordan, Y

    K. Jordan, Y . Jin, V . Boza, J. You, F. Cesista, L. Newhouse, and J. Bernstein. Muon: An optimizer for hidden layers in neural networks. https://kellerjordan.github.io/posts/muon/,

  16. [16]

    Accessed: 2026-04-27

  17. [17]

    D. P. Kingma and J. Ba. Adam: A method for stochastic optimization. InInternational Conference on Learning Representations (ICLR), 2015

  18. [18]

    Z. Kovarik. Some iterative methods for improving orthonormality.SIAM Journal on Numerical Analysis, 7(3):386–389, 1970

  19. [19]

    J. Li, A. Fang, G. Smyrnis, M. Ivgi, M. Jordan, S. Gadre, H. Bansal, E. Guha, S. Keh, K. Arora, S. Garg, R. Xin, N. Muennighoff, R. Heckel, J. Mercat, M. Chen, S. Gururangan, M. Wortsman, A. Albalak, Y . Bitton, M. Nezhurina, A. Abbas, C.-Y . Hsieh, D. Ghosh, J. Gardner, M. Kilian, H. Zhang, R. Shao, S. Pratt, S. Sanyal, G. Ilharco, G. Daras, K. Marathe, ...

  20. [20]

    Lidin, A

    J. Lidin, A. Sarfi, E. Miahi, Q. Anthony, S. Chauhan, E. Pappas, B. Thérien, E. Belilovsky, and S. Dare. Covenant-72b: Pre-training a 72b llm with trustless peers over-the-internet, 2026

  21. [21]

    J. Liu, J. Su, X. Yao, Z. Jiang, G. Lai, Y . Du, Y . Qin, W. Xu, E. Lu, J. Yan, Y . Chen, H. Zheng, Y . Liu, S. Liu, B. Yin, W. He, H. Zhu, Y . Wang, J. Wang, M. Dong, Z. Zhang, Y . Kang, H. Zhang, X. Xu, Y . Zhang, Y . Wu, X. Zhou, and Z. Yang. Muon is scalable for llm training, 2025

  22. [22]

    Loshchilov and F

    I. Loshchilov and F. Hutter. Decoupled weight decay regularization. InInternational Conference on Learning Representations (ICLR), 2019

  23. [23]

    Marczak, S

    D. Marczak, S. Magistri, S. Cygert, B. Twardowski, A. D. Bagdanov, and J. van de Weijer. No task left behind: Isotropic model merging with common and task-specific subspaces. In Forty-second International Conference on Machine Learning, 2025

  24. [24]

    McMahan, E

    B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. y. Arcas. Communication- Efficient Learning of Deep Networks from Decentralized Data. In A. Singh and J. Zhu, editors, Proceedings of the 20th International Conference on Artificial Intelligence and Statistics, volume 54 ofProceedings of Machine Learning Research, pages 1273–1282. PMLR, 20–22 Apr 2017

  25. [25]

    Nesterov

    Y . Nesterov. A method for solving the convex programming problem with convergence rate O(1/k2).Proceedings of the USSR Academy of Sciences, 269:543–547, 1983

  26. [26]

    Neyshabur, H

    B. Neyshabur, H. Sedghi, and C. Zhang. What is being transferred in transfer learning? In H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, editors,Advances in Neural Information Processing Systems (NeurIPS), volume 33, pages 512–523. Curran Associates, Inc., 2020

  27. [27]

    Paszke, S

    A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Te- jani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala. PyTorch: An imperative style, high-performance deep learning library. In H. Wallach, H. Larochelle, A. Beygelzimer, ...

  28. [28]

    Pfeiffer, A

    J. Pfeiffer, A. Rücklé, C. Poth, A. Kamath, I. Vuli ´c, S. Ruder, K. Cho, and I. Gurevych. AdapterHub: A framework for adapting transformers. In Q. Liu and D. Schlangen, editors, Proceedings of the 2020 Conference on Empirical Methods in Natural Language Process- ing: System Demonstrations, pages 46–54, Online, Oct. 2020. Association for Computational Lin...

  29. [29]

    S. J. Reddi, Z. Charles, M. Zaheer, Z. Garrett, K. Rush, J. Koneˇcný, S. Kumar, and H. B. McMa- han. Adaptive federated optimization. InInternational Conference on Learning Representations, 2021

  30. [30]

    Sarfi, B

    A. Sarfi, B. Thérien, J. Lidin, and E. Belilovsky. Communication efficient llm pre-training with sparseloco.https://arxiv.org/pdf/2508.15706, 2025

  31. [31]

    Sutskever, J

    I. Sutskever, J. Martens, G. Dahl, and G. Hinton. On the importance of initialization and mo- mentum in deep learning. In S. Dasgupta and D. McAllester, editors,International Conference on Machine Learning (ICML), volume 28 ofProceedings of Machine Learning Research, pages 1139–1147, Atlanta, Georgia, USA, 17–19 Jun 2013. PMLR

  32. [32]

    Thérien, X

    B. Thérien, X. Huang, A. Defazio, I. Rish, and E. Belilovsky. Muloco: Muon is a practical inner optimizer for diloco, 2026

  33. [33]

    Touvron, L

    H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y . Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V . Goswami, N. Goyal, A. Hartshorn, S. Hosseini, R. Hou, H. Inan, M. Kardas, V . Kerkez, M. Khabsa, I. Kloumann, A. Koren...

  34. [34]

    T. Wolf, L. Debut, V . Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, and J. Brew. Huggingface’s transformers: State-of-the-art natural language processing.CoRR, abs/1910.03771, 2019

  35. [35]

    Yadav, D

    P. Yadav, D. Tam, L. Choshen, C. A. Raffel, and M. Bansal. Ties-merging: Resolving interfer- ence when merging models. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors,Advances in Neural Information Processing Systems (NeurIPS), volume 36, pages 7093–7115. Curran Associates, Inc., 2023

  36. [36]

    E. Yang, L. Shen, G. Guo, X. Wang, X. Cao, J. Zhang, and D. Tao. Model merging in LLMs, MLLMs, and beyond: Methods, theories, applications, and opportunities.ACM Comput. Surv., 58(8), Feb. 2026

  37. [37]

    G. Yang, J. B. Simon, and J. Bernstein. A spectral condition for feature learning, 2024

  38. [38]

    L. Yu, B. Yu, H. Yu, F. Huang, and Y . Li. Language models are super mario: Absorbing abilities from homologous models as a free lunch. In R. Salakhutdinov, Z. Kolter, K. Heller, A. Weller, N. Oliver, J. Scarlett, and F. Berkenkamp, editors,International Conference on Machine Learning (ICML), volume 235 ofProceedings of Machine Learning Research, pages 57...

  39. [39]

    spectral

    L. Zhou, D. Solombrino, D. Crisostomi, M. S. Bucarelli, G. A. D’Inverno, F. Silvestri, and E. Rodolà. On task vectors and gradients. InUniReps: 3rd Edition of the Workshop on Unifying Representations in Neural Models, 2025. 14 A Optimal Hyperparameters We report the optimal hyperparameters used for each main-text result table: inner learning rate ηin, out...