REVIEW 5 minor 39 references
Isotropic merging of worker updates turns DiLoCo into a low-communication trainer that degrades far less as worker count grows.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-12 05:27 UTC pith:FOCFNAR2
load-bearing objection Clean, useful bridge from model merging to DiLoCo outer steps; IsoLoCo is a practical win that scales better with workers, and the evidence holds inside the stated regime.
Can Model Merging Improve Aggregation in DiLoCo?
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Replacing DiLoCo's pseudo-gradient average with Iso-C-style isotropic spectral correction, then equipping the corrected update with Nesterov momentum, yields IsoLoCo: a low-communication outer step that substantially narrows the performance gap to data-parallel training as worker count and inner-step count increase.
What carries the argument
IsoLoCo: after averaging worker pseudo-gradients, replace every matrix's singular values by their mean (or RMS proxy) so the update becomes isotropic, then feed the corrected gradient into a Nesterov momentum outer optimizer.
Load-bearing premise
The ranking of methods rests on single-run validation loss under a FLOP-matched protocol that shrinks each worker's batch as workers increase; if multi-seed variance, different data, or larger models reverse that ranking, the claim collapses.
What would settle it
Train matched 178 M or 512 M models with DiLoCo versus IsoLoCo at R=64 or R=128 under identical FLOP and global-batch budgets; if IsoLoCo no longer shows a clear validation-loss advantage that grows with R, the central claim fails.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper draws a structural analogy between DiLoCo’s outer pseudo-gradient aggregation and task-arithmetic model merging (Eqs. 1–2, §2.3), then systematically substitutes several merging methods for DiLoCo’s outer step. Iso-C is identified as the strongest drop-in aggregator (Table 1). The authors equip it with Nesterov momentum to obtain IsoLoCo (Alg. 2) and evaluate it on Llama-style pre-training (178 M–1 B, DCLM, Chinchilla 20× tokens) under a FLOP-matched, fixed-global-batch protocol. IsoLoCo consistently lowers validation loss relative to DiLoCo, with the gap widening as the number of workers grows to R=128 (Tables 2–3, Fig. 1); the advantage also holds across inner-step counts H and composes with a Muon inner optimizer. A Newton–Schulz + RMS approximation yields a practical speed-up with negligible loss degradation.
Significance. If the empirical ranking holds, the work supplies a concrete, low-overhead outer-step improvement for the increasingly important DiLoCo family of low-communication optimizers. The explicit merging–distributed-training bridge is novel and immediately actionable; the public code, extensive hyper-parameter tables (Appendix A), and ablations (momentum type, orthogonalization target, CLIP_HIGH/LOW, Muon composition) raise the bar for reproducibility in this area. The result is of practical interest to groups training across heterogeneous or geo-distributed clusters and of conceptual interest to both the model-merging and federated-optimization communities.
minor comments (5)
- All main tables report single-run validation losses. Even a brief multi-seed check at one (R,H) setting would strengthen confidence that the ranking is not seed-dependent.
- Figure 1 and Tables 2–4 would benefit from a short note clarifying that the AdamW DP baseline uses the same global batch and token budget; the FLOP-matched claim is stated in §3.1 but is easy to miss when reading the figures alone.
- In §4.3 the RMS proxy is introduced without an explicit inequality relating it to the arithmetic mean; a one-line remark that RMS ≥ mean (with equality only for flat spectra) would make the approximation transparent.
- Appendix B’s comparison of IsoLoCo versus Muon-as-outer-optimizer is useful; a pointer to it from the main-text Muon discussion (§4.4) would help readers who skip the appendix.
- A few typographical inconsistencies remain (e.g., “dimen-sions”, “matriecs” in the appendix). A light copy-edit pass would polish the camera-ready version.
Circularity Check
No significant circularity: purely empirical ranking of aggregation rules under FLOP-matched training; no prediction is forced by construction.
full rationale
The paper's central claim is that IsoLoCo (Iso-C isotropic correction of averaged pseudo-gradients plus Nesterov outer momentum) yields lower held-out validation loss than DiLoCo across worker counts, model sizes, and inner-step counts. That claim is established by training runs whose final losses are measured on a held-out DCLM subset (Tables 1–4, Figure 1). The only formal step is the structural analogy in Section 2.3 (Eqs. 1–2) showing that DiLoCo's SGD outer step is task-arithmetic merging; this is an identity of update forms, not a derivation that forces IsoLoCo's superiority. Iso-C and DiLoCo are taken from independent prior literature and re-implemented; hyper-parameters are swept (Appendix A); ablations (Section 4.5) and the Newton–Schulz variant are also measured, not fitted-then-predicted. No equation, uniqueness theorem, or self-citation chain makes the reported ranking true by construction. Score 0 is therefore the correct outcome.
Axiom & Free-Parameter Ledger
free parameters (3)
- outer learning rate η_out =
setting-dependent (e.g. 1.8 for IsoLoCo R=8)
- outer momentum μ =
0.7–0.9
- inner learning rate η_in =
setting-dependent
axioms (3)
- domain assumption FLOP-matched comparison with fixed global batch size (inner batch = global / R) is the correct fairness criterion for low-communication methods.
- domain assumption Pseudo-gradients after H local AdamW steps behave sufficiently like task vectors for merging methods designed for multi-task fine-tuning to transfer.
- domain assumption Single-run validation loss on a 100 M-token DCLM subset is a reliable ranking signal.
invented entities (1)
-
IsoLoCo outer optimizer
no independent evidence
read the original abstract
Model merging techniques, which aggregate independently finetuned models into one to combine their capabilities, have become a topic of significant interest in recent years, with a broad array of methods having been proposed to tackle this problem. Simultaneously, an emerging trend in distributed learning has been the use of methods such as local SGD and DiLoCo, which greatly reduce communication costs by periodically aggregating the independently trained local models. However, these communication-efficient methods have been shown to degrade in performance relative to the FLOP-matched data-parallel gold standard as the number of independent local models grows and as the number of local training steps before global communication is increased. In this work, we draw an explicit analogy between the pseudo-gradient aggregation step in local SGD/DiLoCo and task arithmetic-based model merging, establishing a straightforward way to utilize merging methods in the context of distributed optimization. We then evaluate multiple state-of-the-art model merging methods in this setting and identify one method in particular, Iso-C, as a promising approach for improving DiLoCo. We find that DiLoCo SGD with Iso-C aggregation outperforms not only simple pseudo-gradient averaging but even the momentum-based DiLoCo, despite lacking a momentum mechanism itself. Building on this finding, we propose IsoLoCo, which adapts Iso-C for distributed training by equipping it with Nesterov momentum. Our empirical evaluations on language model pre-training across varying numbers of local workers show that IsoLoCo significantly outperforms DiLoCo, with the gap between them widening as the number of workers increases. This advantage remains present across model sizes and inner step counts, confirming that merging-inspired aggregation is an effective strategy for low-communication distributed training.
Figures
Reference graph
Works this paper leans on
-
[1]
K. Ahn, B. Xu, N. Abreu, Y . Fan, G. Magakyan, P. Sharma, Z. Zhan, and J. Langford. Dion: Distributed orthonormalized updates, 2025
2025
-
[2]
AI, :, I
E. AI, :, I. Shah, A. M. Polloreno, K. Stratos, P. Monk, A. Chaluvaraju, A. Hojel, A. Ma, A. Thomas, A. Tanwer, D. J. Shah, K. Nguyen, K. Smith, M. Callahan, M. Pust, M. Parmar, P. Rushton, P. Mazarakis, R. Kapila, S. Srivastava, S. Singla, T. Romanski, Y . Vanjani, and A. Vaswani. Practical efficiency of muon for pretraining, 2025
2025
-
[3]
Björck and C
Å. Björck and C. Bowie. An iterative algorithm for computing the best estimate of an orthogonal matrix.SIAM Journal on Numerical Analysis, 8(2):358–364, 1971
1971
-
[4]
Charles, G
Z. Charles, G. Teston, L. M. Dery, J. K. Rush, N. Fallen, Z. Garrett, A. Szlam, and A. Douillard. Communication-efficient language model training scales reliably and robustly: Scaling laws for diloco. InThe Thirty-ninth Annual Conference on Neural Information Processing Systems, 2026
2026
-
[5]
Davari and E
M. Davari and E. Belilovsky. Model breadcrumbs: Scaling multi-task model merging with sparse masks. InComputer Vision – ECCV 2024: 18th European Conference, Milan, Italy, September 29–October 4, 2024, Proceedings, Part LXXV, page 270–287, Berlin, Heidelberg,
2024
-
[6]
Douillard, Q
A. Douillard, Q. Feng, A. A. Rusu, R. Chhaparia, Y . Donchev, A. Kuncoro, M. Ranzato, A. Szlam, and J. Shen. Diloco: Distributed low-communication training of language models, 2024
2024
-
[7]
Draxler, K
F. Draxler, K. Veschgini, M. Salmhofer, and F. Hamprecht. Essentially no barriers in neural network energy landscape. In J. Dy and A. Krause, editors,International Conference on Machine Learning (ICML), volume 80 ofProceedings of Machine Learning Research, pages 1309–1318. PMLR, 10–15 Jul 2018
2018
-
[8]
Frankle, G
J. Frankle, G. K. Dziugaite, D. Roy, and M. Carbin. Linear mode connectivity and the lottery ticket hypothesis. In H. D. III and A. Singh, editors,International Conference on Machine Learning (ICML), volume 119 ofProceedings of Machine Learning Research, pages 3259–3269. PMLR, 13–18 Jul 2020
2020
-
[9]
A. A. Gargiulo, D. Crisostomi, M. S. Bucarelli, S. Scardapane, F. Silvestri, and E. Rodolà. Task singular vectors: Reducing task interference in model merging. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 18695–18705, June 2025
2025
-
[10]
Garipov, P
T. Garipov, P. Izmailov, D. Podoprikhin, D. P. Vetrov, and A. G. Wilson. Loss surfaces, mode connectivity, and fast ensembling of DNNs. In S. Bengio, H. Wallach, H. Larochelle, K. Grau- man, N. Cesa-Bianchi, and R. Garnett, editors,Advances in Neural Information Processing Systems (NeurIPS), volume 31. Curran Associates, Inc., 2018
2018
-
[11]
Hoffmann, S
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, J. W. Rae, O. Vinyals, and L. Sifre. Training compute-optimal large language models, 2022
2022
-
[12]
S. Horoi, G. Wolf, E. Belilovsky, and G. K. Dziugaite. Less is more: Undertraining experts improves model upcycling.arXiv preprint arXiv:2506.14126, 2025
Pith/arXiv arXiv 2025
-
[13]
Ilharco, M
G. Ilharco, M. T. Ribeiro, M. Wortsman, L. Schmidt, H. Hajishirzi, and A. Farhadi. Editing models with task arithmetic. InInternational Conference on Learning Representations (ICLR), 2023
2023
-
[14]
S. Jaghouar, J. M. Ong, M. Basra, F. Obeid, J. Straube, M. Keiblinger, E. Bakouch, L. Atkins, M. Panahi, C. Goddard, M. Ryabinin, and J. Hagemann. INTELLECT-1 technical report.CoRR, abs/2412.01152, 2024. 12
Pith/arXiv arXiv 2024
-
[15]
Jordan, Y
K. Jordan, Y . Jin, V . Boza, J. You, F. Cesista, L. Newhouse, and J. Bernstein. Muon: An optimizer for hidden layers in neural networks. https://kellerjordan.github.io/posts/muon/,
-
[16]
Accessed: 2026-04-27
2026
-
[17]
D. P. Kingma and J. Ba. Adam: A method for stochastic optimization. InInternational Conference on Learning Representations (ICLR), 2015
2015
-
[18]
Z. Kovarik. Some iterative methods for improving orthonormality.SIAM Journal on Numerical Analysis, 7(3):386–389, 1970
1970
-
[19]
J. Li, A. Fang, G. Smyrnis, M. Ivgi, M. Jordan, S. Gadre, H. Bansal, E. Guha, S. Keh, K. Arora, S. Garg, R. Xin, N. Muennighoff, R. Heckel, J. Mercat, M. Chen, S. Gururangan, M. Wortsman, A. Albalak, Y . Bitton, M. Nezhurina, A. Abbas, C.-Y . Hsieh, D. Ghosh, J. Gardner, M. Kilian, H. Zhang, R. Shao, S. Pratt, S. Sanyal, G. Ilharco, G. Daras, K. Marathe, ...
2024
-
[20]
Lidin, A
J. Lidin, A. Sarfi, E. Miahi, Q. Anthony, S. Chauhan, E. Pappas, B. Thérien, E. Belilovsky, and S. Dare. Covenant-72b: Pre-training a 72b llm with trustless peers over-the-internet, 2026
2026
-
[21]
J. Liu, J. Su, X. Yao, Z. Jiang, G. Lai, Y . Du, Y . Qin, W. Xu, E. Lu, J. Yan, Y . Chen, H. Zheng, Y . Liu, S. Liu, B. Yin, W. He, H. Zhu, Y . Wang, J. Wang, M. Dong, Z. Zhang, Y . Kang, H. Zhang, X. Xu, Y . Zhang, Y . Wu, X. Zhou, and Z. Yang. Muon is scalable for llm training, 2025
2025
-
[22]
Loshchilov and F
I. Loshchilov and F. Hutter. Decoupled weight decay regularization. InInternational Conference on Learning Representations (ICLR), 2019
2019
-
[23]
Marczak, S
D. Marczak, S. Magistri, S. Cygert, B. Twardowski, A. D. Bagdanov, and J. van de Weijer. No task left behind: Isotropic model merging with common and task-specific subspaces. In Forty-second International Conference on Machine Learning, 2025
2025
-
[24]
McMahan, E
B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. y. Arcas. Communication- Efficient Learning of Deep Networks from Decentralized Data. In A. Singh and J. Zhu, editors, Proceedings of the 20th International Conference on Artificial Intelligence and Statistics, volume 54 ofProceedings of Machine Learning Research, pages 1273–1282. PMLR, 20–22 Apr 2017
2017
-
[25]
Nesterov
Y . Nesterov. A method for solving the convex programming problem with convergence rate O(1/k2).Proceedings of the USSR Academy of Sciences, 269:543–547, 1983
1983
-
[26]
Neyshabur, H
B. Neyshabur, H. Sedghi, and C. Zhang. What is being transferred in transfer learning? In H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, editors,Advances in Neural Information Processing Systems (NeurIPS), volume 33, pages 512–523. Curran Associates, Inc., 2020
2020
-
[27]
Paszke, S
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Te- jani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala. PyTorch: An imperative style, high-performance deep learning library. In H. Wallach, H. Larochelle, A. Beygelzimer, ...
2019
-
[28]
Pfeiffer, A
J. Pfeiffer, A. Rücklé, C. Poth, A. Kamath, I. Vuli ´c, S. Ruder, K. Cho, and I. Gurevych. AdapterHub: A framework for adapting transformers. In Q. Liu and D. Schlangen, editors, Proceedings of the 2020 Conference on Empirical Methods in Natural Language Process- ing: System Demonstrations, pages 46–54, Online, Oct. 2020. Association for Computational Lin...
2020
-
[29]
S. J. Reddi, Z. Charles, M. Zaheer, Z. Garrett, K. Rush, J. Koneˇcný, S. Kumar, and H. B. McMa- han. Adaptive federated optimization. InInternational Conference on Learning Representations, 2021
2021
- [30]
-
[31]
Sutskever, J
I. Sutskever, J. Martens, G. Dahl, and G. Hinton. On the importance of initialization and mo- mentum in deep learning. In S. Dasgupta and D. McAllester, editors,International Conference on Machine Learning (ICML), volume 28 ofProceedings of Machine Learning Research, pages 1139–1147, Atlanta, Georgia, USA, 17–19 Jun 2013. PMLR
2013
-
[32]
Thérien, X
B. Thérien, X. Huang, A. Defazio, I. Rish, and E. Belilovsky. Muloco: Muon is a practical inner optimizer for diloco, 2026
2026
-
[33]
Touvron, L
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y . Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V . Goswami, N. Goyal, A. Hartshorn, S. Hosseini, R. Hou, H. Inan, M. Kardas, V . Kerkez, M. Khabsa, I. Kloumann, A. Koren...
2023
-
[34]
T. Wolf, L. Debut, V . Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, and J. Brew. Huggingface’s transformers: State-of-the-art natural language processing.CoRR, abs/1910.03771, 2019
Pith/arXiv arXiv 1910
-
[35]
Yadav, D
P. Yadav, D. Tam, L. Choshen, C. A. Raffel, and M. Bansal. Ties-merging: Resolving interfer- ence when merging models. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors,Advances in Neural Information Processing Systems (NeurIPS), volume 36, pages 7093–7115. Curran Associates, Inc., 2023
2023
-
[36]
E. Yang, L. Shen, G. Guo, X. Wang, X. Cao, J. Zhang, and D. Tao. Model merging in LLMs, MLLMs, and beyond: Methods, theories, applications, and opportunities.ACM Comput. Surv., 58(8), Feb. 2026
2026
-
[37]
G. Yang, J. B. Simon, and J. Bernstein. A spectral condition for feature learning, 2024
2024
-
[38]
L. Yu, B. Yu, H. Yu, F. Huang, and Y . Li. Language models are super mario: Absorbing abilities from homologous models as a free lunch. In R. Salakhutdinov, Z. Kolter, K. Heller, A. Weller, N. Oliver, J. Scarlett, and F. Berkenkamp, editors,International Conference on Machine Learning (ICML), volume 235 ofProceedings of Machine Learning Research, pages 57...
2024
-
[39]
spectral
L. Zhou, D. Solombrino, D. Crisostomi, M. S. Bucarelli, G. A. D’Inverno, F. Silvestri, and E. Rodolà. On task vectors and gradients. InUniReps: 3rd Edition of the Workshop on Unifying Representations in Neural Models, 2025. 14 A Optimal Hyperparameters We report the optimal hyperparameters used for each main-text result table: inner learning rate ηin, out...
2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.