REVIEW 4 major objections 5 minor 62 references
Heterogeneous LLM merging can succeed with deterministic dimensional adaptation plus small-ratio weighted averaging, without training or semantic alignment.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 16:18 UTC pith:EJZEQEPW
load-bearing objection Worth a look for the ratio-regime idea, but the headline numbers are per-task maxima over the ratio grid—no single merged model achieves them. the 4 major comments →
Rethinking Heterogeneous LLM Merging: A Weighted Model Averaging Perspective
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper claims that two differently sized LLMs from the same family can be merged by first projecting one into the other's parameter space—expanding the smaller or truncating the larger—and then interpolating with a small mixing ratio. On Qwen-family pairs from 3B to 32B, deterministic expansion largely preserves the smaller model's behavior, and small-ratio interpolation transfers complementary capabilities: for example, union merging a 14B and a 32B checkpoint reaches an average of 0.7525 versus 0.7044 and 0.7424 for the two sources, and intersection merging improves a 3B model from 0.5459 to 0.5716 by injecting a truncated 32B branch with a small coefficient. The paper also reports that
What carries the argument
The central machinery is a pair of deterministic tensor maps followed by convex interpolation. EXPAND places the smaller checkpoint into the larger architecture by copying compatible tensors, zero-filling new attention and MLP coordinates, and initializing added layers as residual identities; TRUNCATE projects the larger checkpoint into the smaller architecture by copying shape-compatible tensors, slicing oversized tensors along head and MLP axes, and dropping extra layers. Interpolation is performed as theta_union(lambda) = (1-lambda) E(theta_small) + lambda theta_large or theta_intersection(mu) = (1-mu) theta_small + mu T(theta_large) with small ratios. This works by keeping the merged poi
Load-bearing premise
The load-bearing premise is that after expansion or truncation, copied and sliced weights still carry out the same computation in both checkpoints; if tensor names, module roles, and head structure do not align, small-ratio averaging loses its semantic basis and the reported gains disappear.
What would settle it
A decisive check: choose lambda and mu on a held-out validation split and only then evaluate on the benchmarks used for reporting; if the selected merged checkpoint no longer beats both source averages, the headline gains are selection artifacts rather than transferred capability.
If this is right
- A zero-training, zero-alignment merging recipe exists for architecture-compatible LLM families: deterministic expansion or truncation plus small-ratio interpolation can yield a single checkpoint that beats both source models on average.
- Union-style merging has two usable endpoint neighborhoods (lambda below about 0.1 and above about 0.9), while intersection-style merging is usable only at very small mu; near-balanced interpolation should be avoided.
- Intersection merging transfers most when the capability gap is large, as when a truncated 32B branch improves a 3B model from 0.5459 to 0.5716, while smaller gaps produce milder gains.
- Aggregate improvements hide task-level regressions, so a merged checkpoint must be validated per task before deployment even when its average score exceeds both sources.
Where Pith is reading between the lines
- If the ratio-sensitivity phase pattern holds across more model families, the collapse boundary could be predicted from representation-similarity or mode-connectivity measures, letting practitioners locate safe mixing ratios without a full benchmark sweep.
- The same expand-and-truncate operators could be applied to checkpoints whose hidden dimensions are permuted rather than aligned; if gains vanish under a fixed permutation, the transfer depends on matching coordinate roles, not merely on adding structured perturbation.
- A testable extension is to use the truncated large model as a parameter-space prior and then briefly fine-tune the merged checkpoint; if small-ratio averaging already captures complementary skills, it may also provide a better starting point for further training.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes training-free heterogeneous LLM merging by deterministic dimensional adaptation (expansion for union-style, truncation for intersection-style) followed by ratio-controlled convex interpolation between the adapted checkpoints. It reports experiments on Qwen2.5/Qwen3 model pairs (3B–32B) across math, code, NLU, commonsense, knowledge, and instruction-following benchmarks, claiming that small-ratio interpolation can improve over both source checkpoints without training, adapters, routing, or semantic alignment. The paper also identifies ratio-sensitive regimes (preservation, transfer, collapse) and a task-level seesaw effect.
Significance. If the central claim were fully supported, the paper would establish a remarkably simple baseline: deterministic tensor-slot expansion/truncation plus small-ratio weighted averaging can merge heterogeneous LLM checkpoints with zero training and zero alignment. The protocol is clearly specified in Eqs. (7)–(8), and the appendix gives unusually detailed per-ratio task-level scores (Tables 10–16), which is a genuine strength. The perturbation controls in Table 19 and the comparison with TIES, SLERP, and DARE in Table 20 are also useful diagnostics. However, the headline quantitative claims are not supported as stated: the 'Merged' columns in Tables 3–9 are per-task maxima over the ratio grid, not the scores of any single checkpoint. When a single fixed-ratio model is considered, the reported average improvements shrink to at most about 0.005–0.01, with no error bars, and the ratio is selected on the same benchmarks later reported as the outcome. The qualitative phenomena may survive, but the paper’s main evidence needs substantial revision.
major comments (4)
- [§5 and Tables 3–9 vs. Appendix A (Tables 10–16)] The 'Merged' columns are not scores of a single checkpoint; they are per-task bests over the ratio grid. For P1, Table 3's gsm8k=0.9045, humanevalplus=0.6494, ifeval=0.5072, and BBH=0.8241 come from different λ values (0.98, 0.96, 0.96, 0.96 respectively), and the average 0.7525 equals the mean of the BEST column in Table 10. The best single-λ average is 0.7481 at λ=0.98 versus source B at 0.7424. Likewise, Table 6's 0.5716 is the mean of per-task bests from Table 13; the best single-μ average is 0.5554 at μ=0.01 versus 0.5459. The abstract and §5 make claims about a 'best merged model' that no actual model realizes. Please re-report all headline results as single fixed-ratio checkpoints, or explicitly relabel these as per-task oracle upper bounds and adjust every conclusion accordingly.
- [§3.4, Eqs. (5)–(6)] The ratio λ (or μ) is selected by minimizing L_eval on the same benchmark suite that is later reported as the outcome, and the 'best' result is then the maximum over the grid. This makes the reported gains in-sample selection artifacts rather than independent predictions. Even when a single ratio is used, choosing it on the test benchmarks overstates expected transfer. The paper needs a proper validation split, or at minimum it must present all grid averages and frame the results as a descriptive upper envelope. The perturbation controls in Table 19 are helpful but do not address this selection issue.
- [§6.2 and Tables 21–22] The claimed phase diagram (λ∈(0,0.1)∪(0.9,1) stable, λ∈(0.1,0.9) collapse; μ∈(0,0.1) stable, μ∈(0.1,1) collapse) is based on one GSM8K pair (Qwen2.5-14B/32B) in the appendix, with no task-level or pair-level corroboration in the main text. The terms 'flexible' and 'strict' in Tables 21–22 are never defined. If ratio sensitivity is a central contribution, it needs definitions, error bars, and evidence across more than one task/model pair before it can be presented as a general regime.
- [§5.2 and Tables 6–9] Once the per-task-best aggregation is corrected, the claimed intersection-style gains mostly vanish. For example, Table 14 shows the best single-μ average for Qwen2.5-14B is 0.7086 (μ=0.04) versus baseline 0.7045, a gain of 0.0041; Table 16 shows Qwen3-4B improving from 0.6277 to 0.6292, a gain of 0.0015. These margins are within evaluation noise, and no error bars or repeated runs are reported for any main table. The statement that truncation injection 'provides a useful performance gain' needs statistical support or must be substantially softened.
minor comments (5)
- [Tables 3–9] The tables contain unexplained parenthetical annotations such as '(+0.03)', '(+0.01)', and '(+0.02)' next to Base B or Base A entries. These appear to be editing remnants and should be removed or explained. Also, 'A vg.' should be 'Avg.'.
- [Appendix C, Tables 21–22] Please define 'flexible' and 'strict' and describe how the reported variance is computed. Without definitions, these diagnostic tables are hard to interpret.
- [Figure 3 and §5.1] The figure claims to clarify 'knowledge-structure compatibility', but no method is given for computing the competence profiles or their overlap. As presented, it is an illustration rather than evidence; please either add a concrete computation or label it as a schematic.
- [Appendix B.2 (last paragraph)] The assumption that tensor names, module roles, and head structure are 'sufficiently comparable' is stated only in the appendix. Since this substantially narrows the meaning of 'heterogeneous', it should be stated prominently in the abstract or introduction as a scope condition.
- [Reproducibility] No code, checkpoint, or artifact link is provided, and no computational details (e.g., evaluation harness versions, number of runs) are given. Please include these for reproducibility.
Circularity Check
Headline 'best merged' numbers are per-task maxima over the searched ratio grid, not scores of any single checkpoint; selection uses the same benchmarks later reported.
specific steps
-
fitted input called prediction
[Section 3.4, Eqs. (5)-(6); Section 4 ratio-grid paragraph; Appendix A Tables 10/13]
"For each protocol, we search over its own ratio grid and select the best checkpoint by downstream validation performance: θ*_∪ = arg min_{λ∈Λ} L_eval(θ_∪(λ)) ... We use fixed ratio grids ... λ∈ {0.02,0.04,0.96,0.98} ... µ∈ {0.01,0.02,0.03,0.04,0.05} ... The merged checkpoints are evaluated on the benchmark families reported for each protocol. Table 10: Task-level best/worst outcomes ... A vg. 0.7044 0.7424 0.7481 0.7479 0.7093 0.7085 0.7525 0.7036."
The 'Merged' columns in Tables 3-5 are the Appendix-A 'BEST' columns, i.e. for each task the maximum over the searched ratios. The reported average 0.7525 is the mean of these per-task maxima; the best fixed-λ average is 0.7481 (λ=0.98). Thus the headline 'best merged model reaches 0.7525' is the maximum over the grid on the same evaluation benchmarks used for selection, not an independent outcome of a single checkpoint. The selection rule in Eqs. (5)-(6) makes the reported best value a fitted maximum by construction.
-
fitted input called prediction
[Section 5; Section 5.2; Table 6 and Appendix A Table 13]
"Intersection-style merging also yields positive gains when the truncated larger checkpoint is injected with a small coefficient: Qwen2.5-3B improves from 0.5459 to 0.5716 ... The merged model is computed as (1−µ) Qwen2.5-3B + µ Qwen2.5-32B-truncated, where µ is selected from {0.01,0.02,0.03,0.04,0.05}. Table 13: A vg. 0.5459 0.5554 0.5551 0.5514 0.5523 0.5459 (+0.03) 0.5716 0.5336."
The claimed improvement from 0.5459 to 0.5716 is the mean of the per-task BEST column over µ∈{0.01,...,0.05}; no single µ achieves it. The best fixed-µ average is 0.5554 (µ=0.01), only 0.0095 above baseline. Therefore the headline intersection gain is a per-task maximum over the grid on the same benchmarks used for selection, i.e. a fitted extreme, not a property of one merged checkpoint.
full rationale
The paper's central quantitative claims are partially circular in a specific, quotable way: Eqs. (5)-(6) select λ/µ by minimizing L_eval on the downstream benchmarks, and the 'Merged' columns in Tables 3-5 and 6-9 are per-task BEST values from Appendix A, i.e., maxima over the searched ratio grid. Consequently, the abstract/Section 5 statements that 'the best merged model reaches 0.7525' (union P1) and that intersection improves Qwen2.5-3B from 0.5459 to 0.5716 are not scores of any single checkpoint; they are the mean of task-wise grid maxima. The best fixed-ratio averages are much smaller (0.7481 for P1 at λ=0.98; 0.5554 for P2 at µ=0.01), and the same grids also produce WORST averages below baseline, confirming that no one ratio yields the reported headline. This is the 'fitted input called prediction' pattern: the reported 'prediction' is forced by the argmin/max selection on the same evaluation data. I found no load-bearing self-citation or imported uniqueness theorem; the references are external and the qualitative small-ratio interpolation results have independent content (e.g., fixed-ratio rows, perturbation controls, baseline comparisons). The issue is specifically that the strongest numbers are maxima over the selection grid and are presented as if they were properties of a merged model. This warrants a score of 7: the headline results reduce by construction, though the paper also contains non-circular smaller claims.
Axiom & Free-Parameter Ledger
free parameters (4)
- union mixing ratio lambda =
lambda in {0.02, 0.04, 0.96, 0.98}; per-task best selected via argmin L_eval (Eq. 5)
- intersection mixing ratio mu =
mu in {0.01, 0.02, 0.03, 0.04, 0.05}; per-task best selected via argmin L_eval (Eq. 6)
- expansion slot map pi_{n->m}(i) = floor(i*m/n) =
deterministic given source and target configs
- normalization copy scale c =
1 or sqrt(h_s/h_t)
axioms (4)
- domain assumption Qwen-family checkpoints share tensor names, module roles, and head structure sufficiently for copy/slice to map corresponding computation
- domain assumption Zero-initialized new branches and residual-identity new layers preserve the source model function
- domain assumption Ratio selection on the evaluation benchmarks is a valid proxy for merging quality and does not overfit
- domain assumption Small-ratio interpolation operates in a locally Lipschitz region around the anchor weights
read the original abstract
Can large language models with substantially different parameter spaces be merged by direct weighted averaging, without training or semantic alignment? Existing heterogeneous fusion methods typically introduce distillation, adapters, learned latent spaces, routing, or feature alignment, leaving open whether a simpler recipe can work for genuinely different billion-parameter checkpoints. We revisit this counterintuitive question through training-free dimensional adaptation followed by ratio-controlled interpolation. In union-style merging, we expand the smaller model into the larger parameter space; in intersection-style merging, we truncate the larger model into the smaller parameter space. Across Qwen-family model pairs and benchmarks covering mathematical reasoning, code generation, language understanding, commonsense reasoning, knowledge, and instruction following, deterministic expansion largely preserves the source model function, and small-ratio interpolation can improve over strong source checkpoints by transferring complementary capabilities. However, near-balanced interpolation often collapses, and task-level results reveal a seesaw effect in which gains on some capabilities coexist with regressions on others. These results show that simple parameter averaging, when paired with lightweight dimensional adaptation and carefully controlled ratios, is a surprisingly strong baseline for heterogeneous LLM merging, suggesting that the limits of direct weighted fusion may also bound what more complex heterogeneous merging methods can achieve at scale.
Figures
Reference graph
Works this paper leans on
-
[1]
Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebron, and Sumit Sanghai. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.298 GQA : Training generalized multi-query transformer models from multi-head checkpoints . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 4895--4901
-
[2]
Ainsworth, Jonathan Hayase, and Siddhartha Srinivasa
Samuel K. Ainsworth, Jonathan Hayase, and Siddhartha Srinivasa. 2023. https://arxiv.org/abs/2209.04836 Git Re-Basin : Merging models modulo permutation symmetries . In Proceedings of the 11th International Conference on Learning Representations
Pith/arXiv arXiv 2023
-
[3]
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016. https://doi.org/10.48550/arXiv.1607.06450 Layer normalization . arXiv preprint arXiv:1607.06450
-
[4]
Loubna Ben Allal, Niklas Muennighoff, Logesh Kumar Umapathi, Ben Lipkin, and Leandro von Werra. 2022. https://github.com/bigcode-project/bigcode-evaluation-harness A framework for the evaluation of code generation models . https://github.com/bigcode-project/bigcode-evaluation-harness
2022
-
[5]
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. 2020. https://ojs.aaai.org/index.php/AAAI/article/view/6239 PIQA : Reasoning about physical commonsense in natural language . In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 7432--7439
2020
-
[6]
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, and 1 others. 2021. https://doi.org/10.48550/arXiv.2107.03374 Evaluating large language models trained on code . arXiv preprint arXiv:2107.03374
-
[7]
Shilian Chen, Jie Zhou, Qin Chen, Wen Wu, Xin Li, Qi Feng, and Liang He. 2026. https://doi.org/10.48550/arXiv.2604.01674 Can heterogeneous language models be fused? arXiv preprint arXiv:2604.01674
-
[8]
Zhijun Chen, Xiaodong Lu, Jingzheng Li, Pengpeng Chen, Zhuoran Li, Kai Sun, Yuankai Luo, Qianren Mao, Ming Li, Likang Xiao, Dingqi Yang, Xiao Huang, Yikun Ban, Hailong Sun, and Philip S. Yu. 2025. https://doi.org/10.48550/arXiv.2502.18036 Harnessing multiple large language models: A survey on LLM ensemble . arXiv preprint arXiv:2502.18036
-
[9]
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019. https://doi.org/10.18653/v1/N19-1300 BoolQ : Exploring the surprising difficulty of natural yes/no questions . In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Te...
-
[10]
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. https://doi.org/10.48550/arXiv.1803.05457 Think you have solved question answering? try ARC , the AI2 reasoning challenge . arXiv preprint arXiv:1803.05457
-
[11]
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. https://doi.org/10.48550/arXiv.2110.14168 Training verifiers to solve math word problems . arXiv preprint arXiv:2110.14168
-
[12]
MohammadReza Davari and Eugene Belilovsky. 2024. https://arxiv.org/abs/2312.06795 Model breadcrumbs: Scaling multi-task model merging with sparse masks . In Proceedings of the European Conference on Computer Vision
Pith/arXiv arXiv 2024
-
[13]
Guodong Du, Zhuo Li, Xuanning Zhou, Junlin Li, Zesheng Shi, Wanyu Lin, Ho-Kin Tang, Xiucheng Li, Fangming Liu, Wenya Wang, Min Zhang, and Jing Li. 2025 a . https://doi.org/10.48550/arXiv.2505.18502 Knowledge fusion of large language models via modular SkillPacks . arXiv preprint arXiv:2505.18502
-
[14]
Yiyang Du, Xiaochen Wang, Chi Chen, Jiabo Ye, Yiru Wang, Peng Li, Ming Yan, Ji Zhang, Fei Huang, Zhifang Sui, Maosong Sun, and Yang Liu. 2025 b . https://doi.org/10.48550/arXiv.2503.23733 AdaMMS : Model merging for heterogeneous multimodal large language models with unsupervised coefficient optimization . arXiv preprint arXiv:2503.23733
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2503.23733 2025
-
[15]
Jonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, and Michael Carbin. 2020. https://arxiv.org/abs/1912.05671 Linear mode connectivity and the lottery ticket hypothesis . In Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pages 3259--3269
Pith/arXiv arXiv 2020
-
[16]
Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac'h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, and 5 others. 2023. https://doi.org/10.5281/zenodo.10256836 A framework...
-
[17]
Vetrov, and Andrew Gordon Wilson
Timur Garipov, Pavel Izmailov, Dmitrii Podoprikhin, Dmitry P. Vetrov, and Andrew Gordon Wilson. 2018. https://arxiv.org/abs/1802.10026 Loss surfaces, mode connectivity, and fast ensembling of DNN s . In Advances in Neural Information Processing Systems, volume 31
Pith/arXiv arXiv 2018
-
[18]
Jian Gu, Aldeida Aleti, Chunyang Chen, and Hongyu Zhang. 2025. https://doi.org/10.48550/arXiv.2505.20144 SeMe : Training-free language model merging via semantic alignment . arXiv preprint arXiv:2505.20144
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2505.20144 2025
-
[19]
Stefan Hackmann. 2024. https://doi.org/10.48550/arXiv.2409.19173 HM3 : Heterogeneous multi-class model merging . arXiv preprint arXiv:2409.19173
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2409.19173 2024
-
[20]
Yifei He, Yuzheng Hu, Yong Lin, Tong Zhang, and Han Zhao. 2024. https://doi.org/10.48550/arXiv.2408.13656 Localize-and-stitch: Efficient model merging via sparse task arithmetic . Transactions on Machine Learning Research
-
[21]
Yifei He, Siqi Zeng, Yuzheng Hu, Rui Yang, Tong Zhang, and Han Zhao. 2025. https://doi.org/10.48550/arXiv.2505.10833 MergeBench : A benchmark for merging domain-specialized LLM s . In Advances in Neural Information Processing Systems
-
[22]
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020. https://doi.org/10.48550/arXiv.2009.03300 Measuring massive multitask language understanding . arXiv preprint arXiv:2009.03300
-
[23]
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basu, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021. https://doi.org/10.48550/arXiv.2103.03874 Measuring mathematical problem solving with the MATH dataset . arXiv preprint arXiv:2103.03874
-
[24]
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015. https://doi.org/10.48550/arXiv.1503.02531 Distilling the knowledge in a neural network . arXiv preprint arXiv:1503.02531
-
[25]
O g uz Ka g an Hitit, Leander Girrbach, and Zeynep Akata. 2026. https://doi.org/10.48550/arXiv.2511.21437 A systematic study of in-the-wild model merging for large language models . Transactions on Machine Learning Research
-
[26]
Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Suchin Gururangan, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi. 2023. https://arxiv.org/abs/2212.04089 Editing models with task arithmetic . In Proceedings of the 11th International Conference on Learning Representations
Pith/arXiv arXiv 2023
-
[27]
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, L \'e lio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, and 7 others. 2024...
-
[28]
Xisen Jin, Xiang Ren, Daniel Preotiuc-Pietro, and Pengxiang Cheng. 2023. https://arxiv.org/abs/2212.09849 Dataless knowledge fusion by merging weights of language models . In Proceedings of the 11th International Conference on Learning Representations
Pith/arXiv arXiv 2023
-
[29]
Keller Jordan, Hanie Sedghi, Olga Saukh, Rahim Entezari, and Behnam Neyshabur. 2023. https://arxiv.org/abs/2211.08403 REPAIR : Renormalizing permuted activations for interpolation repair . In Proceedings of the 11th International Conference on Learning Representations
Pith/arXiv arXiv 2023
-
[30]
Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Geoffrey Hinton. 2019. https://arxiv.org/abs/1905.00414 Similarity of neural network representations revisited . In Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pages 3519--3529
Pith/arXiv arXiv 2019
-
[31]
Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. 2024. https://arxiv.org/abs/2305.20050 Let's verify step by step . In Proceedings of the 12th International Conference on Learning Representations
Pith/arXiv arXiv 2024
-
[32]
Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang. 2023. https://arxiv.org/abs/2305.01210 Is your code generated by ChatGPT really correct? rigorous evaluation of large language models for code generation . In Advances in Neural Information Processing Systems, volume 36
Pith/arXiv arXiv 2023
-
[33]
Jinliang Lu, Ziliang Pang, Min Xiao, Yaochen Zhu, Rui Xia, and Jiajun Zhang. 2024. https://doi.org/10.48550/arXiv.2407.06089 Merge, ensemble, and cooperate! A survey on collaborative strategies in the era of large language models . arXiv preprint arXiv:2407.06089
-
[34]
Michael Matena and Colin Raffel. 2022. https://doi.org/10.48550/arXiv.2111.09832 Merging models with fisher-weighted averaging . In Advances in Neural Information Processing Systems, volume 35
-
[35]
Jonas Pfeiffer, Aishwarya Kamath, Andreas R \"u ckl \'e , Kyunghyun Cho, and Iryna Gurevych. 2021. https://doi.org/10.18653/v1/2021.eacl-main.39 AdapterFusion : Non-destructive task composition for transfer learning . In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 487--503
-
[36]
Maithra Raghu, Justin Gilmer, Jason Yosinski, and Jascha Sohl-Dickstein. 2017. https://arxiv.org/abs/1706.05806 SVCCA : Singular vector canonical correlation analysis for deep learning dynamics and interpretability . In Advances in Neural Information Processing Systems, volume 30
Pith/arXiv arXiv 2017
-
[37]
Melissa Roemmele, Cosmin Adrian Bejan, and Andrew S. Gordon. 2011. https://commonsensereasoning.org/2011/papers/Roemmele.pdf Choice of plausible alternatives: An evaluation of commonsense causal reasoning . In AAAI Spring Symposium: Logical Formalizations of Commonsense Reasoning
2011
-
[38]
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2020. https://ojs.aaai.org/index.php/AAAI/article/view/6399 WinoGrande : An adversarial Winograd schema challenge at scale . In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 8732--8740
2020
-
[39]
Ken Shoemake. 1985. https://doi.org/10.1145/325165.325242 Animating rotation with quaternion curves . In Proceedings of the 12th Annual Conference on Computer Graphics and Interactive Techniques, pages 245--254
arXiv 1985
-
[40]
Mingyang Song and Mao Zheng. 2026. https://doi.org/10.48550/arXiv.2603.09938 Model merging in the era of large language models: Methods, applications, and future directions . arXiv preprint arXiv:2603.09938
-
[41]
Bedionita Soro, Aoxuan Silvia Zhang, Bruno Andreis, Jaehyeong Jo, Song Chong, and Sung Ju Hwang. 2026. https://openreview.net/forum?id=VSDV0SWwOC LS-Merge : Merging language models in latent space . In Proceedings of the 14th International Conference on Learning Representations
2026
-
[42]
George Stoica, Daniel Bolya, Jakob Bjorner, Pratik Ramesh, Taylor Hearn, and Judy Hoffman. 2024. https://arxiv.org/abs/2305.03053 ZipIt ! merging models from different tasks without training . In Proceedings of the 12th International Conference on Learning Representations
Pith/arXiv arXiv 2024
-
[43]
Mirac Suzgun, Nathan Scales, Nathanael Sch \"a rli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V. Le, Ed H. Chi, Denny Zhou, and Jason Wei. 2022. https://doi.org/10.48550/arXiv.2210.09261 Challenging BIG-Bench tasks and whether chain-of-thought can solve them . arXiv preprint arXiv:2210.09261
-
[44]
Derek Tam, Mohit Bansal, and Colin Raffel. 2024. https://doi.org/10.48550/arXiv.2312.04339 Merging by matching models in task parameter subspaces . Transactions on Machine Learning Research
-
[45]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. https://arxiv.org/abs/1706.03762 Attention is all you need . In Advances in Neural Information Processing Systems, volume 30
Pith/arXiv arXiv 2017
-
[46]
Fanqi Wan, Xinting Huang, Deng Cai, Xiaojun Quan, Wei Bi, and Shuming Shi. 2024 a . https://doi.org/10.48550/arXiv.2401.10491 Knowledge fusion of large language models . arXiv preprint arXiv:2401.10491
-
[47]
Fanqi Wan, Longguang Zhong, Ziyi Yang, Ruijun Chen, and Xiaojun Quan. 2024 b . https://doi.org/10.48550/arXiv.2408.07990 FuseChat : Knowledge fusion of chat models . arXiv preprint arXiv:2408.07990
-
[48]
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019. https://arxiv.org/abs/1804.07461 GLUE : A multi-task benchmark and analysis platform for natural language understanding . In International Conference on Learning Representations
Pith/arXiv arXiv 2019
-
[49]
Ke Wang, Nikolaos Dimitriadis, Guillermo Ortiz-Jimenez, Fran c ois Fleuret, and Pascal Frossard. 2024. https://doi.org/10.48550/arXiv.2405.07813 Localizing task information for improved model merging and compression . In Proceedings of the 41st International Conference on Machine Learning
-
[50]
Yuanyi Wang, Yanggan Gu, Yiming Zhang, Qi Zhou, Zhaoyi Yan, Congkai Xie, Xinyao Wang, Jianbo Yuan, and Hongxia Yang. 2025. https://doi.org/10.48550/arXiv.2509.24244 Model merging scaling laws in large language models . arXiv preprint arXiv:2509.24244
-
[51]
Adina Williams, Nikita Nangia, and Samuel R. Bowman. 2018. https://doi.org/10.18653/v1/N18-1101 A broad-coverage challenge corpus for sentence understanding through inference . In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1112--1122
-
[52]
Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, and Ludwig Schmidt
Mitchell Wortsman, Gabriel Ilharco, Samir Ya Gadre, Rebecca Roelofs, Raphael Gontijo-Lopes, Ari S. Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, and Ludwig Schmidt. 2022. https://proceedings.mlr.press/v162/wortsman22a.html Model soups: Averaging weights of multiple fine-tuned models improves accuracy without increasing inference ti...
2022
-
[53]
Prateek Yadav, Derek Tam, Leshem Choshen, Colin Raffel, and Mohit Bansal. 2023. https://arxiv.org/abs/2306.01708 TIES-Merging : Resolving interference when merging models . In Advances in Neural Information Processing Systems, volume 36
Pith/arXiv arXiv 2023
-
[54]
Prateek Yadav, Tu Vu, Jonathan Lai, Alexandra Chronopoulou, Manaal Faruqui, Mohit Bansal, and Tsendsuren Munkhdalai. 2025. https://doi.org/10.48550/arXiv.2410.03617 What matters for model merging at scale? arXiv preprint arXiv:2410.03617
-
[55]
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, and 1 others. 2025. https://doi.org/10.48550/arXiv.2505.09388 Qwen3 technical report . arXiv preprint arXiv:2505.09388
-
[56]
Qwen Team. 2025. https://huggingface.co/Qwen/Qwen3-4B-Thinking-2507 Qwen3-4B-Thinking-2507 . Hugging Face model card
2025
-
[57]
An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, and 1 others. 2024. https://doi.org/10.48550/arXiv.2412.15115 Qwen2.5 technical report . arXiv preprint arXiv:2412.15115
-
[58]
Enneng Yang, Li Shen, Guibing Guo, Xingwei Wang, Xiaochun Cao, Jie Zhang, and Dacheng Tao. 2026. https://doi.org/10.1145/3787849 Model merging in LLMs , MLLMs , and beyond: Methods, theories, applications, and opportunities . ACM Computing Surveys, 58(8)
doi:10.1145/3787849 2026
-
[59]
Le Yu, Bowen Yu, Haiyang Yu, Fei Huang, and Yongbin Li. 2024. https://arxiv.org/abs/2311.03099 Language models are Super Mario : Absorbing abilities from homologous models as a free lunch . In Proceedings of the 41st International Conference on Machine Learning
Pith/arXiv arXiv 2024
-
[60]
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. https://doi.org/10.18653/v1/P19-1472 HellaSwag : Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791--4800
-
[61]
Biao Zhang and Rico Sennrich. 2019. https://arxiv.org/abs/1910.07467 Root mean square layer normalization . In Advances in Neural Information Processing Systems, volume 32
Pith/arXiv arXiv 2019
-
[62]
Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, Sujoy Basu, Yi Luan, Denny Zhou, and Le Hou. 2023. https://doi.org/10.48550/arXiv.2311.07911 Instruction-following evaluation for large language models . arXiv preprint arXiv:2311.07911
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.