Pith. sign in

REVIEW 3 cited by

Revisiting Weight Averaging for Model Merging

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.12153 v2 pith:V64TH4WA submitted 2024-12-11 cs.LG cs.AI

classification cs.LGcs.AI
keywords modelperformancetasktasksvectorsacrossaveragingmerging
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Model merging aims to build a multi-task learner by combining the parameters of individually fine-tuned models without additional training. While a straightforward approach is to average model parameters across tasks, this often results in suboptimal performance due to interference among parameters across tasks. In this paper, we present intriguing results that weight averaging implicitly induces task vectors centered around the weight averaging itself and that applying a low-rank approximation to these centered task vectors significantly improves merging performance. Our analysis shows that centering the task vectors effectively reduces task interference and most of task-specific knowledge is concentrated in the top singular vectors. Our method demonstrates robust and scalable performance on vision benchmarks across varying numbers of tasks and model sizes. Furthermore, we observe that our approach is applicable to natural language processing tasks with competitive performance.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Task Alignment: A Simple Proxy for Practical Model Merging Across Diverse Vision Tasks

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    Task alignment serves as an efficient proxy for hyperparameter selection in model merging, accelerating the process by orders of magnitude while preserving performance in vision models with heterogeneous decoders.

  2. Can LoRA Fusion Support Cross-Domain Tasks in Cloud-Edge Collaboration?

    cs.DC 2026-04 conditional novelty 5.0 of 10

    On a new composite cross-domain benchmark, standard LoRA fusion under cloud-edge prune-train-recover often loses to the base model; a shared-subspace conflict gate recovers modest accuracy.

  3. Why Do More Experts Fail? A Theoretical Analysis of Model Merging

    cs.LG 2025-05 reject novelty 4.0 of 10

    The paper claims to prove an upper bound and diminishing returns in model merging, but the proofs are not sound and the heavy-tailed claim is contradicted by its own equations.

Pith tools