Pith. sign in

REVIEW 1 cited by

Wasserstein Barycenter-based Model Fusion and Linear Mode Connectivity of Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.06671 v1 pith:J5TAESLZ submitted 2022-10-13 cs.LG

classification cs.LG
keywords fusionmodelconnectivityframeworklinearmodenetworkbarycenter
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Based on the concepts of Wasserstein barycenter (WB) and Gromov-Wasserstein barycenter (GWB), we propose a unified mathematical framework for neural network (NN) model fusion and utilize it to reveal new insights about the linear mode connectivity of SGD solutions. In our framework, the fusion occurs in a layer-wise manner and builds on an interpretation of a node in a network as a function of the layer preceding it. The versatility of our mathematical framework allows us to talk about model fusion and linear mode connectivity for a broad class of NNs, including fully connected NN, CNN, ResNet, RNN, and LSTM, in each case exploiting the specific structure of the network architecture. We present extensive numerical experiments to: 1) illustrate the strengths of our approach in relation to other model fusion methodologies and 2) from a certain perspective, provide new empirical evidence for recent conjectures which say that two local minima found by gradient-based methods end up lying on the same basin of the loss landscape after a proper permutation of weights is applied to one of the models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Heterogeneous Federated Reinforcement Learning Using Wasserstein Barycenters

    cs.LG 2025-06 conditional novelty 4.0 of 10

    FedWB aggregates local models by computing Wasserstein barycenters of flattened, normalized weight matrices, yielding faster early convergence than FedAvg on MNIST and on heterogeneous CartPole DQN training.

Pith tools