Pith. sign in

REVIEW 2 cited by

Merging Text Transformer Models from Different Initializations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.00986 v3 pith:YQ2P4H6C submitted 2024-03-01 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords modelsmergingminimamodeltransformerworkarchitecturedifferent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent work on permutation-based model merging has shown impressive low- or zero-barrier mode connectivity between models from completely different initializations. However, this line of work has not yet extended to the Transformer architecture, despite its dominant popularity in the language domain. Therefore, in this work, we investigate the extent to which separate Transformer minima learn similar features, and propose a model merging technique to investigate the relationship between these minima in the loss landscape. The specifics of the architecture, like its residual connections, multi-headed attention, and discrete, sequential input, require specific interventions in order to compute model permutations that remain within the same functional equivalence class. In merging these models with our method, we consistently find lower loss barriers between minima compared to model averaging, across models trained on a masked-language modeling task or fine-tuned on a language understanding benchmark. Our results show that the minima of these models are less sharp and isolated than previously understood, and provide a basis for future work on merging separately trained Transformer models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Merging Feed-Forward Sublayers for Compressed Transformers

    cs.CL 2025-01 conditional novelty 6.0 of 10

    Merging aligned feed-forward sublayers into tied weights can remove over a third of a Transformer's feed-forward parameters with only small performance losses, after a short fine-tuning.

  2. Training-free Heterogeneous Model Merging

    cs.LG 2024-12 conditional novelty 6.0 of 10

    A training-free framework merges heterogeneous neural networks of different depth and width by segmenting deeper models and elastically zipping neurons, reaching accuracy close to homogeneous merging.

Pith tools