A two-level permutation strategy over attention heads and their inner units transfers task vectors from an old pretrained Transformer to a new one, data-free.
Reproducible scaling laws for contrastive language-image learning
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Update Your Transformer to the Latest Release: Re-Basin of Task Vectors
A two-level permutation strategy over attention heads and their inner units transfers task vectors from an old pretrained Transformer to a new one, data-free.