Pith. sign in

REVIEW 2 cited by

On the Emergence of Cross-Task Linearity in the Pretraining-Finetuning Paradigm

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.03660 v2 pith:G7L4PZKZ submitted 2024-02-06 cs.LG cs.AI

classification cs.LGcs.AI
keywords finetunedmodelsspacelinearparadigmpretraining-finetuningapproximatelycheckpoint
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The pretraining-finetuning paradigm has become the prevailing trend in modern deep learning. In this work, we discover an intriguing linear phenomenon in models that are initialized from a common pretrained checkpoint and finetuned on different tasks, termed as Cross-Task Linearity (CTL). Specifically, we show that if we linearly interpolate the weights of two finetuned models, the features in the weight-interpolated model are often approximately equal to the linear interpolation of features in two finetuned models at each layer. We provide comprehensive empirical evidence supporting that CTL consistently occurs for finetuned models that start from the same pretrained checkpoint. We conjecture that in the pretraining-finetuning paradigm, neural networks approximately function as linear maps, mapping from the parameter space to the feature space. Based on this viewpoint, our study unveils novel insights into explaining model merging/editing, particularly by translating operations from the parameter space to the feature space. Furthermore, we delve deeper into the root cause for the emergence of CTL, highlighting the role of pretraining.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Orientation, not magnitude: the causal structure of task-vector interference in merged language models

    cs.LG 2026-08 conditional novelty 8.0 of 10

    In merged language models, task-vector interference is causally carried by the orientation of an internal cross-term, not by its magnitude, and instruction wrappers can hide this interference while the model still carries it.

  2. Parameter-Efficient Interventions for Enhanced Model Merging

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Applying ReFT-style interventions at every transformer block of a merged model improves multi-task accuracy beyond post-hoc single-layer repair, and slicing the representation keeps the parameter cost low.

Pith tools