Pith. sign in

REVIEW 6 cited by

ColD Fusion: Collaborative Descent for Distributed Multitask Finetuning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.01378 v2 pith:VS6JPITC submitted 2022-12-02 cs.LG cs.CLcs.DC

classification cs.LGcs.CLcs.DC
keywords coldfusionmultitaskdatasetsmodelmodelsbenefitscontinually
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose a new paradigm to continually evolve pretrained models, denoted ColD Fusion. It provides the benefits of multitask learning but leverages distributed computation with limited communication and eliminates the need for shared data. Consequentially, ColD Fusion can give rise to a synergistic loop, where finetuned models can be recycled to continually improve the pretrained model they are based upon. We show that ColD Fusion yields comparable benefits to multitask training by producing a model that (a) attains strong performance on all of the datasets it was trained on; and (b) is a better starting point for finetuning on unseen datasets. We show that ColD Fusion outperforms RoBERTa and even previous multitask models. Specifically, when training and testing on 35 diverse datasets, ColD Fusion-based model outperforms RoBERTa by 2.33 points on average without any changes to the architecture.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SAFE-Merge: Data-Free Continual Model Merging with General Knowledge Preservation

    cs.LG 2026-08 conditional novelty 6.0 of 10

    SAFE-Merge masks risk-prone parameter updates and recovers lost task information with a constrained low-rank correction, achieving the best H-score in data-free continual model merging benchmarks.

  2. Merge to Mix: Mixing Datasets via Model Merging

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Merge to Mix shows that the performance of a parameter-averaged model predicts the performance of a model fine-tuned on any dataset mixture, enabling fast and accurate dataset mixture selection.

  3. How to Merge Your Multimodal Models Over Time?

    cs.LG 2024-12 conditional novelty 6.0 of 10

    A systematic study of temporal model merging shows that initialization and deployment choices matter far more than the merging technique, with EMA-style weight interpolation as the best practice.

  4. Task Arithmetic Through The Lens Of One-Shot Federated Learning

    cs.LG 2024-11 conditional novelty 5.0 of 10

    Task arithmetic is exactly one-shot FedAvg with outer step size beta = lambda T, and FedNova, FedGMA, Median, and CCLIP can often improve merged model performance.

  5. Beyond Task Vectors: Selective Task Arithmetic Based on Importance Metrics

    cs.LG 2024-11 reject novelty 5.0 of 10

    STA improves task arithmetic by masking task vectors with a first-order Taylor expansion importance metric, raising average fused accuracy to 82.84% on six vision tasks.

  6. CLUES: Collaborative High-Quality Data Selection for LLMs via Training Dynamics

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A collaborative data-selection method that scores each private sample's influence on a public anchor set and filters by a global threshold before federated learning or model merging.

Pith tools