Pith. sign in

REVIEW 1 cited by

Transferring Knowledge from Large Foundation Models to Small Downstream Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.07337 v1 pith:MZZX7DTG submitted 2024-06-11 cs.LG

classification cs.LG
keywords pre-traineddownstreammodelstransferfeaturesinformationmodelmultiple
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

How do we transfer the relevant knowledge from ever larger foundation models into small, task-specific downstream models that can run at much lower costs? Standard transfer learning using pre-trained weights as the initialization transfers limited information and commits us to often massive pre-trained architectures. This procedure also precludes combining multiple pre-trained models that learn complementary information. To address these shortcomings, we introduce Adaptive Feature Transfer (AFT). Instead of transferring weights, AFT operates purely on features, thereby decoupling the choice of the pre-trained model from the smaller downstream model. Rather than indiscriminately compressing all pre-trained features, AFT adaptively transfers pre-trained features that are most useful for performing the downstream task, using a simple regularization that adds minimal overhead. Across multiple vision, language, and multi-modal datasets, AFT achieves significantly better downstream performance compared to alternatives with a similar computational cost. Furthermore, AFT reliably translates improvement in pre-trained models into improvement in downstream performance, even if the downstream model is over $50\times$ smaller, and can effectively transfer complementary information learned by multiple pre-trained models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards Fast, Specialized Machine Learning Force Fields: Distilling Foundation Models via Energy Hessians

    physics.chem-ph 2025-01 conditional novelty 7.0 of 10

    Distilling energy Hessians from foundation model force fields produces specialized student models that are up to 20x faster with equal or better accuracy and improved energy conservation.

Pith tools