Pith. sign in

Context-Aware Multimodal Pretraining

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Large-scale multimodal representation learning successfully optimizes for zero-shot transfer at test time. Yet the standard pretraining paradigm (contrastive learning on large amounts of image-text data) does not explicitly encourage representations to support few-shot adaptation. In this work, we propose a simple, but carefully designed extension to multimodal pretraining which enables representations to accommodate additional context. Using this objective, we show that vision-language models can be trained to exhibit significantly increased few-shot adaptation: across 21 downstream tasks, we find up to four-fold improvements in test-time sample efficiency, and average few-shot adaptation gains of over 5%, while retaining zero-shot generalization performance across model scales and training durations. In particular, equipped with simple, training-free, metric-based adaptation mechanisms, our representations easily surpass more complex and expensive optimization-based schemes, vastly simplifying generalization to new domains.

citation-role summary

background 1

citation-polarity summary

fields

cs.LG 1

years

2024 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

unclear 1

representative citing papers

How to Merge Your Multimodal Models Over Time?

cs.LG · 2024-12-09 · conditional · novelty 6.0

A systematic study of temporal model merging shows that initialization and deployment choices matter far more than the merging technique, with EMA-style weight interpolation as the best practice.

citing papers explorer

Showing 1 of 1 citing paper.

  • How to Merge Your Multimodal Models Over Time? cs.LG · 2024-12-09 · conditional · none · ref 66 · internal anchor

    A systematic study of temporal model merging shows that initialization and deployment choices matter far more than the merging technique, with EMA-style weight interpolation as the best practice.