Pith. sign in

REVIEW 2 cited by

APE: Aligning Pretrained Encoders to Quickly Learn Aligned Multimodal Representations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.03927 v1 pith:WSBJCZTV submitted 2022-10-08 cs.LG

classification cs.LG
keywords datalessencoderstrainingachievealignedaligningalignment
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advances in learning aligned multimodal representations have been primarily driven by training large neural networks on massive, noisy paired-modality datasets. In this work, we ask whether it is possible to achieve similar results with substantially less training time and data. We achieve this by taking advantage of existing pretrained unimodal encoders and careful curation of alignment data relevant to the downstream task of interest. We study a natural approach to aligning existing encoders via small auxiliary functions, and we find that this method is competitive with (or outperforms) state of the art in many settings while being less prone to overfitting, less costly to train, and more robust to distribution shift. With a properly chosen alignment distribution, our method surpasses prior state of the art for ImageNet zero-shot classification on public data while using two orders of magnitude less time and data and training 77% fewer parameters.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. QuARI: Query Adaptive Retrieval Improvement

    cs.CV 2025-05 conditional novelty 7.0 of 10

    A hypernetwork predicts a query-specific low-rank linear projection that reshapes frozen VLM embeddings, improving retrieval on ILIAS and INQUIRE.

  2. (Almost) Free Modality Stitching of Foundation Models

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A hypernetwork that generates connector weights for all image-text model pairs can rank pairs like grid search at about 10x lower training cost, but the best connector lags grid search by a few points.

Pith tools