Pith. sign in

REVIEW 1 cited by

CSA: Data-efficient Mapping of Unimodal Features to Multimodal Features

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.07610 v4 pith:H2XSYXZY submitted 2024-10-10 cs.LG cs.AIcs.CVcs.IR

classification cs.LGcs.AIcs.CVcs.IR
keywords multimodalunimodaldataencodersfeaturesclassificationclipimage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Multimodal encoders like CLIP excel in tasks such as zero-shot image classification and cross-modal retrieval. However, they require excessive training data. We propose canonical similarity analysis (CSA), which uses two unimodal encoders to replicate multimodal encoders using limited data. CSA maps unimodal features into a multimodal space, using a new similarity score to retain only the multimodal information. CSA only involves the inference of unimodal encoders and a cubic-complexity matrix decomposition, eliminating the need for extensive GPU-based model training. Experiments show that CSA outperforms CLIP while requiring $50,000\times$ fewer multimodal data pairs to bridge the modalities given pre-trained unimodal encoders on ImageNet classification and misinformative news caption detection. CSA surpasses the state-of-the-art method to map unimodal features to multimodal features. We also demonstrate the ability of CSA with modalities beyond image and text, paving the way for future modality pairs with limited paired multimodal data but abundant unpaired unimodal data, such as lidar and text.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Any2Any: Incomplete Multimodal Retrieval with Conformal Prediction

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Any2Any uses two-stage conformal prediction to align incomparable cross-modal similarity scores, enabling retrieval with incomplete multimodal queries and references.

Pith tools