Pith. sign in

REVIEW 1 cited by

ModalChorus: Visual Probing and Alignment of Multi-modal Embeddings via Modal Fusion Map

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.12315 v2 pith:CBYDTD4H submitted 2024-07-17 cs.CV cs.AIcs.HCcs.IR

classification cs.CVcs.AIcs.HCcs.IR
keywords embeddingsfusionmodalchorusalignmentcross-modalmulti-modalprobingclip
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multi-modal embeddings form the foundation for vision-language models, such as CLIP embeddings, the most widely used text-image embeddings. However, these embeddings are vulnerable to subtle misalignment of cross-modal features, resulting in decreased model performance and diminished generalization. To address this problem, we design ModalChorus, an interactive system for visual probing and alignment of multi-modal embeddings. ModalChorus primarily offers a two-stage process: 1) embedding probing with Modal Fusion Map (MFM), a novel parametric dimensionality reduction method that integrates both metric and nonmetric objectives to enhance modality fusion; and 2) embedding alignment that allows users to interactively articulate intentions for both point-set and set-set alignments. Quantitative and qualitative comparisons for CLIP embeddings with existing dimensionality reduction (e.g., t-SNE and MDS) and data fusion (e.g., data context map) methods demonstrate the advantages of MFM in showcasing cross-modal features over common vision-language datasets. Case studies reveal that ModalChorus can facilitate intuitive discovery of misalignment and efficient re-alignment in scenarios ranging from zero-shot classification to cross-modal retrieval and generation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Human-Guided Image Generation for Expanding Small-Scale Training Image Datasets

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A human-guided image-generation tool with contrastive multi-modal projection and sample-level prompt feedback lifted classification accuracy from 48.45% to 81.80% in a 10-class pet case study.

Pith tools