Pith. sign in

REVIEW 1 cited by

Equivariant Similarity for Vision-Language Foundation Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.14465 v2 pith:ULC63RNS submitted 2023-03-25 cs.CV

classification cs.CV
keywords equivariancesimilarityvlmseqbenexistingimage-textpairschallenging
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This study explores the concept of equivariance in vision-language foundation models (VLMs), focusing specifically on the multimodal similarity function that is not only the major training objective but also the core delivery to support downstream tasks. Unlike the existing image-text similarity objective which only categorizes matched pairs as similar and unmatched pairs as dissimilar, equivariance also requires similarity to vary faithfully according to the semantic changes. This allows VLMs to generalize better to nuanced and unseen multimodal compositions. However, modeling equivariance is challenging as the ground truth of semantic change is difficult to collect. For example, given an image-text pair about a dog, it is unclear to what extent the similarity changes when the pixel is changed from dog to cat? To this end, we propose EqSim, a regularization loss that can be efficiently calculated from any two matched training pairs and easily pluggable into existing image-text retrieval fine-tuning. Meanwhile, to further diagnose the equivariance of VLMs, we present a new challenging benchmark EqBen. Compared to the existing evaluation sets, EqBen is the first to focus on "visual-minimal change". Extensive experiments show the lack of equivariance in current VLMs and validate the effectiveness of EqSim. Code is available at https://github.com/Wangt-CN/EqBen.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Insect-Foundation: A Foundation Model and Large Multimodal Dataset for Vision-Language Insect Understanding

    cs.CV 2025-02 conditional novelty 5.0 of 10

    An insect-specific vision-language assistant trained on a new 1M-image multimodal insect dataset with patch-matching self-supervision reports improved insect classification and VQA accuracy over LLaVA and prior self-s...

Pith tools