Pith. sign in

REVIEW 1 cited by

MLIP: Efficient Multi-Perspective Language-Image Pretraining with Exhaustive Data Utilization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.01460 v2 pith:3NB2H4VY submitted 2024-06-03 cs.CV cs.AI

MLIP: Efficient Multi-Perspective Language-Image Pretraining with Exhaustive Data Utilization

classification cs.CV cs.AI
keywords clipsupervisionfrequencylanguage-imagemlippretrainingtokensadditionally
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Contrastive Language-Image Pretraining (CLIP) has achieved remarkable success, leading to rapid advancements in multimodal studies. However, CLIP faces a notable challenge in terms of inefficient data utilization. It relies on a single contrastive supervision for each image-text pair during representation learning, disregarding a substantial amount of valuable information that could offer richer supervision. Additionally, the retention of non-informative tokens leads to increased computational demands and time costs, particularly in CLIP's ViT image encoder. To address these issues, we propose Multi-Perspective Language-Image Pretraining (MLIP). In MLIP, we leverage the frequency transform's sensitivity to both high and low-frequency variations, which complements the spatial domain's sensitivity limited to low-frequency variations only. By incorporating frequency transforms and token-level alignment, we expand CILP's single supervision into multi-domain and multi-level supervision, enabling a more thorough exploration of informative image features. Additionally, we introduce a token merging method guided by comprehensive semantics from the frequency and spatial domains. This allows us to merge tokens to multi-granularity tokens with a controllable compression rate to accelerate CLIP. Extensive experiments validate the effectiveness of our design.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Multi-Alignment Contrastive Learning for Enzyme--Reaction Retrieval

    q-bio.BM 2025-12 conditional novelty 5.0

    FGW-CLIP, a contrastive method that aligns enzymes and reactions while also aligning within-domain EC structure with a Gromov-Wasserstein regularizer, reports state-of-the-art retrieval on EnzymeMap and ReactZyme.