Pith. sign in

REVIEW 1 cited by

Contrastive Vision-Language Pre-training with Limited Resources

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2112.09331 v3 pith:VIXUBPVW submitted 2021-12-17 cs.CV cs.MM

classification cs.CVcs.MM
keywords datapre-trainingresourcescontrastivelimitedmethodszerovldual-encoder
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Pioneering dual-encoder pre-training works (e.g., CLIP and ALIGN) have revealed the potential of aligning multi-modal representations with contrastive learning. However, these works require a tremendous amount of data and computational resources (e.g., billion-level web data and hundreds of GPUs), which prevent researchers with limited resources from reproduction and further exploration. To this end, we propose a stack of novel methods, which significantly cut down the heavy resource dependency and allow us to conduct dual-encoder multi-modal representation alignment with limited resources. Besides, we provide a reproducible baseline of competitive results, namely ZeroVL, with only 14M publicly accessible academic datasets and 8 V100 GPUs. Additionally, we collect 100M web data for pre-training, and achieve comparable or superior results than state-of-the-art methods, further proving the effectiveness of our methods on large-scale data. We hope that this work will provide useful data points and experience for future research in contrastive vision-language pre-training. Code is available at https://github.com/zerovl/ZeroVL.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. RoofNet: A Global Multimodal Dataset for Roof Material Identification from Earth Observation

    cs.CE 2025-05 conditional novelty 7.0 of 10

    RoofNet is a multimodal dataset pairing high-resolution Earth observation imagery with roof material annotations from diverse global locations to support vision-language models for hazard exposure mapping.

Pith tools