Pith. sign in

WIT: wikipedia-based image text dataset for multimodal multilingual machine learning

2 Pith papers cite this work. Polarity classification is still indexing.

2 Pith papers citing it

fields

cs.CV 2

years

2026 1 2023 1

representative citing papers

Sigmoid Loss for Language Image Pre-Training

cs.CV · 2023-03-27 · conditional · novelty 6.0

SigLIP replaces softmax-based contrastive loss with a simple pairwise sigmoid loss for vision-language pre-training, decoupling batch size from normalization and reaching strong zero-shot performance with limited compute.

citing papers explorer

Showing 2 of 2 citing papers.

  • Sigmoid Loss for Language Image Pre-Training cs.CV · 2023-03-27 · conditional · none · ref 41

    SigLIP replaces softmax-based contrastive loss with a simple pairwise sigmoid loss for vision-language pre-training, decoupling batch size from normalization and reaching strong zero-shot performance with limited compute.

  • ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP cs.CV · 2026-06-25 · unverdicted · none · ref 77

    ReasonCLIP-58M applies continual pretraining with visually grounded reasoning captions on 58M examples to improve CLIP-style models on commonsense and compositional reasoning tasks.