The abstract claims a new multi-rationale benchmark and a training-free contrastive framework for explainable object recognition, but the submitted full text is a different paper on M-valley twisted TMDs, so none of the claimed results are present to evaluate.
On the Difference of BERT-style and CLIP-style Text Encoders
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Masked language modeling (MLM) has been one of the most popular pretraining recipes in natural language processing, e.g., BERT, one of the representative models. Recently, contrastive language-image pretraining (CLIP) has also attracted attention, especially its vision models that achieve excellent performance on a broad range of vision tasks. However, few studies are dedicated to studying the text encoders learned by CLIP. In this paper, we analyze the difference between BERT-style and CLIP-style text encoders from three experiments: (i) general text understanding, (ii) vision-centric text understanding, and (iii) text-to-image generation. Experimental analyses show that although CLIP-style text encoders underperform BERT-style ones for general text understanding tasks, they are equipped with a unique ability, i.e., synesthesia, for the cross-modal association, which is more similar to the senses of humans.
fields
cs.CV 1years
2025 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
Multi-Rationale Explainable Object Recognition via Contrastive Conditional Inference
The abstract claims a new multi-rationale benchmark and a training-free contrastive framework for explainable object recognition, but the submitted full text is a different paper on M-valley twisted TMDs, so none of the claimed results are present to evaluate.