Pith. sign in

REVIEW 3 cited by

On the Difference of BERT-style and CLIP-style Text Encoders

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.03678 v1 pith:GVMRZE6O submitted 2023-06-06 cs.CL

classification cs.CL
keywords textencodersbert-styleclip-styleunderstandingclipdifferencegeneral
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Masked language modeling (MLM) has been one of the most popular pretraining recipes in natural language processing, e.g., BERT, one of the representative models. Recently, contrastive language-image pretraining (CLIP) has also attracted attention, especially its vision models that achieve excellent performance on a broad range of vision tasks. However, few studies are dedicated to studying the text encoders learned by CLIP. In this paper, we analyze the difference between BERT-style and CLIP-style text encoders from three experiments: (i) general text understanding, (ii) vision-centric text understanding, and (iii) text-to-image generation. Experimental analyses show that although CLIP-style text encoders underperform BERT-style ones for general text understanding tasks, they are equipped with a unique ability, i.e., synesthesia, for the cross-modal association, which is more similar to the senses of humans.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Virtual Dosimetrists: A Radiotherapy Training "Flight Simulator"

    physics.med-ph 2025-05 conditional novelty 6.0 of 10

    A dose-prediction network conditioned on CLIP text embeddings can generate and iteratively revise radiotherapy dose distributions from simple language prompts, creating a fast training simulator for plan quality review.

  2. Unified Framework for Open-World Compositional Zero-shot Learning

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A single-stream vision-language transformer with top-K text selection and a sparse compositor achieves state-of-the-art open-world compositional zero-shot learning on MIT-States, C-GQA, and VAW-CZSL.

  3. Multi-Rationale Explainable Object Recognition via Contrastive Conditional Inference

    cs.CV 2025-08 unverdicted novelty 4.0 of 10

    The abstract claims a new multi-rationale benchmark and a training-free contrastive framework for explainable object recognition, but the submitted full text is a different paper on M-valley twisted TMDs, so none of t...

Pith tools