Pith. sign in

PubMedClip: How much does CLIP benefit visual question answering in the medical domain? In Findings of the Association for Computational Linguistics: EACL 2023, pages 1181–1193,

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

citation-role summary

background 1

citation-polarity summary

fields

cs.CV 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

unclear 1

representative citing papers

GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures

cs.CV · 2025-07-24 · conditional · novelty 4.0

Adding LLM-style GEGLU, RMSNorm, and rotary position embeddings to CoCa's vision encoder reduced contrastive loss, perplexity, and CoCa loss on one pretraining and three fine-tuning datasets, compared with an internally trained control.

citing papers explorer

Showing 1 of 1 citing paper.

  • GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures cs.CV · 2025-07-24 · conditional · none · ref 12

    Adding LLM-style GEGLU, RMSNorm, and rotary position embeddings to CoCa's vision encoder reduced contrastive loss, perplexity, and CoCa loss on one pretraining and three fine-tuning datasets, compared with an internally trained control.