Adding LLM-style GEGLU, RMSNorm, and rotary position embeddings to CoCa's vision encoder reduced contrastive loss, perplexity, and CoCa loss on one pretraining and three fine-tuning datasets, compared with an internally trained control.
PubMedClip: How much does CLIP benefit visual question answering in the medical domain? In Findings of the Association for Computational Linguistics: EACL 2023, pages 1181–1193,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures
Adding LLM-style GEGLU, RMSNorm, and rotary position embeddings to CoCa's vision encoder reduced contrastive loss, perplexity, and CoCa loss on one pretraining and three fine-tuning datasets, compared with an internally trained control.