Contrastive training on randomly laid-out, alphabet-labeled synthetic diagrams makes CLIP encoders detect edge direction and caption diagrams, beating pretrained CLIP and zero-shot GPT-4o on the synthetic test set.
Neural codes for image retrieval
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Can Visual Encoder Learn to See Arrows?
Contrastive training on randomly laid-out, alphabet-labeled synthetic diagrams makes CLIP encoders detect edge direction and caption diagrams, beating pretrained CLIP and zero-shot GPT-4o on the synthetic test set.