The pre-normalized CLIP space consists of two linearly separable, offset ellipsoid shells, and cosine similarity to the modality mean closely estimates how typical an image or caption is.
Crosspoint: Self-supervised cross-modal contrastive learning for 3d point cloud understanding
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
The Double-Ellipsoid Geometry of CLIP
The pre-normalized CLIP space consists of two linearly separable, offset ellipsoid shells, and cosine similarity to the modality mean closely estimates how typical an image or caption is.