UniverSat is a ViT-style model with a universal patch encoder enabling self-supervised training on heterogeneous multimodal Earth observation data from varying resolutions and sensors.
Cluster and predict latents patches for improved masked image modeling.arXiv preprint arXiv:2502.08769
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2verdicts
UNVERDICTED 2representative citing papers
TC-JEPA conditions masked feature prediction on text captions via sparse cross-attention to produce more semantically rich visual representations and outperforms contrastive methods on fine-grained tasks.
citing papers explorer
-
UniverSat: Resolution- and Modality-Agnostic Transformers for Earth Observation
UniverSat is a ViT-style model with a universal patch encoder enabling self-supervised training on heterogeneous multimodal Earth observation data from varying resolutions and sensors.
-
Text-Conditional JEPA for Learning Semantically Rich Visual Representations
TC-JEPA conditions masked feature prediction on text captions via sparse cross-attention to produce more semantically rich visual representations and outperforms contrastive methods on fine-grained tasks.