A single JEPA objective combining photometric and temporal prediction in one shared latent space matches or beats task-specific JEPAs on image, video, and planning benchmarks, with one loss hyperparameter.
Implementation Details We use a ViT-Small/16 encoder (15M parameters) for control and ViT-Large for large-scale image/video representation
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling
A single JEPA objective combining photometric and temporal prediction in one shared latent space matches or beats task-specific JEPAs on image, video, and planning benchmarks, with one loss hyperparameter.