The paper formulates JEPA pretraining as conditional spectral graph learning equivalent to low-rank factorization of an action-conditioned co-occurrence matrix and derives a finite-sample generalization bound connecting pretraining error to downstream planning regret.
When Does LeJEPA Learn a World Model?
5 Pith papers cite this work. Polarity classification is still indexing.
abstract
A representation that scrambles the true degrees of freedom of the world cannot support reliable planning or compositional generalization. We prove that LeJEPA (alignment plus Gaussian regularization) linearly recovers the world's latent variables from nonlinear observations, a property known as linear identifiability, in a broad class of worlds where latents evolve under stationary, additive-noise transitions. Our main result is that among all such worlds, the Gaussian is the unique latent distribution for which this guarantee holds. The forward direction rests on a spectral decomposition in which each degree of nonlinearity is strictly penalized by alignment, making the linear map the optimum; the converse rules out every non-Gaussian alternative. We further prove an approximate identifiability result where the guarantee degrades gracefully, and show that linear, orthogonal identifiability enables optimal latent-space planning. We validate the theory with experiments ranging from 2D examples to 1024-dimensional latents, including distributional ablations and pixel-based robotic control. Our theory turns an empirically successful recipe into a mathematical guarantee, providing the foundation for building World Models that provably recover the structure of the world.
years
2026 5verdicts
UNVERDICTED 5representative citing papers
Physics-Grounded Symbolic Architectures achieve exact linear identifiability and near-infinite temporal consistency for any latent distribution, while statistical world models cannot for non-Gaussian dynamics.
Exact equivariance preserved through training renders one-step relMSE invariant across the symmetry group, enabling zero-shot generalization from a restricted training slice.
ILL rules on PMFs are marginal laws on deterministic quotient variables; the resulting constraint sets define log-linear factor graphs whose factors are indexed by learned abstractions, positioning ILL as interpretable PGM structure learning.
Proposes DCGWM architecture that partitions latent space into physical and behavioral subspaces with isolated gradient flows to structurally prevent objective interference collapse in grounded JEPA world models.
citing papers explorer
-
A Generalization Theory for JEPA-Based World Models
The paper formulates JEPA pretraining as conditional spectral graph learning equivalent to low-rank factorization of an action-conditioned co-occurrence matrix and derives a finite-sample generalization bound connecting pretraining error to downstream planning regret.
-
Identifiability Without Gaussianity: Symbolic World Models and Near-Infinite Temporal Consistency
Physics-Grounded Symbolic Architectures achieve exact linear identifiability and near-infinite temporal consistency for any latent distribution, while statistical world models cannot for non-Gaussian dynamics.
-
Exact equivariance, kept through training, buys zero-shot generalisation across the symmetry group
Exact equivariance preserved through training renders one-step relMSE invariant across the symmetry group, enabling zero-shot generalization from a restricted training slice.
-
Information Lattice Learning as Probabilistic Graphical Model Structure Learning
ILL rules on PMFs are marginal laws on deterministic quotient variables; the resulting constraint sets define log-linear factor graphs whose factors are indexed by learned abstractions, positioning ILL as interpretable PGM structure learning.
-
Dual-Channel Grounded World Modeling (DCGWM): Structural Prevention of Objective Interference Collapse via Heterogeneous External Grounding with Inward-Only Gradient Flow
Proposes DCGWM architecture that partitions latent space into physical and behavioral subspaces with isolated gradient flows to structurally prevent objective interference collapse in grounded JEPA world models.