LeJEPA achieves linear identifiability of latent variables uniquely when the latents are Gaussian in worlds with stationary additive-noise transitions.
Causal- jepa: Learning world models through object-level latent interventions
12 Pith papers cite this work. Polarity classification is still indexing.
abstract
World models require robust relational understanding to support prediction, reasoning, and control. While object-centric representations provide a useful abstraction, they are not sufficient to capture interaction-dependent dynamics. We therefore propose C-JEPA, a simple and flexible object-centric world model that extends masked joint embedding prediction from image patches to object-centric representations. By masking object-level latents and requiring each masked object state to be inferred from the surrounding context, C-JEPA imposes structured partial observability during training, creating counterfactual-like prediction queries that discourage shortcut solutions and make interaction-dependent prediction necessary under the learning objective. Empirically, C-JEPA leads to consistent gains in visual question answering, with an absolute improvement of about 20% in counterfactual reasoning over the same architecture without object-level masking. On agent control tasks, C-JEPA enables substantially more efficient planning by using only 1% of the total latent input features required by patch-based world models, while achieving comparable performance. Finally, we provide a formal analysis demonstrating that object-level masking induces useful inductive bias by controlling observability. Our code is available at https://github.com/galilai-group/cjepa.
citation-role summary
citation-polarity summary
years
2026 12roles
background 1polarities
background 1representative citing papers
Goal-conditioned world models transcribe instructions instead of perceiving spatial relations when the instruction names the scored quantity, and removing the goal from the dynamics fixes it.
VLWMs learn variable-length action-conditioned dynamics in latent space with curriculum training, yielding 13% average gains over prior latent world models on long-horizon tasks.
Cross-trajectory negative sampling in contrastive predictive objectives causes encoding of slow noise over dynamics; intra-trajectory sampling eliminates the shortcut and recovers dynamical variables even under strong noise.
World models succeed when their latent states are built to meet task-specific sufficiency constraints rather than preserving the maximum amount of information.
Neuro-JEPA, a JEPA-and-MoE transformer pretrained on 1.55M brain MRI scans, outperformed prior neuroimaging foundation models and was the only one to beat a simple CNN baseline on average.
Loop-OWM uses color-prototype slots, demonstration-conditioned task summaries, and looped transitions to model ARC rules as visual-symbolic state changes and outperforms baselines on ARC-1 and ARC-2.
Slot-MPC learns slot representations to build a differentiable object-centric dynamics model that supports efficient gradient-based MPC for robotic manipulation in novel situations.
LeWM is a ~15M-parameter JEPA world model that trains end-to-end from pixels with only next-embedding prediction plus a Gaussian latent regularizer, cutting loss hyperparameters to one.
Einstein World Models integrate visual rollouts from a callable world-module into LLM reasoning traces to support complex thought beyond language.
CausalVAE plug-in for world models preserves factual prediction and boosts counterfactual retrieval, with large gains on physics benchmarks and recovered physical interaction trends.
citing papers explorer
-
When Does LeJEPA Learn a World Model?
LeJEPA achieves linear identifiability of latent variables uniquely when the latents are Gaussian in worlds with stationary additive-noise transitions.
-
Grounding Spatial Relations in a Compact World Model: Instruction Leakage and a Goal-Free Dynamics Fix
Goal-conditioned world models transcribe instructions instead of perceiving spatial relations when the instruction names the scored quantity, and removing the goal from the dynamics fixes it.
-
Beyond the Next Step: Variable-Length Latent World Models for Long-Horizon Planning
VLWMs learn variable-length action-conditioned dynamics in latent space with curriculum training, yielding 13% average gains over prior latent world models on long-horizon tasks.
-
Contrast encodes inductive bias: separating slow noise from dynamics in predictive representation learning
Cross-trajectory negative sampling in contrastive predictive objectives causes encoding of slow noise over dynamics; intra-trajectory sampling eliminates the shortcut and recovers dynamical variables even under strong noise.
-
Latent State Design for World Models under Sufficiency Constraints
World models succeed when their latent states are built to meet task-specific sufficiency constraints rather than preserving the maximum amount of information.
-
Learning Sparse Latent Predictive Foundation Model for Multimodal Neuroimaging
Neuro-JEPA, a JEPA-and-MoE transformer pretrained on 1.55M brain MRI scans, outperformed prior neuroimaging foundation models and was the only one to beat a simple CNN baseline on average.
-
Slots, Transitions, Loops: Learning Composable World Models for ARC
Loop-OWM uses color-prototype slots, demonstration-conditioned task summaries, and looped transitions to model ARC rules as visual-symbolic state changes and outperforms baselines on ARC-1 and ARC-2.
-
Slot-MPC: Goal-Conditioned Model Predictive Control with Object-Centric Representations
Slot-MPC learns slot representations to build a differentiable object-centric dynamics model that supports efficient gradient-based MPC for robotic manipulation in novel situations.
-
LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels
LeWM is a ~15M-parameter JEPA world model that trains end-to-end from pixels with only next-embedding prediction plus a Gaussian latent regularizer, cutting loss hyperparameters to one.
-
Einstein World Models
Einstein World Models integrate visual rollouts from a callable world-module into LLM reasoning traces to support complex thought beyond language.
-
CausalVAE as a Plug-in for World Models: Towards Reliable Counterfactual Dynamics
CausalVAE plug-in for world models preserves factual prediction and boosts counterfactual retrieval, with large gains on physics benchmarks and recovered physical interaction trends.
- Trivium: Temporal Regret as a First-Class Objective for Causal-Memory Controllers