LeJEPA achieves linear identifiability of latent variables uniquely when the latents are Gaussian in worlds with stationary additive-noise transitions.
hub Canonical reference
VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning
Canonical reference. 86% of citing Pith papers cite this work as background.
abstract
Recent self-supervised methods for image representation learning are based on maximizing the agreement between embedding vectors from different views of the same image. A trivial solution is obtained when the encoder outputs constant vectors. This collapse problem is often avoided through implicit biases in the learning architecture, that often lack a clear justification or interpretation. In this paper, we introduce VICReg (Variance-Invariance-Covariance Regularization), a method that explicitly avoids the collapse problem with a simple regularization term on the variance of the embeddings along each dimension individually. VICReg combines the variance term with a decorrelation mechanism based on redundancy reduction and covariance regularization, and achieves results on par with the state of the art on several downstream tasks. In addition, we show that incorporating our new variance term into other methods helps stabilize the training and leads to performance improvements.
hub tools
citation-role summary
citation-polarity summary
representative citing papers
A JEPA-based hypernetwork maps lattice field theory couplings to flow-model weights, and the geometry of those weights recovers the phase transition, intrinsic dimension, and Ising critical exponent of 2D scalar field theory without supervised physics labels.
A hybrid evolution-strategy and gradient-descent framework maximizes a non-differentiable 'surprise score' to discover non-random features for non-parametric self-supervised image clustering.
UR-JEPA applies uniform rectifiability regularization via a smoothed Carleson square function to JEPA training, producing embeddings with 4-5 order PCA spectral drop at dimension 20-25 and lower seed variance than Gaussian regularization on Inet10, Galaxy10, and EuroSAT.
Spectral Guidance learns singular functions via self-supervised objective to project guidance signals onto diffusion sampling trajectories, enabling stable control without retraining or backpropagation and improving CIFAR-10 accuracy by 37 points with 4x faster sampling.
NTM models each generative reverse step as a conditional normalizing flow with a hybrid shallow-deep architecture, enabling exact-likelihood training and strong four-step sampling performance on text-to-image tasks.
PairAlign learns compact variable-length token sequences for audio via self-alignment on paired content-preserving views, achieving 55% fewer archive tokens than VQ while preserving edit-distance retrieval at 12.71 tokens/s.
Action-conditioned JEPA models treat pathology as a transition vector on latent states to simulate cardiac dynamics, outperforming supervised learning by over 0.05 AUROC in low-resource regimes on MIMIC-IV-ECG.
CoReDi coevolves semantic representations with the diffusion model via a jointly learned linear projection stabilized by stop-gradient, normalization, and regularization, yielding faster convergence and higher sample quality than fixed-representation baselines.
MOSAIC learns overlap-aware shared-specific representations, fits a first-stage predictor on overlapping data, and calibrates the gap using target-pattern samples, with non-asymptotic error bounds decomposing overlap size, calibration gap, and representation error.
Equivariant Poincaré ResNets combine hyperbolic geometry with C4 and D4 group symmetries via specialized reshaping, permutations, and batch norm to reduce optimization space and speed convergence while staying inside the Poincaré ball.
RT-SFOD adapts dual-head detectors like YOLOv10 for source-free object detection via DHF pseudo-label fusion and MARD loss, delivering 1.4-3.5% mAP gains with 1.3x higher throughput and ~2x fewer parameters than prior SFOD methods.
Delta-JEPA augments latent forward prediction with a Latent Difference Action Decoder that reconstructs actions from embedding displacements, yielding action-sensitive world models that improve planning on four visual continuous-control tasks over JEPA baselines.
ScaleAware-JEPA combines Constrained Diffusion Decomposition with a scale-tied JEPA objective to learn label-free latent coordinates that recover coherent morphology in multiscale fields such as MHD turbulence and interstellar gas.
RBFN projection heads serve as competitive replacements for MLP heads in SSL and enable SNS, a label-free metric from RBF parameters that correlates strongly with logistic regression evaluation.
DES Y3 weak lensing analysis with hybrid map-level statistics and simulation-based inference yields S8 = 0.808 ± 0.017, Ωm = 0.325 ± 0.024, and w < -0.766, improving the figure of merit by 60% over prior state-of-the-art.
Feedback alignment in deep networks is limited by low-rank error signals; orthogonal weight updates and activity normalization raise effective rank and boost performance.
DALE-CT, a 2D LeJEPA model with depth-aware dual supervision, reaches 0.833 Macro AUROC on multi-abnormality detection in CT and approaches 3D SOTA performance using less data and no textual supervision.
A block-rotation predictor inside a JEPA model imposes the circular geometry of modular arithmetic on latent representations of MNIST digits and yields strong zero-shot generalization to unseen operations.
GEARS is a geometry-first generative framework that learns domain-invariant encoders and permutation-equivariant diffusion generators to reconstruct intrinsic 2D cell coordinates and distance matrices from unpaired scRNA-seq guided by ST.
Introduces LOES, a constructive spectral method to select task-discriminative subspaces from intermediate layer embeddings, and GeoReg for enforcing simplicial class geometry during fine-tuning, with reported gains increasing with model depth across modalities.
Layerwise self-supervised local rules learn the hierarchical structure of the Random Hierarchy Model as data-efficiently as supervised backpropagation, while direct feedback approximations fail due to missing masking nonlinearities.
RC-aux corrects spatiotemporal mismatch in reconstruction-free latent world models by adding multi-horizon prediction and reachability supervision, improving planning performance on goal-conditioned pixel-control tasks.
Self-supervised learning can be understood as latent distribution matching, and under a Gaussian predictive model this yields identifiable representations up to affine transformations.
citing papers explorer
-
When Does LeJEPA Learn a World Model?
LeJEPA achieves linear identifiability of latent variables uniquely when the latents are Gaussian in worlds with stationary additive-noise transitions.
-
Weight-Space Physics: Interpretable Hypernetworks for Lattice Quantum Field Theories
A JEPA-based hypernetwork maps lattice field theory couplings to flow-model weights, and the geometry of those weights recovers the phase transition, intrinsic dimension, and Ising critical exponent of 2D scalar field theory without supervised physics labels.
-
Converge to Surprise: Evolutionary Self-supervised Image Clustering
A hybrid evolution-strategy and gradient-descent framework maximizes a non-differentiable 'surprise score' to discover non-random features for non-parametric self-supervised image clustering.
-
UR-JEPA: Uniform Rectifiability as a Regularizer for Joint-Embedding Predictive Architectures
UR-JEPA applies uniform rectifiability regularization via a smoothed Carleson square function to JEPA training, producing embeddings with 4-5 order PCA spectral drop at dimension 20-25 and lower seed variance than Gaussian regularization on Inet10, Galaxy10, and EuroSAT.
-
Spectral Guidance for Flexible and Efficient Control of Diffusion Models
Spectral Guidance learns singular functions via self-supervised objective to project guidance signals onto diffusion sampling trajectories, enabling stable control without retraining or backpropagation and improving CIFAR-10 accuracy by 37 points with 4x faster sampling.
-
Normalizing Trajectory Models
NTM models each generative reverse step as a conditional normalizing flow with a hybrid shallow-deep architecture, enabling exact-likelihood training and strong four-step sampling performance on text-to-image tasks.
-
PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization
PairAlign learns compact variable-length token sequences for audio via self-alignment on paired content-preserving views, achieving 55% fewer archive tokens than VQ while preserving edit-distance retrieval at 12.71 tokens/s.
-
Beyond Patient Invariance: Learning Cardiac Dynamics via Action-Conditioned JEPAs
Action-conditioned JEPA models treat pathology as a transition vector on latent states to simulate cardiac dynamics, outperforming supervised learning by over 0.05 AUROC in low-resource regimes on MIMIC-IV-ECG.
-
Coevolving Representations in Joint Image-Feature Diffusion
CoReDi coevolves semantic representations with the diffusion model via a jointly learned linear projection stabilized by stop-gradient, normalization, and regularization, yielding faster convergence and higher sample quality than fixed-representation baselines.
-
Pattern-Calibrated Multimodal Prediction under Blockwise Missingness
MOSAIC learns overlap-aware shared-specific representations, fits a first-stage predictor on overlapping data, and calibrates the gap using target-pattern samples, with non-asymptotic error bounds decomposing overlap size, calibration gap, and representation error.
-
Group-Equivariant Poincar\'e Convolutional Networks
Equivariant Poincaré ResNets combine hyperbolic geometry with C4 and D4 group symmetries via specialized reshaping, permutations, and batch norm to reduce optimization space and speed convergence while staying inside the Poincaré ball.
-
Real-Time Source-Free Object Detection
RT-SFOD adapts dual-head detectors like YOLOv10 for source-free object detection via DHF pseudo-label fusion and MARD loss, delivering 1.4-3.5% mAP gains with 1.3x higher throughput and ~2x fewer parameters than prior SFOD methods.
-
Delta-JEPA: Learning Action-Sensitive World Models via Latent Difference Decoding
Delta-JEPA augments latent forward prediction with a Latent Difference Action Decoder that reconstructs actions from embedding displacements, yielding action-sensitive world models that improve planning on four visual continuous-control tasks over JEPA baselines.
-
ScaleAware-JEPA: Latent Representation for Discovery in Multiscale Physical Fields
ScaleAware-JEPA combines Constrained Diffusion Decomposition with a scale-tied JEPA objective to learn label-free latent coordinates that recover coherent morphology in multiscale fields such as MHD turbulence and interstellar gas.
-
Radial Basis Function Networks as Projection Heads in Self-Supervised Learning
RBFN projection heads serve as competitive replacements for MLP heads in SSL and enable SNS, a label-free metric from RBF parameters that correlates strongly with logistic regression evaluation.
-
Dark Energy Survey Year 3 results: optimized $w$CDM simulation-based inference with weak lensing map-level hybrid statistics
DES Y3 weak lensing analysis with hybrid map-level statistics and simulation-based inference yields S8 = 0.808 ± 0.017, Ωm = 0.325 ± 0.024, and w < -0.766, improving the figure of merit by 60% over prior state-of-the-art.
-
Overcoming Rank Collapse in Feedback Alignment
Feedback alignment in deep networks is limited by low-rank error signals; orthogonal weight updates and activity normalization raise effective rank and boost performance.
-
DALE-CT: Depth-Aware Foundation Models for Computed Tomography
DALE-CT, a 2D LeJEPA model with depth-aware dual supervision, reaches 0.833 Macro AUROC on multi-abnormality detection in CT and approaches 3D SOTA performance using less data and no textual supervision.
-
BRo-JEPA: Learning Modular Arithmetic in Latent Space
A block-rotation predictor inside a JEPA model imposes the circular geometry of modular arithmetic on latent representations of MNIST digits and yields strong zero-shot generalization to unseen operations.
-
Geometry-First Generative Spatial Single-Cell Reconstruction
GEARS is a geometry-first generative framework that learns domain-invariant encoders and permutation-equivariant diffusion generators to reconstruct intrinsic 2D cell coordinates and distance matrices from unpaired scRNA-seq guided by ST.
-
Uncovering the Latent Potential of Deep Intermediate Representations
Introduces LOES, a constructive spectral method to select task-discriminative subspaces from intermediate layer embeddings, and GeoReg for enforcing simplicial class geometry during fine-tuning, with reported gains increasing with model depth across modalities.
-
Self-supervised local learning rules learn the hidden hierarchical structure of high-dimensional data
Layerwise self-supervised local rules learn the hierarchical structure of the Random Hierarchy Model as data-efficiently as supervised backpropagation, while direct feedback approximations fail due to missing masking nonlinearities.
-
Predictive but Not Plannable: RC-aux for Latent World Models
RC-aux corrects spatiotemporal mismatch in reconstruction-free latent world models by adding multi-horizon prediction and reachability supervision, improving planning performance on goal-conditioned pixel-control tasks.
-
Understanding Self-Supervised Learning via Latent Distribution Matching
Self-supervised learning can be understood as latent distribution matching, and under a Gaussian predictive model this yields identifiable representations up to affine transformations.
-
Understanding DNNs in Feature Interaction Models: A Dimensional Collapse Perspective
DNNs mitigate dimensional collapse of embeddings in feature interaction models, shown via parallel and stacked experiments plus gradient analysis.
-
Monitoring Neural Training with Topology: A Footprint-Predictable Collapse Index
A composite Collapse Index based on incremental discrete Morse homology provides low-latency early warning of representational collapse during neural network training.
-
Self-Supervised Representation Learning via Hyperspherical Density Shaping
HyDeS introduces hyperspherical density shaping with a von Mises-Fisher estimator to create theoretically grounded self-supervised representations that focus on foreground features.
-
RankUp: Towards High-rank Representations for Large Scale Advertising Recommender Systems
RankUp raises effective rank of representations in deep MetaFormer recommenders via randomized splitting and multi-embeddings, delivering 2-5% GMV gains in production deployments at Weixin.
-
HSG: Hyperbolic Scene Graph
Hyperbolic Scene Graph (HSG) learns embeddings in hyperbolic space for better hierarchical structure in scene graphs, achieving graph IoU of 33.51 versus 25.37 for the best Euclidean baseline.
-
Rapidly deploying on-device eye tracking by distilling visual foundation models
DistillGaze reduces median gaze error by 58.62% on a 2000+ participant dataset by distilling foundation models into a 256K-parameter on-device model using synthetic labeled data and unlabeled real data.
-
Dreamer-CDP: Improving Reconstruction-free World Models Via Continuous Deterministic Representation Prediction
Dreamer-CDP achieves reconstruction-free world modeling via a JEPA-style predictor on continuous deterministic representations and matches Dreamer's performance on Crafter.
-
LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
LeJEPA derives an optimal isotropic Gaussian target for embeddings and enforces it via sketched regularization to deliver scalable, heuristics-free self-supervised pretraining with 79% ImageNet linear accuracy on ViT-H/14.
-
Learning General Representation of 12-Lead Electrocardiogram with a Joint-Embedding Predictive Architecture
ECG-JEPA applies a joint-embedding predictive architecture with Cross-Pattern Attention to learn semantic representations from unlabeled 12-lead ECG data and reports state-of-the-art results on diagnostic classification, feature extraction, and segmentation.
-
Revisiting Feature Prediction for Learning Visual Representations from Video
V-JEPA models trained only on feature prediction from 2 million public videos achieve 81.9% on Kinetics-400, 72.2% on Something-Something-v2, and 77.9% on ImageNet-1K using frozen ViT-H/16 backbones.
-
Vision Transformers Need Registers
Adding register tokens to Vision Transformers eliminates high-norm background artifacts and raises state-of-the-art performance on dense visual prediction tasks.
-
STST-JEPA: Shallow-Target Spatio-Temporal Joint Embedding Prediction Architecture For EEG Self-Supervised Learning
A JEPA-style EEG foundation model with shallow EMA targets plus light reconstruction reaches strong multi-task transfer and 3.06-year validation age MAE on a large multi-site corpus.
-
AGE: Adaptive-masking for Graph Embedding in Graph Retrieval-Augmented Generation
AGE applies adaptive masking via a learnable sampler in Transformer-based SSL to align graph and text embeddings, yielding higher accuracy on four GraphQA benchmarks for non-parametric GraphRAG.
-
SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning
SingGuard introduces a policy-adaptive multimodal LLM guardrail with dynamic reasoning regimes and SingGuard-Bench, reporting SOTA F1 scores across 35 datasets and improved policy-following accuracy under runtime shifts.
-
FetSelect: Task-Specific Architectures and Self-Supervised Learning for Automated Fetal Ultrasound Frame Selection
FetSelect pairs a frozen vision foundation model with a hybrid multi-head design and BYOL pretraining on ultrasound data to select quality fetal frames, reporting mean AUROC 0.956 on expert-labeled test data.
-
AEF-Econ: Toward Plug-and-Play Socioeconomic Foundation Embeddings from AlphaEarth for Urban Remote Sensing
AEF-Econ introduces the CHN-Econ benchmark and Capacity-Adaptive Reconstruction to produce socioeconomic foundation embeddings that raise cross-region R² from 0.301 to 0.848 and cross-tier R² from 0.160 to 0.693 over baseline AEF.
-
Building The Ph(ysical)AI Layer Of Machine Intelligence
A principle-driven RF encoder achieves 77.7% average accuracy across 15 cross-modal tasks, performing better on physically grounded tasks than semantic ones.
-
HQ-JEPA: Hybrid Quantum Joint-Embedding Predictive Architecture for Cross-Modal Remote Sensing Representation Learning
HQ-JEPA combines JEPA-style predictive self-supervision with cross-modal alignment and a SWAP-test-based quantum fidelity loss for learning representations from paired remote sensing imagery, reporting competitive results on GeoBench tasks.
-
Back to Parsimonious Latents: Learning Task-Centric World Models from Visual Foundations
TC-WM converts foundation-model visual embeddings into parsimonious task-sufficient world model latents via linear projection, contrastive physical-state alignment, and embedding reconstruction, with a theoretical identification guarantee.
-
Mind Dreamer: Untethering Imagination via Active Causal Intervention on Latent Manifolds
Mind Dreamer uses active causal intervention via an adversarial initial-state generator and relay value functions to untether imagination in MBRL, claiming 1.67x average and up to 8.8x sparse-reward speedups over DreamerV3.
-
Representation Without Reward: A JEPA Audit for LLM Fine-Tuning
An empirical audit of 22 JEPA-style training auxiliaries on Llama-3.2-1B fine-tuning for regex generation finds no statistically significant task improvement after multiple-testing correction, even when auxiliaries visibly alter hidden-state geometry.
-
MER-DG: Modality-Entropy Regularization for Multimodal Domain Generalization
MER-DG applies modality-entropy regularization to reduce fusion overfitting in multimodal domain generalization, reporting average gains of 5% over standard fusion and 2% over prior methods on EPIC-Kitchens and HAC benchmarks.
-
Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference Perspective
CmIR uses causal inference to separate invariant causal representations from spurious ones in multimodal data, improving generalization under distribution shifts and noise via invariance, mutual information, and reconstruction constraints.
-
Hierarchical Planning with Latent World Models
Hierarchical latent world models with macro-actions solve long-horizon visual planning (70% Franka pick-and-place vs 0% flat planning) with up to 3× less compute.
-
Diverse via bounded Agreement: Geometric Regularization for Multimodal Fusion
Adding a dispersion loss plus a bounded cross-modal drift penalty to intermediate embeddings improves unimodal and multimodal accuracy across audio-visual, image-text, and RF benchmarks.
-
ZooClaw-FashionSigLIP2: Distilled Fine-tuning for Robust Fashion Retrieval
ZooClaw-FashionSigLIP2 applies distilled full fine-tuning plus WiseFT interpolation to SigLIP2-base and reports outperforming LoRA, larger backbones, and external data on fashion retrieval benchmarks while releasing a new benchmark and bias analysis.