Pith. sign in

REVIEW 32 cited by

Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.08243 v3 pith:GD2CKQXV submitted 2023-01-19 cs.CV cs.AIcs.LGeess.IV

Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture

classification cs.CV cs.AIcs.LGeess.IV
keywords i-jepalearningrepresentationssemanticapproacharchitectureblockblocks
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

This paper demonstrates an approach for learning highly semantic image representations without relying on hand-crafted data-augmentations. We introduce the Image-based Joint-Embedding Predictive Architecture (I-JEPA), a non-generative approach for self-supervised learning from images. The idea behind I-JEPA is simple: from a single context block, predict the representations of various target blocks in the same image. A core design choice to guide I-JEPA towards producing semantic representations is the masking strategy; specifically, it is crucial to (a) sample target blocks with sufficiently large scale (semantic), and to (b) use a sufficiently informative (spatially distributed) context block. Empirically, when combined with Vision Transformers, we find I-JEPA to be highly scalable. For instance, we train a ViT-Huge/14 on ImageNet using 16 A100 GPUs in under 72 hours to achieve strong downstream performance across a wide range of tasks, from linear classification to object counting and depth prediction.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 32 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. What Can Latent World Models Know? Physical Parameter Identifiability in Multimodal Predictive Representations

    cs.LG 2026-07 conditional novelty 7.0

    Prediction targets, not fused inputs, decide which physical parameters enter a latent world model; slow ratio-type parameters like drag remain largely unacquired despite high recoverability certificates.

  2. UR-JEPA: Uniform Rectifiability as a Regularizer for Joint-Embedding Predictive Architectures

    cs.LG 2026-05 unverdicted novelty 7.0

    UR-JEPA applies uniform rectifiability regularization via a smoothed Carleson square function to JEPA training, producing embeddings with 4-5 order PCA spectral drop at dimension 20-25 and lower seed variance than Gau...

  3. CRONOS: Benchmarking Counterfactual Physical Consistency in Video Models

    cs.CV 2026-05 unverdicted novelty 7.0

    CRONOS benchmark shows recent open-source video generators fail to preserve physical consistency under controlled changes to viewpoint, scene, object category, and appearance.

  4. Seeking the Unfamiliar but Memorable: Conceptual Creativity as Meta-Learning

    cs.LG 2026-05 unverdicted novelty 7.0

    Creativity is defined as meta-learning where a frozen diffusion creator optimizes candidates for rapid improvement by an adapting appraiser such as an autoencoder or CLIP adapter.

  5. Normalizing Trajectory Models

    cs.CV 2026-05 unverdicted novelty 7.0

    NTM models each generative reverse step as a conditional normalizing flow with a hybrid shallow-deep architecture, enabling exact-likelihood training and strong four-step sampling performance on text-to-image tasks.

  6. Normalizing Trajectory Models

    cs.CV 2026-05 unverdicted novelty 7.0

    NTM uses per-step conditional normalizing flows plus a trajectory-wide predictor to achieve exact-likelihood 4-step sampling that matches or exceeds baselines on text-to-image tasks.

  7. ProteinJEPA: Latent prediction complements protein language models

    cs.LG 2026-05 unverdicted novelty 7.0

    Masked-position MLM plus JEPA latent prediction outperforms MLM-only pretraining on 10-11 of 16 downstream tasks for 35M-150M protein models while JEPA alone fails.

  8. Latent State Design for World Models under Sufficiency Constraints

    cs.AI 2026-05 unverdicted novelty 7.0

    World models succeed when their latent states are built to meet task-specific sufficiency constraints rather than preserving the maximum amount of information.

  9. ScaleAware-JEPA: Latent Representation for Discovery in Multiscale Physical Fields

    cs.LG 2026-06 unverdicted novelty 6.0

    ScaleAware-JEPA combines Constrained Diffusion Decomposition with a scale-tied JEPA objective to learn label-free latent coordinates that recover coherent morphology in multiscale fields such as MHD turbulence and int...

  10. Neural Voxel Dynamics: Learning Implicit 3D Physics via Volumetric Feature Advection

    cs.CV 2026-06 unverdicted novelty 6.0

    A self-supervised framework learns implicit 3D physics by lifting V-JEPA features into voxels and performing volumetric feature advection conditioned on actions.

  11. Learning from Semantic Dictionaries: Discriminative Codebook Contrastive Learning for Unified Visual Representation and Generation

    cs.CV 2026-05 unverdicted novelty 6.0

    LEASE achieves state-of-the-art unified performance on ImageNet-1K by combining masked token reconstruction and codebook contrast losses in a one-time precomputed discrete token space.

  12. SpectralEarth-FM: Bringing Hyperspectral Imagery into Multimodal Earth Observation Pretraining

    cs.CV 2026-05 unverdicted novelty 6.0

    SpectralEarth-FM is a multisensor hierarchical transformer pretrained on a 40TB co-located HSI-MSI-SAR dataset using a JEPA-style objective and reports state-of-the-art results on hyperspectral and standard EO benchmarks.

  13. Semantic Generative Tuning for Unified Multimodal Models

    cs.CV 2026-05 unverdicted novelty 6.0

    Semantic Generative Tuning uses image segmentation as a generative proxy to align misaligned representation spaces in unified multimodal models and improve both perception and generative layout fidelity.

  14. Entity-Centric World Models: Interaction-Aware Masking for Causal Video Prediction

    cs.CV 2026-05 unverdicted novelty 6.0

    IA-JEPA applies motion-centric masking in JEPA to focus on entity interactions, reporting 14.26% causal reasoning accuracy on CLEVRER versus 3.22% for standard baselines plus higher latent entropy and R²=0.43 energy l...

  15. Entity-Centric World Models: Interaction-Aware Masking for Causal Video Prediction

    cs.CV 2026-05 unverdicted novelty 6.0

    IA-JEPA applies interaction-aware masking to JEPA, raising causal reasoning accuracy on CLEVRER from 3.22% to 14.26% while producing a higher-entropy latent space that better aligns with physical energy.

  16. Physically Native World Models: A Hamiltonian Perspective on Generative World Modeling

    cs.AI 2026-05 unverdicted novelty 6.0

    Hamiltonian World Models structure latent dynamics around energy-conserving Hamiltonian evolution to produce physically grounded, action-controllable predictions for embodied decision making.

  17. Surprised by Attention: Predictable Query Dynamics for Time Series Anomaly Detection

    cs.LG 2026-03 conditional novelty 6.0

    Predicting multi-head attention queries from history and scoring cosine mismatch against an EMA target, combined with reconstruction error, improves unsupervised multivariate anomaly ranking and localization.

  18. ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model

    cs.CV 2026-03 conditional novelty 5.5

    Dual-temporal VLM guidance injected into a JEPA predictor via multi-layer pyramid features improves hand-manipulation trajectory forecasting over VLM-only and JEPA-only baselines.

  19. STST-JEPA: Shallow-Target Spatio-Temporal Joint Embedding Prediction Architecture For EEG Self-Supervised Learning

    cs.LG 2026-07 conditional novelty 5.0

    A JEPA-style EEG foundation model with shallow EMA targets plus light reconstruction reaches strong multi-task transfer and 3.06-year validation age MAE on a large multi-site corpus.

  20. LeNEPA: No-Augmentation Next-Latent Prediction for Time-Series Representation Learning

    cs.LG 2026-07 unverdicted novelty 5.0

    LeNEPA proposes a no-augmentation next-latent prediction recipe that maintains frozen-probe performance across ECG and synthetic diagnostic time-series datasets under fixed-recipe conditions where a tuned JEPA baselin...

  21. Do Video Foundation Models Understand Intuitive Physics? A Layerwise Probing Analysis

    cs.CV 2026-06 unverdicted novelty 5.0

    Video foundation models encode intuitive physics knowledge that is strongest in V-JEPA at intermediate-to-late layers and depends on pretraining type and probe design.

  22. Quantifying the Pre-training Dividend: Generative versus Latent Self-Supervised Learning for Time Series Foundation Models

    cs.LG 2026-05 unverdicted novelty 5.0

    Self-supervised pre-training delivers large gains up to 375% on time series anomaly detection and classification but only marginal benefits for forecasting, driven by a precision-invariance trade-off in the learned re...

  23. Semantic Generative Tuning for Unified Multimodal Models

    cs.CV 2026-05 unverdicted novelty 5.0

    Semantic Generative Tuning applies segmentation-based generative proxies during post-training to align and improve both understanding and generation in unified multimodal models.

  24. Representation Without Reward: A JEPA Audit for LLM Fine-Tuning

    cs.LG 2026-05 conditional novelty 5.0

    An empirical audit of 22 JEPA-style training auxiliaries on Llama-3.2-1B fine-tuning for regex generation finds no statistically significant task improvement after multiple-testing correction, even when auxiliaries vi...

  25. Physically Native World Models: A Hamiltonian Perspective on Generative World Modeling

    cs.AI 2026-05 unverdicted novelty 5.0

    Proposes Hamiltonian World Models as a physically grounded framework encoding observations into latent phase space and evolving them via Hamiltonian dynamics with control and dissipation for embodied prediction and planning.

  26. Physically Native World Models: A Hamiltonian Perspective on Generative World Modeling

    cs.AI 2026-05 unverdicted novelty 5.0

    The paper introduces Hamiltonian World Models by encoding observations into structured latent phase space and evolving states via Hamiltonian-inspired dynamics for physically meaningful rollouts in embodied AI.

  27. Weak-to-Strong Knowledge Distillation Accelerates Visual Learning

    cs.CV 2026-04 unverdicted novelty 5.0

    Weak-to-strong knowledge distillation applied early and then turned off accelerates convergence to target performance in visual learning tasks by factors of 1.7-4.8x.

  28. The Cartesian Cut in Agentic AI

    cs.AI 2026-04 unverdicted novelty 5.0

    LLM agents use a Cartesian split between learned prediction and engineered control, enabling modularity but creating sensitivity and bottlenecks unlike integrated biological systems.

  29. PANC: Prior-Aware Normalized Cut via Anchor-Augmented Token Graphs

    cs.CV 2026-02 unverdicted novelty 5.0

    PANC augments Normalized Cut with anchor-augmented token graphs using priors to steer spectral partitions, yielding mIoU gains of 2.3-8.7% over baselines on DUTS-TE, DUT-OMRON, and CrackForest.

  30. Bridging the Usability Gap: Lessons from Interpreting Studies for Machine Interpreting Design

    cs.CL 2026-06 unverdicted novelty 4.0

    Machine interpreting should shift from fidelity metrics to three design priorities—agency, grounding, and experience—drawn from interpreting studies to close the usability gap with human-mediated communication.

  31. World Action Models: A Survey

    cs.RO 2026-06 unverdicted novelty 3.0

    A survey that clarifies boundaries and organizes World Action Models by generation requirements and predictive substrates, identifying a trend toward generating less of the future.

  32. Research progress on quantum neural networks and quantum machine learning

    quant-ph 2026-05 unverdicted novelty 2.0

    Survey summarizing performance metrics of fully connected QNNs, quantum CNNs, equivariant QNNs, quantum Hopfield networks, quantum Boltzmann machines, quantum reservoir computing, and composite networks for reinforcem...