Pith. sign in

hub Mixed citations

V-JEPA 2.1: Unlocking dense features in video self-supervised learning

Mixed citation behavior. Most common role is background (57%).

37 Pith papers citing it
Background 57% of classified citations
abstract

We present V-JEPA 2.1, a family of self-supervised models that learn dense, high-quality visual representations for both images and videos while retaining strong global scene understanding. The approach combines four key components. First, a dense predictive loss uses a masking-based objective in which both visible and masked tokens contribute to the training signal, encouraging explicit spatial and temporal grounding. Second, deep self-supervision applies the self-supervised objective hierarchically across multiple intermediate encoder layers to improve representation quality. Third, multi-modal tokenizers enable unified training across images and videos. Finally, the model benefits from effective scaling in both model capacity and training data. Together, these design choices produce representations that are spatially structured, semantically coherent, and temporally consistent. Empirically, V-JEPA 2.1 achieves state-of-the-art performance on several challenging benchmarks, including 7.71 mAP on Ego4D for short-term object-interaction anticipation and 40.8 Recall@5 on EPIC-KITCHENS for high-level action anticipation, as well as a 20-point improvement in real-robot grasping success rate over V-JEPA-2 AC. The model also demonstrates strong performance in robotic navigation (5.687 ATE on TartanDrive), depth estimation (0.307 RMSE on NYUv2 with a linear probe), and global recognition (77.7 on Something-Something-V2). These results show that V-JEPA 2.1 significantly advances the state of the art in dense visual understanding and world modeling.

hub tools

citation-role summary

background 4 method 3

citation-polarity summary

years

2026 37

representative citing papers

Vision Pretraining for Dense Spatial Perception

cs.CV · 2026-07-06 · conditional · novelty 7.0

A boundary-forcing masked modeling paradigm for self-supervised vision pretraining yields a 1B model rivaling 7B models on dense spatial perception tasks.

GEAR: Guided End-to-End AutoRegression for Image Synthesis

cs.CV · 2026-06-30 · unverdicted · novelty 7.0

GEAR jointly trains VQ tokenizer and AR generator end-to-end via dual hard/soft read-out and representation alignment, achieving up to 10x faster ImageNet gFID convergence than LlamaGen-REPA while generalizing across quantizers and to text-to-image.

OctoSense: Self-Supervised Learning for Multimodal Robot Perception

cs.CV · 2026-06-25 · unverdicted · novelty 7.0

OctoSense supplies a large multimodal robotics dataset and a late-fusion masked autoencoder that runs fast and outperforms image-only models on optical flow, depth, segmentation, and ego-motion tasks while remaining robust under sensor degradation.

Gen4U: Unifying Video Generation and Understanding via Diffusion

cs.CV · 2026-07-07 · conditional · novelty 6.0

Frozen video diffusion models, probed at optimal depth and noise levels, produce representations competitive with discriminative encoders across semantic and geometric video tasks in a single forward pass.

Active Inference as the Test-Time Scaling Law for Physical AI Agents

cs.AI · 2026-06-22 · unverdicted · novelty 6.0

Active inference supplies a test-time scaling law for physical AI agents by framing policy updates as soft Bayesian inference that reduces expected prediction errors, with a variational solution shown to outperform Q-learning and Bayesian RL in driving simulations.

Latent Video Prediction Learns Better World Models

cs.CV · 2026-05-15 · unverdicted · novelty 6.0

Latent prediction video models exhibit a distinct robustness profile across corruption, occlusion, fine-grained discrimination, and temporal sensitivity compared to other self-supervised video models when used as world models.

EgoExo-WM: Unlocking Exo Video for Ego World Models

cs.CV · 2026-05-14 · unverdicted · novelty 6.0 · 2 refs

Method converts exocentric videos to egocentric format via body-pose extraction and kinematics to improve egocentric world-model prediction and planning.

Lifting Embodied World Models for Planning and Control

cs.CV · 2026-04-28 · conditional · novelty 6.0

Planning with a frozen egocentric world model is lifted from 48-dim joint actions to a few 2D goal waypoints via a trained policy, reducing CEM error reduction by 3.8x.

Einstein World Models

cs.AI · 2026-06-25 · unverdicted · novelty 5.0

Einstein World Models integrate visual rollouts from a callable world-module into LLM reasoning traces to support complex thought beyond language.

citing papers explorer

Showing 37 of 37 citing papers.