Pith. sign in

Scaling 4d representations

6 Pith papers cite this work. Polarity classification is still indexing.

6 Pith papers citing it
abstract

Scaling has not yet been convincingly demonstrated for pure self-supervised learning from video. However, prior work has focused evaluations on semantic-related tasks $\unicode{x2013}$ action classification, ImageNet classification, etc. In this paper we focus on evaluating self-supervised learning on non-semantic vision tasks that are more spatial (3D) and temporal (+1D = 4D), such as camera pose estimation, point and object tracking, and depth estimation. We show that by learning from very large video datasets, masked auto-encoding (MAE) with transformer video models actually scales, consistently improving performance on these 4D tasks, as model size increases from 20M all the way to the largest by far reported self-supervised video model $\unicode{x2013}$ 22B parameters. Rigorous apples-to-apples comparison with many recent image and video models demonstrates the benefits of scaling 4D representations. Pretrained models are available at https://github.com/google-deepmind/representations4d .

fields

cs.CV 5 cs.AI 1

years

2026 4 2025 2

representative citing papers

Gen4U: Unifying Video Generation and Understanding via Diffusion

cs.CV · 2026-07-07 · conditional · novelty 6.0

Frozen video diffusion models, probed at optimal depth and noise levels, produce representations competitive with discriminative encoders across semantic and geometric video tasks in a single forward pass.

LA-Pose: Latent Action Pretraining Meets Pose Estimation

cs.CV · 2026-04-30 · unverdicted · novelty 6.0

LA-Pose achieves over 10% higher pose accuracy than recent feed-forward methods on Waymo and PandaSet benchmarks by repurposing latent actions from self-supervised inverse-dynamics pretraining while using orders of magnitude less labeled 3D data.

Frozen Forecasting: A Unified Evaluation

cs.CV · 2025-07-18 · unverdicted · novelty 6.0

A new evaluation framework using latent diffusion on frozen vision backbones shows video-pretrained models consistently outperform image-based ones in forecasting entire trajectories across abstraction levels.

citing papers explorer

Showing 6 of 6 citing papers.

  • Gen4U: Unifying Video Generation and Understanding via Diffusion cs.CV · 2026-07-07 · conditional · none · ref 4 · internal anchor

    Frozen video diffusion models, probed at optimal depth and noise levels, produce representations competitive with discriminative encoders across semantic and geometric video tasks in a single forward pass.

  • LA-Pose: Latent Action Pretraining Meets Pose Estimation cs.CV · 2026-04-30 · unverdicted · none · ref 8

    LA-Pose achieves over 10% higher pose accuracy than recent feed-forward methods on Waymo and PandaSet benchmarks by repurposing latent actions from self-supervised inverse-dynamics pretraining while using orders of magnitude less labeled 3D data.

  • Frozen Forecasting: A Unified Evaluation cs.CV · 2025-07-18 · unverdicted · none · ref 4

    A new evaluation framework using latent diffusion on frozen vision backbones shows video-pretrained models consistently outperform image-based ones in forecasting entire trajectories across abstraction levels.

  • V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning cs.AI · 2025-06-11 · unverdicted · none · ref 12

    V-JEPA 2 pre-trained on massive unlabeled video achieves strong results on motion understanding and action anticipation, SOTA video QA at 8B scale, and enables zero-shot robotic planning on Franka arms using only 62 hours of unlabeled robot video.

  • Where Do We (Not) Need Temporal Context in Low-Resource Video Task Adaptation? cs.CV · 2026-06-02 · unverdicted · none · ref 9

    Systematic empirical comparison of temporal context placement across backbone, PEFT modules, and probes for low-resource video task adaptation on appearance, motion, and dense tasks.

  • Towards Data-Efficient Video Pre-training with Frozen Image Foundation Models cs.CV · 2026-05-18 · unverdicted · none · ref 4

    Freezing an image foundation model and training only a recurrent temporal module yields strong temporal performance on video tasks without large-scale video pre-training.