Frozen video diffusion models, probed at optimal depth and noise levels, produce representations competitive with discriminative encoders across semantic and geometric video tasks in a single forward pass.
Scaling 4d representations
6 Pith papers cite this work. Polarity classification is still indexing.
abstract
Scaling has not yet been convincingly demonstrated for pure self-supervised learning from video. However, prior work has focused evaluations on semantic-related tasks $\unicode{x2013}$ action classification, ImageNet classification, etc. In this paper we focus on evaluating self-supervised learning on non-semantic vision tasks that are more spatial (3D) and temporal (+1D = 4D), such as camera pose estimation, point and object tracking, and depth estimation. We show that by learning from very large video datasets, masked auto-encoding (MAE) with transformer video models actually scales, consistently improving performance on these 4D tasks, as model size increases from 20M all the way to the largest by far reported self-supervised video model $\unicode{x2013}$ 22B parameters. Rigorous apples-to-apples comparison with many recent image and video models demonstrates the benefits of scaling 4D representations. Pretrained models are available at https://github.com/google-deepmind/representations4d .
representative citing papers
LA-Pose achieves over 10% higher pose accuracy than recent feed-forward methods on Waymo and PandaSet benchmarks by repurposing latent actions from self-supervised inverse-dynamics pretraining while using orders of magnitude less labeled 3D data.
A new evaluation framework using latent diffusion on frozen vision backbones shows video-pretrained models consistently outperform image-based ones in forecasting entire trajectories across abstraction levels.
V-JEPA 2 pre-trained on massive unlabeled video achieves strong results on motion understanding and action anticipation, SOTA video QA at 8B scale, and enables zero-shot robotic planning on Franka arms using only 62 hours of unlabeled robot video.
Systematic empirical comparison of temporal context placement across backbone, PEFT modules, and probes for low-resource video task adaptation on appearance, motion, and dense tasks.
Freezing an image foundation model and training only a recurrent temporal module yields strong temporal performance on video tasks without large-scale video pre-training.
citing papers explorer
-
Gen4U: Unifying Video Generation and Understanding via Diffusion
Frozen video diffusion models, probed at optimal depth and noise levels, produce representations competitive with discriminative encoders across semantic and geometric video tasks in a single forward pass.
-
LA-Pose: Latent Action Pretraining Meets Pose Estimation
LA-Pose achieves over 10% higher pose accuracy than recent feed-forward methods on Waymo and PandaSet benchmarks by repurposing latent actions from self-supervised inverse-dynamics pretraining while using orders of magnitude less labeled 3D data.
-
Frozen Forecasting: A Unified Evaluation
A new evaluation framework using latent diffusion on frozen vision backbones shows video-pretrained models consistently outperform image-based ones in forecasting entire trajectories across abstraction levels.
-
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
V-JEPA 2 pre-trained on massive unlabeled video achieves strong results on motion understanding and action anticipation, SOTA video QA at 8B scale, and enables zero-shot robotic planning on Franka arms using only 62 hours of unlabeled robot video.
-
Where Do We (Not) Need Temporal Context in Low-Resource Video Task Adaptation?
Systematic empirical comparison of temporal context placement across backbone, PEFT modules, and probes for low-resource video task adaptation on appearance, motion, and dense tasks.
-
Towards Data-Efficient Video Pre-training with Frozen Image Foundation Models
Freezing an image foundation model and training only a recurrent temporal module yields strong temporal performance on video tasks without large-scale video pre-training.