Pith. sign in

REVIEW 9 cited by

Scaling 4D Representations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.15212 v2 pith:CE3YOYL2 submitted 2024-12-19 cs.CV cs.AIcs.LG

Scaling 4D Representations

classification cs.CV cs.AIcs.LG
keywords videolearningmodelsscalingself-supervisedtasksclassificationestimation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Scaling has not yet been convincingly demonstrated for pure self-supervised learning from video. However, prior work has focused evaluations on semantic-related tasks $\unicode{x2013}$ action classification, ImageNet classification, etc. In this paper we focus on evaluating self-supervised learning on non-semantic vision tasks that are more spatial (3D) and temporal (+1D = 4D), such as camera pose estimation, point and object tracking, and depth estimation. We show that by learning from very large video datasets, masked auto-encoding (MAE) with transformer video models actually scales, consistently improving performance on these 4D tasks, as model size increases from 20M all the way to the largest by far reported self-supervised video model $\unicode{x2013}$ 22B parameters. Rigorous apples-to-apples comparison with many recent image and video models demonstrates the benefits of scaling 4D representations. Pretrained models are available at https://github.com/google-deepmind/representations4d .

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Unique Lives, Shared World: Learning from Single-Life Videos

    cs.CV 2025-12 conditional novelty 7.0

    Vision models trained independently on single egocentric lives converge to aligned geometric representations, and about 30 hours of one life matches 30 hours of diverse video for depth-estimation pretraining.

  2. Self-Supervised Learning of Structured Dynamics from Videos

    cs.CV 2026-07 conditional novelty 6.0

    A two-token 'primary/residual' future-feature predictor separates camera from object motion, outpacing frozen-feature baselines and matching larger supervised models on several probes.

  3. SeeSE3: Emergence of 3D Space in Vision Features

    cs.CV 2026-07 conditional novelty 6.0

    Self-supervised vision features, especially DINOv2, contain a subspace that a small trained adapter can map to 3D camera motion, enabling pose estimation and latent-space navigation without explicit 3D reconstruction.

  4. Gen4U: Unifying Video Generation and Understanding via Diffusion

    cs.CV 2026-07 conditional novelty 6.0

    Frozen video diffusion models, probed at optimal depth and noise levels, produce representations competitive with discriminative encoders across semantic and geometric video tasks in a single forward pass.

  5. LA-Pose: Latent Action Pretraining Meets Pose Estimation

    cs.CV 2026-04 unverdicted novelty 6.0

    LA-Pose achieves over 10% higher pose accuracy than recent feed-forward methods on Waymo and PandaSet benchmarks by repurposing latent actions from self-supervised inverse-dynamics pretraining while using orders of ma...

  6. Frozen Forecasting: A Unified Evaluation

    cs.CV 2025-07 unverdicted novelty 6.0

    A new evaluation framework using latent diffusion on frozen vision backbones shows video-pretrained models consistently outperform image-based ones in forecasting entire trajectories across abstraction levels.

  7. V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

    cs.AI 2025-06 unverdicted novelty 6.0

    V-JEPA 2 pre-trained on massive unlabeled video achieves strong results on motion understanding and action anticipation, SOTA video QA at 8B scale, and enables zero-shot robotic planning on Franka arms using only 62 h...

  8. Where Do We (Not) Need Temporal Context in Low-Resource Video Task Adaptation?

    cs.CV 2026-06 unverdicted novelty 5.0

    Systematic empirical comparison of temporal context placement across backbone, PEFT modules, and probes for low-resource video task adaptation on appearance, motion, and dense tasks.

  9. Towards Data-Efficient Video Pre-training with Frozen Image Foundation Models

    cs.CV 2026-05 unverdicted novelty 5.0

    Freezing an image foundation model and training only a recurrent temporal module yields strong temporal performance on video tasks without large-scale video pre-training.