Pith. sign in

Revisiting feature prediction for learning visual rep- resentations from video

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.CV 1

years

2024 1

verdicts

CONDITIONAL 1

representative citing papers

Scaling 4D Representations

cs.CV · 2024-12-19 · conditional · novelty 7.0

Scaling masked autoencoding video transformers from 20M to 22B parameters steadily improved camera pose, tracking, and depth estimation, while language-supervised and image-only models lagged on these tasks.

citing papers explorer

Showing 1 of 1 citing paper.

  • Scaling 4D Representations cs.CV · 2024-12-19 · conditional · none · ref 8

    Scaling masked autoencoding video transformers from 20M to 22B parameters steadily improved camera pose, tracking, and depth estimation, while language-supervised and image-only models lagged on these tasks.