SignMAE uses segmentation-driven masking in a mask-and-reconstruct self-supervised task to learn fine-grained sign representations, achieving state-of-the-art accuracy on WLASL, NMFs-CSL, and Slovo with fewer frames and modalities.
Videomae v2: Scaling video masked autoencoders with dual masking
3 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
fields
cs.CV 3years
2026 3roles
background 1polarities
background 1representative citing papers
A motion-only embedding trained on synthetic point tracks matches or beats large appearance-based video models on temporal tasks and improves them when combined.
Video foundation models encode intuitive physics knowledge that is strongest in V-JEPA at intermediate-to-late layers and depends on pretraining type and probe design.
citing papers explorer
-
SignMAE: Segmentation-Driven Self-Supervised Learning for Sign Language Recognition
SignMAE uses segmentation-driven masking in a mask-and-reconstruct self-supervised task to learn fine-grained sign representations, achieving state-of-the-art accuracy on WLASL, NMFs-CSL, and Slovo with fewer frames and modalities.
-
The TIME Machine: On The Power of Motion for Efficient Perception
A motion-only embedding trained on synthetic point tracks matches or beats large appearance-based video models on temporal tasks and improves them when combined.
-
Do Video Foundation Models Understand Intuitive Physics? A Layerwise Probing Analysis
Video foundation models encode intuitive physics knowledge that is strongest in V-JEPA at intermediate-to-late layers and depends on pretraining type and probe design.