Pith. sign in

REVIEW 2 cited by

Unmasked Teacher: Towards Training-Efficient Video Foundation Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.16058 v2 pith:FKHDMYK3 submitted 2023-03-28 cs.CV

classification cs.CV
keywords videomodelsunmaskeddatafoundationteachervfmsconvergence
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Video Foundation Models (VFMs) have received limited exploration due to high computational costs and data scarcity. Previous VFMs rely on Image Foundation Models (IFMs), which face challenges in transferring to the video domain. Although VideoMAE has trained a robust ViT from limited data, its low-level reconstruction poses convergence difficulties and conflicts with high-level cross-modal alignment. This paper proposes a training-efficient method for temporal-sensitive VFMs that integrates the benefits of existing methods. To increase data efficiency, we mask out most of the low-semantics video tokens, but selectively align the unmasked tokens with IFM, which serves as the UnMasked Teacher (UMT). By providing semantic guidance, our method enables faster convergence and multimodal friendliness. With a progressive pre-training framework, our model can handle various tasks including scene-related, temporal-related, and complex video-language understanding. Using only public sources for pre-training in 6 days on 32 A100 GPUs, our scratch-built ViT-L/16 achieves state-of-the-art performances on various video tasks. The code and models will be released at https://github.com/OpenGVLab/unmasked_teacher.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SMART-Vision: Survey of Modern Action Recognition Techniques in Vision

    cs.CV 2025-01 conditional novelty 4.0 of 10

    The SMART-Vision survey organizes vision-based human action recognition into a hybrid Venn-diagram taxonomy and reviews the emerging open-set/open-world HAR literature.

  2. A Survey on Efficiency Optimization Techniques for DNN-based Video Analytics: Process Systems, Algorithms, and Applications

    cs.CV 2025-07 conditional novelty 3.0 of 10

    A bottom-up survey of efficiency optimization techniques for DNN-based video analytics, spanning storage, computing, algorithms, and applications.

Pith tools