Pith. sign in

REVIEW 4 cited by

VideoPrism: A Foundational Visual Encoder for Video Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.13217 v3 pith:V7Y5M63P submitted 2024-02-20 cs.CV cs.AI

classification cs.CVcs.AI
keywords videovideoprismunderstandingencodertaskstextachievinganswering
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce VideoPrism, a general-purpose video encoder that tackles diverse video understanding tasks with a single frozen model. We pretrain VideoPrism on a heterogeneous corpus containing 36M high-quality video-caption pairs and 582M video clips with noisy parallel text (e.g., ASR transcripts). The pretraining approach improves upon masked autoencoding by global-local distillation of semantic video embeddings and a token shuffling scheme, enabling VideoPrism to focus primarily on the video modality while leveraging the invaluable text associated with videos. We extensively test VideoPrism on four broad groups of video understanding tasks, from web video question answering to CV for science, achieving state-of-the-art performance on 31 out of 33 video understanding benchmarks. Our models are released at https://github.com/google-deepmind/videoprism.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Prompting-MammAlps: Fine-Grained Text-to-Video Retrieval for Camera-Trap Data

    cs.CV 2026-07 accept novelty 7.0 of 10

    A camera-trap TVR benchmark of 135 ethology queries plus an interpretable SALMA-to-JSON plus constrained-LLM-parser pipeline yields 34% set F1, beating zero-shot VLMs at 18%.

  2. A non-invasive video-based method for individual identification of wildlife using gait dynamics

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A SAM3 + ResNet18 + VideoPrism pipeline produces gait embeddings that cluster individual animals of five species with high intra- and lower inter-individual cosine similarity under lateral walking views.

  3. SurgBench: A Unified Large-Scale Benchmark for Surgical Video Analysis

    cs.CV 2025-06 conditional novelty 5.0 of 10

    SurgBench assembles a 53-million-frame surgical video pretraining corpus from 16 sources plus a 72-task evaluation benchmark, and shows continual pretraining with VideoMAE improves accuracy by 7.9% top-3 over Kinetics...

  4. Infinite Video Understanding

    cs.CV 2025-07 conditional novelty 3.0 of 10

    The paper argues that video understanding research should aim at processing streams of arbitrary, unbounded duration and outlines the challenges, directions, and metrics needed.

Pith tools