Pith. sign in

REVIEW 7 cited by

Perception Test: A Diagnostic Benchmark for Multimodal Video Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.13786 v2 pith:O74H2BXZ submitted 2023-05-23 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords perceptiontestvideobenchmarkmodelsmultimodalavailablebaseline
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We propose a novel multimodal video benchmark - the Perception Test - to evaluate the perception and reasoning skills of pre-trained multimodal models (e.g. Flamingo, SeViLA, or GPT-4). Compared to existing benchmarks that focus on computational tasks (e.g. classification, detection or tracking), the Perception Test focuses on skills (Memory, Abstraction, Physics, Semantics) and types of reasoning (descriptive, explanatory, predictive, counterfactual) across video, audio, and text modalities, to provide a comprehensive and efficient evaluation tool. The benchmark probes pre-trained models for their transfer capabilities, in a zero-shot / few-shot or limited finetuning regime. For these purposes, the Perception Test introduces 11.6k real-world videos, 23s average length, designed to show perceptually interesting situations, filmed by around 100 participants worldwide. The videos are densely annotated with six types of labels (multiple-choice and grounded video question-answers, object and point tracks, temporal action and sound segments), enabling both language and non-language evaluations. The fine-tuning and validation splits of the benchmark are publicly available (CC-BY license), in addition to a challenge server with a held-out test split. Human baseline results compared to state-of-the-art video QA models show a substantial gap in performance (91.4% vs 46.2%), suggesting that there is significant room for improvement in multimodal video understanding. Dataset, baseline code, and challenge server are available at https://github.com/deepmind/perception_test

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors

    cs.CV 2025-08 conditional novelty 6.0 of 10

    LangDC compresses video tokens dynamically by converting clips into captions from a small language model, cutting compute by 49% with near-parity accuracy.

  2. CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    CausalVQA provides 793 paired real-video causal reasoning questions on which the best multimodal model scores 61.66% versus 84.78% for humans, with the largest gaps on anticipation and hypothetical questions.

  3. Correspondence of high-dimensional emotion structures elicited by video clips between humans and Multimodal LLMs

    cs.AI 2025-05 conditional novelty 6.0 of 10

    Multimodal LLMs capture the coarse category-level structure of human emotional responses to videos, but not fine-grained item-level emotion structure.

  4. J-EDI QA: Benchmark for deep-sea organism-specific multimodal LLM

    cs.CV 2024-12 conditional novelty 6.0 of 10

    J-EDI QA is a new 100-image Japanese multiple-choice benchmark for deep-sea organism identification; OpenAI o1 scored 50%, GPT-4o 39%, and non-expert humans about 40%.

  5. LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Inserting a pixel-shuffle plus residual patch-merge layer inside the vision encoder compresses visual tokens more efficiently than post-encoder compression, at modest accuracy cost.

  6. MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

    cs.CL 2024-12 conditional novelty 5.0 of 10

    A fully open pipeline that rewrites multimodal instruction data into CoT-style rationales yields a 12M dataset and an 8B model with strong benchmark gains, though some evaluation benchmarks overlap the training data.

  7. Movie2Story: A framework for understanding videos and telling stories in the form of novel text

    cs.CV 2024-12 reject novelty 4.0 of 10

    MSBench evaluates video-plus-audio to novel-style story generation; the M2S pipeline combines existing video, speech, emotion, and speaker tools with an LLM and reportedly beats video-only baselines.

Pith tools