Pith. sign in

REVIEW 8 cited by

CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.04744 v2 pith:FRNJJDKF submitted 2019-10-10 cs.CV

classification cs.CV
keywords videocaterdatasetdatasetsobjectspatiotemporalarchitecturescontrollable
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Computer vision has undergone a dramatic revolution in performance, driven in large part through deep features trained on large-scale supervised datasets. However, much of these improvements have focused on static image analysis; video understanding has seen rather modest improvements. Even though new datasets and spatiotemporal models have been proposed, simple frame-by-frame classification methods often still remain competitive. We posit that current video datasets are plagued with implicit biases over scene and object structure that can dwarf variations in temporal structure. In this work, we build a video dataset with fully observable and controllable object and scene bias, and which truly requires spatiotemporal understanding in order to be solved. Our dataset, named CATER, is rendered synthetically using a library of standard 3D objects, and tests the ability to recognize compositions of object movements that require long-term reasoning. In addition to being a challenging dataset, CATER also provides a plethora of diagnostic tools to analyze modern spatiotemporal video architectures by being completely observable and controllable. Using CATER, we provide insights into some of the most recent state of the art deep video architectures.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new 23-dimension benchmark finds that even the best VLMs score near random on motion trajectory, temporal extension, and several prediction tasks, far below humans, suggesting weak internal world models.

  2. Finding the Trigger: Causal Abductive Reasoning on Video Events

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A new video benchmark and graph-based method for identifying the root-cause trigger event behind a target event, using counterfactually generated labels.

  3. Video Representation Learning with Joint-Embedding Predictive Architectures

    cs.CV 2024-12 conditional novelty 6.0 of 10

    VJ-VCR applies variance-covariance regularization to a video joint-embedding predictive architecture and beats a generative baseline at probing dynamics from frozen representations.

  4. IMBench: A Benchmark for Intuitive Robotic Manipulation

    cs.RO 2026-07 conditional novelty 5.0 of 10

    IMBench is a 35-task robosuite benchmark with a three-stage evaluation showing current VLMs and robot policies fail to convert physical reasoning into executable manipulation.

  5. An Empirical Study of Autoregressive Pre-training from Videos

    cs.CV 2025-01 conditional novelty 5.0 of 10

    Autoregressive next-token prediction on video and image tokens yields competitive visual representations across recognition, tracking, and robotics benchmarks, with scaling laws that are slower than those of language models.

  6. VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models

    cs.CV 2024-11 conditional novelty 5.0 of 10

    A new controllable synthetic-video benchmark shows that even state-of-the-art video-language models struggle with abstract and symbolic video cognition, with accuracy falling as task difficulty rises.

  7. Video Understanding by Design: How Datasets Shape Video Models

    cs.CV 2025-09 reject novelty 4.0 of 10

    A dataset-centric framework that explains video architectures as responses to structural properties of benchmark datasets.

  8. Do Language Models Understand Time?

    cs.CV 2024-12 conditional novelty 3.0 of 10

    A survey arguing that video-LLMs rely on pretrained encoders and short-biased datasets, leaving them weak at long-term temporal reasoning such as causality and event progression.

Pith tools