Pith. sign in

REVIEW 3 cited by

The AVA-Kinetics Localized Human Actions Video Dataset

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2005.00214 v2 pith:SOEYGFE5 submitted 2020-05-01 cs.CV cs.LGeess.IV

classification cs.CVcs.LGeess.IV
keywords datasetactionava-kineticsvideoactionsannotatedannotationclips
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper describes the AVA-Kinetics localized human actions video dataset. The dataset is collected by annotating videos from the Kinetics-700 dataset using the AVA annotation protocol, and extending the original AVA dataset with these new AVA annotated Kinetics clips. The dataset contains over 230k clips annotated with the 80 AVA action classes for each of the humans in key-frames. We describe the annotation process and provide statistics about the new dataset. We also include a baseline evaluation using the Video Action Transformer Network on the AVA-Kinetics dataset, demonstrating improved performance for action classification on the AVA test set. The dataset can be downloaded from https://research.google.com/ava/

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DeepFake Doctor: Diagnosing and Treating Audio-Video Fake Detection

    cs.MM 2025-06 conditional novelty 6.0 of 10

    Proposes new cross-manipulation evaluation protocols for FakeAVCeleb and DeepSpeak v1, shows temporal jittering mitigates a leading-silence shortcut, and introduces the SIMBA baseline.

  2. From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach

    cs.CV 2025-07 conditional novelty 5.0 of 10

    The paper presents GEST, an event-graph representation of videos that is converted automatically into natural language and is also used as a teacher to pre-train end-to-end video captioning models.

  3. Video Understanding by Design: How Datasets Shape Video Models

    cs.CV 2025-09 reject novelty 4.0 of 10

    A dataset-centric framework that explains video architectures as responses to structural properties of benchmark datasets.

Pith tools