REVIEW 8 cited by
The AVA-Kinetics Localized Human Actions Video Dataset
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
This paper describes the AVA-Kinetics localized human actions video dataset. The dataset is collected by annotating videos from the Kinetics-700 dataset using the AVA annotation protocol, and extending the original AVA dataset with these new AVA annotated Kinetics clips. The dataset contains over 230k clips annotated with the 80 AVA action classes for each of the humans in key-frames. We describe the annotation process and provide statistics about the new dataset. We also include a baseline evaluation using the Video Action Transformer Network on the AVA-Kinetics dataset, demonstrating improved performance for action classification on the AVA test set. The dataset can be downloaded from https://research.google.com/ava/
Forward citations
Cited by 8 Pith papers
-
DeepFake Doctor: Diagnosing and Treating Audio-Video Fake Detection
Proposes new cross-manipulation evaluation protocols for FakeAVCeleb and DeepSpeak v1, shows temporal jittering mitigates a leading-silence shortcut, and introduces the SIMBA baseline.
-
Visual WetlandBirds Dataset: Bird Species Identification and Behavior Recognition in Videos
A new video dataset with per-frame bird species and behavior annotations, covering 13 species and 7 behaviors, is released with baseline recognition results.
-
Multi-Modal Self-Supervised Learning for Surgical Feedback Effectiveness Assessment
A multimodal model combining transcribed surgical feedback and video clips predicts trainee behavior change with AUROC 0.70, modestly outperforming text alone.
-
From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach
The paper presents GEST, an event-graph representation of videos that is converted automatically into natural language and is also used as a teacher to pre-train end-to-end video captioning models.
-
ShotVL: Human-Centric Highlight Frame Retrieval via Language Queries
ShotVL, a fine-tuned InternVL, improves zero-shot highlight-frame retrieval on the new BestShot benchmark, but its zero-shot claim is weakened by a small benchmark-derived training set.
-
Video Understanding by Design: How Datasets Shape Video Models
A dataset-centric framework that explains video architectures as responses to structural properties of benchmark datasets.
-
Action Recognition based Industrial Safety Violation Detection
An action-aware PPE violation detector, built from SlowFast and YOLOv9, claims a 23% F1 improvement over generic PPE checks on a small private industrial dataset.
-
Enhancing Video Understanding: Deep Neural Networks for Spatiotemporal Analysis
A narrative review of spatiotemporal deep neural networks for video understanding, with tables of benchmark datasets and reported model results.
Discussion (0). Continue with ORCID to comment.