Pith. sign in

REVIEW 4 cited by

A Comprehensive Study of Deep Video Action Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2012.06567 v1 pith:LFXRB7TC submitted 2020-12-11 cs.CV cs.MM

classification cs.CVcs.MM
keywords videoactionrecognitiondeepdatasetslearningmodelscomprehensive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Video action recognition is one of the representative tasks for video understanding. Over the last decade, we have witnessed great advancements in video action recognition thanks to the emergence of deep learning. But we also encountered new challenges, including modeling long-range temporal information in videos, high computation costs, and incomparable results due to datasets and evaluation protocol variances. In this paper, we provide a comprehensive survey of over 200 existing papers on deep learning for video action recognition. We first introduce the 17 video action recognition datasets that influenced the design of models. Then we present video action recognition models in chronological order: starting with early attempts at adapting deep learning, then to the two-stream networks, followed by the adoption of 3D convolutional kernels, and finally to the recent compute-efficient models. In addition, we benchmark popular methods on several representative datasets and release code for reproducibility. In the end, we discuss open problems and shed light on opportunities for video action recognition to facilitate new research ideas.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Efficient Video Dataset Distillation via Cluster-Guided Prototype Blending

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A construction-based video distillation pipeline selects teacher-confident clips, allocates slots to feature-space clusters, and blends prototype-anchor pairs with matched soft labels, avoiding gradient updates of sto...

  2. What to Do Next? Memorizing skills from Egocentric Instructional Video

    cs.LG 2025-07 conditional novelty 6.0 of 10

    TAMFormer combines a topological affordance memory with a transformer decoder to plan high-level actions from egocentric video and replan when execution deviates.

  3. Leveraging Vision-Language Large Models for Interpretable Video Action Recognition with Semantic Tokenization

    cs.CV 2025-09 reject novelty 3.0 of 10

    LVLM-VAR transforms video into 'semantic action tokens' and uses a LoRA-tuned vision-language model to classify actions and generate explanations, reporting 94.1% on NTU RGB+D X-Sub.

  4. A Survey on Video Temporal Grounding with Multimodal Large Language Model

    cs.CV 2025-08 unverdicted novelty 3.0 of 10

    A taxonomized review of video temporal grounding with multimodal large language models, covering model roles, training paradigms, feature processing, benchmarks, and open problems.

Pith tools