Pith. sign in

REVIEW 1 cited by

A Survey on Backbones for Deep Video Action Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.05584 v1 pith:PUDTMEBV submitted 2024-05-09 cs.CV cs.AI

A Survey on Backbones for Deep Video Action Recognition

classification cs.CV cs.AI
keywords methodsactionrecognitiondeepnetworksvideobackbonesintroduce
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Action recognition is a key technology in building interactive metaverses. With the rapid development of deep learning, methods in action recognition have also achieved great advancement. Researchers design and implement the backbones referring to multiple standpoints, which leads to the diversity of methods and encountering new challenges. This paper reviews several action recognition methods based on deep neural networks. We introduce these methods in three parts: 1) Two-Streams networks and their variants, which, specifically in this paper, use RGB video frame and optical flow modality as input; 2) 3D convolutional networks, which make efforts in taking advantage of RGB modality directly while extracting different motion information is no longer necessary; 3) Transformer-based methods, which introduce the model from natural language processing into computer vision and video understanding. We offer objective sights in this review and hopefully provide a reference for future research.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Boosting Skeleton-based Zero-Shot Action Recognition with Training-Free Test-Time Adaptation

    cs.CV 2025-12 conditional novelty 5.0

    A training-free cache of structured skeleton descriptors, fused with LLM-generated per-class weights, boosts zero-shot skeleton action recognition on NTU and PKU-MMD benchmarks by several points.