REVIEW 2 cited by
Human Action Recognition and Prediction: A Survey
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Derived from rapid advances in computer vision and machine learning, video analysis tasks have been moving from inferring the present state to predicting the future state. Vision-based action recognition and prediction from videos are such tasks, where action recognition is to infer human actions (present state) based upon complete action executions, and action prediction to predict human actions (future state) based upon incomplete action executions. These two tasks have become particularly prevalent topics recently because of their explosively emerging real-world applications, such as visual surveillance, autonomous driving vehicle, entertainment, and video retrieval, etc. Many attempts have been devoted in the last a few decades in order to build a robust and effective framework for action recognition and prediction. In this paper, we survey the complete state-of-the-art techniques in action recognition and prediction. Existing models, popular algorithms, technical difficulties, popular action databases, evaluation protocols, and promising future directions are also provided with systematic discussions.
Forward citations
Cited by 2 Pith papers
-
Fitting, Comparison, and Alignment of Trajectories on Positive Semi-Definite Matrices with Application to Action Recognition
A skeleton-only action recognition method on PSD matrices with a quotient metric, Bezier curve fitting, and global alignment kernel achieves 97.99% on UTKinect, 96.16% on KTH, and 92.44% on UAV-Gesture.
-
PISEP^2: Pseudo Image Sequence Evolution based 3D Pose Prediction
A non-recursive encoder-dynamics-decoder network, fed with 3D joint coordinates rearranged into small pseudo-images, predicts future poses and outperforms two baselines on G3D and a filtered NTU dataset.
Discussion (0). Continue with ORCID to comment.