Pith. sign in

REVIEW 1 cited by

Deep Keyframe Detection in Human Action Videos

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1804.10021 v1 pith:BQKBXZBE submitted 2018-04-26 cs.CV

classification cs.CV
keywords framesactionhumanvideovideosconvnetdetectionaimed
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Detecting representative frames in videos based on human actions is quite challenging because of the combined factors of human pose in action and the background. This paper addresses this problem and formulates the key frame detection as one of finding the video frames that optimally maximally contribute to differentiating the underlying action category from all other categories. To this end, we introduce a deep two-stream ConvNet for key frame detection in videos that learns to directly predict the location of key frames. Our key idea is to automatically generate labeled data for the CNN learning using a supervised linear discriminant method. While the training data is generated taking many different human action videos into account, the trained CNN can predict the importance of frames from a single video. We specify a new ConvNet framework, consisting of a summarizer and discriminator. The summarizer is a two-stream ConvNet aimed at, first, capturing the appearance and motion features of video frames, and then encoding the obtained appearance and motion features for video representation. The discriminator is a fitting function aimed at distinguishing between the key frames and others in the video. We conduct experiments on a challenging human action dataset UCF101 and show that our method can detect key frames with high accuracy.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Temporal Action Localization with Cross Layer Task Decoupling and Refinement

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A temporal action localization model that decouples classification and localization across feature pyramid layers and fuses instant, local, and global temporal features achieves state-of-the-art results on five benchmarks.

Pith tools