REVIEW 5 cited by
Two-Stream Convolutional Networks for Action Recognition in Videos
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We investigate architectures of discriminatively trained deep Convolutional Networks (ConvNets) for action recognition in video. The challenge is to capture the complementary information on appearance from still frames and motion between frames. We also aim to generalise the best performing hand-crafted features within a data-driven learning framework. Our contribution is three-fold. First, we propose a two-stream ConvNet architecture which incorporates spatial and temporal networks. Second, we demonstrate that a ConvNet trained on multi-frame dense optical flow is able to achieve very good performance in spite of limited training data. Finally, we show that multi-task learning, applied to two different action classification datasets, can be used to increase the amount of training data and improve the performance on both. Our architecture is trained and evaluated on the standard video actions benchmarks of UCF-101 and HMDB-51, where it is competitive with the state of the art. It also exceeds by a large margin previous attempts to use deep nets for video classification.
Forward citations
Cited by 5 Pith papers
-
Sign Language Question Answering: A New Task, Benchmark, and Baseline for Sign Language Understanding
Sign Language QA benchmarks are introduced from PHOENIX14T and CSL-Daily via template-generated questions, and a question-conditioned baseline outperforms video-language and cascaded baselines.
-
MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning
On MarineEVT, an event-centric 20K-pair marine video QA benchmark, EVT-R1 with tool-integrated RL scores 48.89 average accuracy, 5.22 points above the best untuned open-source VLM and 8.54 above the best tool-using co...
-
Neutrino Fingerprints: Image-Based Encodings of IceCube Events for CNN Direction Reconstruction
IceCube events are encoded as 72x72x3 images and processed by ResNet18 to reach 1.10 rad mean angular error in neutrino direction reconstruction.
-
Predicting Soccer Penalty Kick Direction Using Human Action Recognition
A HAR-based two-stream classifier predicts penalty kick direction from kicker motion with 63.9% binary accuracy, surpassing the goalkeeper baseline of 54.2%.
-
EEvAct: Early Event-Based Action Recognition with High-Rate Two-Stream Spiking Neural Networks
A high-rate two-stream spiking network with a lightweight gated fusion unit achieves 94.9% on THU EACT-50 and enables early prediction within 100 ms.
Discussion (0). Sign in to comment.