Pith. sign in

REVIEW 5 cited by

Two-Stream Convolutional Networks for Action Recognition in Videos

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1406.2199 v2 pith:M3ZEIEXK submitted 2014-06-09 cs.CV

classification cs.CV
keywords actionnetworkstrainedvideoarchitectureclassificationconvnetconvolutional
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We investigate architectures of discriminatively trained deep Convolutional Networks (ConvNets) for action recognition in video. The challenge is to capture the complementary information on appearance from still frames and motion between frames. We also aim to generalise the best performing hand-crafted features within a data-driven learning framework. Our contribution is three-fold. First, we propose a two-stream ConvNet architecture which incorporates spatial and temporal networks. Second, we demonstrate that a ConvNet trained on multi-frame dense optical flow is able to achieve very good performance in spite of limited training data. Finally, we show that multi-task learning, applied to two different action classification datasets, can be used to increase the amount of training data and improve the performance on both. Our architecture is trained and evaluated on the standard video actions benchmarks of UCF-101 and HMDB-51, where it is competitive with the state of the art. It also exceeds by a large margin previous attempts to use deep nets for video classification.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 5,369 citations worldwide. Full citation record

  1. Sign Language Question Answering: A New Task, Benchmark, and Baseline for Sign Language Understanding

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Sign Language QA benchmarks are introduced from PHOENIX14T and CSL-Daily via template-generated questions, and a question-conditioned baseline outperforms video-language and cascaded baselines.

  2. MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning

    cs.CV 2026-07 conditional novelty 6.0 of 10

    On MarineEVT, an event-centric 20K-pair marine video QA benchmark, EVT-R1 with tool-integrated RL scores 48.89 average accuracy, 5.22 points above the best untuned open-source VLM and 8.54 above the best tool-using co...

  3. Neutrino Fingerprints: Image-Based Encodings of IceCube Events for CNN Direction Reconstruction

    astro-ph.IM 2026-06 unverdicted novelty 6.0 of 10

    IceCube events are encoded as 72x72x3 images and processed by ResNet18 to reach 1.10 rad mean angular error in neutrino direction reconstruction.

  4. Predicting Soccer Penalty Kick Direction Using Human Action Recognition

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A HAR-based two-stream classifier predicts penalty kick direction from kicker motion with 63.9% binary accuracy, surpassing the goalkeeper baseline of 54.2%.

  5. EEvAct: Early Event-Based Action Recognition with High-Rate Two-Stream Spiking Neural Networks

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A high-rate two-stream spiking network with a lightweight gated fusion unit achieves 94.9% on THU EACT-50 and enables early prediction within 100 ms.

Pith tools