Pith. sign in

REVIEW 4 cited by

Learning Spatiotemporal Features with 3D Convolutional Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1412.0767 v4 pith:MIKPHBH5 submitted 2014-12-02 cs.CV

Learning Spatiotemporal Features with 3D Convolutional Networks

classification cs.CV
keywords convnetsconvolutionalfeatureslearningsimplespatiotemporalbenchmarksbest
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X LinkedIn Reddit HN
read the original abstract

We propose a simple, yet effective approach for spatiotemporal feature learning using deep 3-dimensional convolutional networks (3D ConvNets) trained on a large scale supervised video dataset. Our findings are three-fold: 1) 3D ConvNets are more suitable for spatiotemporal feature learning compared to 2D ConvNets; 2) A homogeneous architecture with small 3x3x3 convolution kernels in all layers is among the best performing architectures for 3D ConvNets; and 3) Our learned features, namely C3D (Convolutional 3D), with a simple linear classifier outperform state-of-the-art methods on 4 different benchmarks and are comparable with current best methods on the other 2 benchmarks. In addition, the features are compact: achieving 52.8% accuracy on UCF101 dataset with only 10 dimensions and also very efficient to compute due to the fast inference of ConvNets. Finally, they are conceptually very simple and easy to train and use.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. AULLM++: Structured-Token-Conditioned Large Language Models for Micro-Expression Action Unit Detection

    cs.CV 2026-03 conditional novelty 5.5

    AULLM++ fuses multi-granularity visual tokens with FACS-prior AU graph instructions into an LLM prompt and uses counterfactual consistency training to improve micro-expression AU detection and cross-domain Macro-F1.

  2. Unsupervised learning for the systematic identification of nondispersive wave packets in driven helium

    quant-ph 2026-05 unverdicted novelty 5.0

    Unsupervised CNN embedding and clustering of Floquet states recovers known nondispersive wave packet regimes in driven helium without labels.

  3. Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders

    cs.CV 2025-09 conditional novelty 5.0

    Matching video diffusion transformer tokens to concatenated DINOv2 and SAM2 features improves FVD/FID and speeds convergence, e.g., 400K-step fusion beats 1M-step baseline on UCF-101.

  4. Spatio-Temporal Wildfire Spread Prediction in Canada using a Video Swin-Hybrid-U-Net and Satellite Imagery

    cs.CV 2026-06 unverdicted novelty 4.0

    Hybrid Video Swin-U-Net forecasts next-day fire incidence maps from spatio-temporal satellite and meteorological sequences for major Canadian wildfires 2014-2023.