A fully binarized Conv3D-LSTM model for video inference runs gesture recognition on Jester with 1.01 MB weights and 6.34 GBOPs, with an 8-9 point accuracy drop versus compact full-precision baselines.
The Jester dataset: A large-scale video dataset of human gestures,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
BILLNET: A Binarized Conv3D-LSTM Network with Logic-gated residual architecture for hardware-efficient video inference
A fully binarized Conv3D-LSTM model for video inference runs gesture recognition on Jester with 1.01 MB weights and 6.34 GBOPs, with an 8-9 point accuracy drop versus compact full-precision baselines.