A two-branch 2D/3D CNN with attention fusion and a bidirectional ConvLSTM yields a marginal LRW accuracy gain, but its LRW-1000 score is below the cited state of the art.
Lipreading from color video
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2019 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Multi-Grained Spatio-temporal Modeling for Lip-reading
A two-branch 2D/3D CNN with attention fusion and a bidirectional ConvLSTM yields a marginal LRW accuracy gain, but its LRW-1000 score is below the cited state of the art.