REVIEW 3 cited by
Combining Residual Networks with LSTMs for Lipreading
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We propose an end-to-end deep learning architecture for word-level visual speech recognition. The system is a combination of spatiotemporal convolutional, residual and bidirectional Long Short-Term Memory networks. We train and evaluate it on the Lipreading In-The-Wild benchmark, a challenging database of 500-size target-words consisting of 1.28sec video excerpts from BBC TV broadcasts. The proposed network attains word accuracy equal to 83.0, yielding 6.8 absolute improvement over the current state-of-the-art, without using information about word boundaries during training or testing.
Forward citations
Cited by 3 Pith papers
-
Text2Lip: Progressive Lip-Synced Talking Face Generation from Text via Viseme-Guided Rendering
A text-only talking face generator maps text to visemes, hallucinates pseudo-audio, and renders lip-synced video; the paper's headline metrics are partially contradicted by its own tables.
-
Detecting COPD Through Speech Analysis: A Dataset of Danish Speech and Machine Learning Approach
A new Danish speech dataset with 96 participants yields 67% COPD detection accuracy with openSMILE features and logistic regression, suggesting speech as a cross-linguistic screening signal.
-
Multi-Grained Spatio-temporal Modeling for Lip-reading
A two-branch 2D/3D CNN with attention fusion and a bidirectional ConvLSTM yields a marginal LRW accuracy gain, but its LRW-1000 score is below the cited state of the art.
Discussion (0). Continue with ORCID to comment.