Pith. sign in

REVIEW 3 cited by

Combining Residual Networks with LSTMs for Lipreading

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1703.04105 v4 pith:Y4C6URW7 submitted 2017-03-12 cs.CV

classification cs.CV
keywords lipreadingnetworksresidualwordabsoluteaccuracyarchitectureattains
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We propose an end-to-end deep learning architecture for word-level visual speech recognition. The system is a combination of spatiotemporal convolutional, residual and bidirectional Long Short-Term Memory networks. We train and evaluate it on the Lipreading In-The-Wild benchmark, a challenging database of 500-size target-words consisting of 1.28sec video excerpts from BBC TV broadcasts. The proposed network attains word accuracy equal to 83.0, yielding 6.8 absolute improvement over the current state-of-the-art, without using information about word boundaries during training or testing.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Text2Lip: Progressive Lip-Synced Talking Face Generation from Text via Viseme-Guided Rendering

    cs.CV 2025-08 reject novelty 6.0 of 10

    A text-only talking face generator maps text to visemes, hallucinates pseudo-audio, and renders lip-synced video; the paper's headline metrics are partially contradicted by its own tables.

  2. Detecting COPD Through Speech Analysis: A Dataset of Danish Speech and Machine Learning Approach

    cs.SD 2025-08 unverdicted novelty 5.0 of 10

    A new Danish speech dataset with 96 participants yields 67% COPD detection accuracy with openSMILE features and logistic regression, suggesting speech as a cross-linguistic screening signal.

  3. Multi-Grained Spatio-temporal Modeling for Lip-reading

    cs.CV 2019-08 conditional novelty 4.0 of 10

    A two-branch 2D/3D CNN with attention fusion and a bidirectional ConvLSTM yields a marginal LRW accuracy gain, but its LRW-1000 score is below the cited state of the art.

Pith tools