Pith. sign in

REVIEW 1 cited by

Deep Learning for Lip Reading using Audio-Visual Information for Urdu Language

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1802.05521 v1 pith:5XV4TW6R submitted 2018-02-15 cs.CV cs.SDeess.AS

classification cs.CVcs.SDeess.AS
keywords learningvisualwordsaudio-visualdeepinformationlanguagelip-reading
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Human lip-reading is a challenging task. It requires not only knowledge of underlying language but also visual clues to predict spoken words. Experts need certain level of experience and understanding of visual expressions learning to decode spoken words. Now-a-days, with the help of deep learning it is possible to translate lip sequences into meaningful words. The speech recognition in the noisy environments can be increased with the visual information [1]. To demonstrate this, in this project, we have tried to train two different deep-learning models for lip-reading: first one for video sequences using spatiotemporal convolution neural network, Bi-gated recurrent neural network and Connectionist Temporal Classification Loss, and second for audio that inputs the MFCC features to a layer of LSTM cells and output the sequence. We have also collected a small audio-visual dataset to train and test our model. Our target is to integrate our both models to improve the speech recognition in the noisy environment

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Integrating Persian Lip Reading in Surena-V Humanoid Robot for Human-Robot Interaction

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A custom 7-word Persian lip-reading dataset is used to train an LSTM that reports 89% accuracy and is deployed on the Surena-V humanoid robot for real-time command recognition.

Pith tools