A custom 7-word Persian lip-reading dataset is used to train an LSTM that reports 89% accuracy and is deployed on the Surena-V humanoid robot for real-time command recognition.
Word-level Persian Lipreading Dataset
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Lip-reading has made impressive progress in recent years, driven by advances in deep learning. Nonetheless, the prerequisite such advances is a suitable dataset. This paper provides a new in-the-wild dataset for Persian word-level lipreading containing 244,000 videos from approximately 1,800 speakers. We evaluated the state-of-the-art method in this field and used a novel approach for word-level lip-reading. In this method, we used the AV-HuBERT model for feature extraction and obtained significantly better performance on our dataset.
citation-role summary
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
dataset 1polarities
support 1representative citing papers
citing papers explorer
-
Integrating Persian Lip Reading in Surena-V Humanoid Robot for Human-Robot Interaction
A custom 7-word Persian lip-reading dataset is used to train an LSTM that reports 89% accuracy and is deployed on the Surena-V humanoid robot for real-time command recognition.