VSRo-200 is the first large Romanian sentence-level lip-reading corpus; human labels beat pseudo-labels at matched size, but scaling pseudo-labels closes the gap and yields strong transfer to word recognition.
Lip reading in the wild
2 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.CV 2years
2026 2representative citing papers
SDTalk proposes a generalizable one-shot 3DGS talking head method that uses structured facial priors for complete reconstruction and dual-branch motion fields for dynamics, outperforming prior identity-specific approaches.
citing papers explorer
-
VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness
VSRo-200 is the first large Romanian sentence-level lip-reading corpus; human labels beat pseudo-labels at matched size, but scaling pseudo-labels closes the gap and yields strong transfer to word recognition.
-
SDTalk: Structured Facial Priors and Dual-Branch Motion Fields for Generalizable Gaussian Talking Head Synthesis
SDTalk proposes a generalizable one-shot 3DGS talking head method that uses structured facial priors for complete reconstruction and dual-branch motion fields for dynamics, outperforming prior identity-specific approaches.