VSRo-200 is the first large Romanian sentence-level lip-reading corpus; human labels beat pseudo-labels at matched size, but scaling pseudo-labels closes the gap and yields strong transfer to word recognition.
Pyannote. audio: neural building blocks for speaker diarization
2 Pith papers cite this work. Polarity classification is still indexing.
verdicts
CONDITIONAL 2representative citing papers
ZipVoice-Dialog is a flow-matching non-autoregressive model for zero-shot spoken dialogue generation that uses curriculum learning and speaker-turn embeddings, paired with a new 6.8k-hour OpenDialog dataset, and reports better speed and quality than autoregressive baselines.
citing papers explorer
-
VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness
VSRo-200 is the first large Romanian sentence-level lip-reading corpus; human labels beat pseudo-labels at matched size, but scaling pseudo-labels closes the gap and yields strong transfer to word recognition.
-
ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching
ZipVoice-Dialog is a flow-matching non-autoregressive model for zero-shot spoken dialogue generation that uses curriculum learning and speaker-turn embeddings, paired with a new 6.8k-hour OpenDialog dataset, and reports better speed and quality than autoregressive baselines.