A localized P.808 ACR listening-test pipeline for four languages is described, and URGENT 2025 analysis shows that subjective MOS and reference-free metrics can miss hallucinations, motivating additional phone-fidelity checks.
Vibravox: A Dataset of French Speech Captured with Body-conduction Audio Sensors
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Vibravox is a dataset compliant with the General Data Protection Regulation (GDPR) containing audio recordings using five different body-conduction audio sensors: two in-ear microphones, two bone conduction vibration pickups, and a laryngophone. The dataset also includes audio data from an airborne microphone used as a reference. The Vibravox corpus contains 45 hours per sensor of speech samples and physiological sounds recorded by 188 participants under different acoustic conditions imposed by a high order ambisonics 3D spatializer. Annotations about the recording conditions and linguistic transcriptions are also included in the corpus. We conducted a series of experiments on various speech-related tasks, including speech recognition, speech enhancement, and speaker verification. These experiments were carried out using state-of-the-art models to evaluate and compare their performances on signals captured by the different audio sensors offered by the Vibravox dataset, with the aim of gaining a better grasp of their individual characteristics.
citation-role summary
citation-polarity summary
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1roles
dataset 1polarities
use dataset 1representative citing papers
citing papers explorer
-
P.808 Multilingual Speech Enhancement Testing: Approach and Results of URGENT 2025 Challenge
A localized P.808 ACR listening-test pipeline for four languages is described, and URGENT 2025 analysis shows that subjective MOS and reference-free metrics can miss hallucinations, motivating additional phone-fidelity checks.