REVIEW 42 cited by
VoxCeleb: a large-scale speaker identification dataset
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Most existing datasets for speaker identification contain samples obtained under quite constrained conditions, and are usually hand-annotated, hence limited in size. The goal of this paper is to generate a large scale text-independent speaker identification dataset collected 'in the wild'. We make two contributions. First, we propose a fully automated pipeline based on computer vision techniques to create the dataset from open-source media. Our pipeline involves obtaining videos from YouTube; performing active speaker verification using a two-stream synchronization Convolutional Neural Network (CNN), and confirming the identity of the speaker using CNN based facial recognition. We use this pipeline to curate VoxCeleb which contains hundreds of thousands of 'real world' utterances for over 1,000 celebrities. Our second contribution is to apply and compare various state of the art speaker identification techniques on our dataset to establish baseline performance. We show that a CNN based architecture obtains the best performance for both identification and verification.
Forward citations
Cited by 42 Pith papers
-
EdgeFaaS: A Function-based Framework for Edge Computing
EdgeFaaS unifies IoT, edge, and cloud resources behind a function-as-a-service interface with workflow and storage virtualization, demonstrated on three workloads across 100+ devices.
-
PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction
Proxy-supervised joint fine-tuning of a BSRNN separator with ASR, speaker-similarity, VAD and DNSMOS losses on a new 71k real-conversation corpus yields the best SIM and timing F1 on REAL-T.
-
A Camera-Native Talking-Head Video Dataset for Various Computer Vision Tasks
A 847-clip camera-native talking-head dataset with MOS annotations and a 120-clip stratified benchmark shows content type and background processing significantly change modern codec BD-rate savings.
-
Query-Based Asymmetric Modeling with Decoupled Input-Output Rates for Speech Restoration
TF-Restormer restores degraded speech at arbitrary input-output sampling rates in a single model, using a heavy encoder and a lightweight query-based decoder to generate missing high-frequency bands.
-
LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model
LLaSO releases a 3.8B speech-language model, 25.5M training instances, and an evaluation benchmark, claiming a normalized score of 0.72.
-
Audio-Assisted Face Video Restoration with Temporal and Identity Complementary Learning
GAVN outperforms state-of-the-art face video restoration on compression artifact removal, deblurring, and super-resolution by fusing audio, landmark, and temporal identity features.
-
SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation
SpeakerVid-5M provides 5.2 million audio-visual human clips (8,743 hours) with rich annotations and a dyadic interaction benchmark for training interactive virtual humans.
-
SPBA: Utilizing Speech Large Language Model for Backdoor Attacks on Speech Classification Models
A speech backdoor attack uses SLLM-generated timbre and emotion triggers with MGDA-balanced training to implant multiple effective backdoors in speech classifiers.
-
Revisiting Active Speaker Detection: An In-the-Wild Benchmark for Generalization and Robustness
UniTalk is a larger and more diverse active speaker detection benchmark on which state-of-the-art models underperform, and it improves cross-dataset generalization when used as a training source.
-
Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits
A unified multi-trait speech benchmark, Vox-Profile, evaluates foundation models on age, sex, accent, emotion, voice quality, fluency, and expressiveness, and demonstrates downstream uses in ASR analysis and speech ge...
-
Multi-band Frequency Reconstruction for Neural Psychoacoustic Coding
MUFFIN is a neural audio codec that quantizes separate latent frequency bands with dedicated codebooks, claiming SOTA reconstruction quality and a competitive 12.5 Hz ultra-low-rate variant.
-
Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis
Continuous autoregressive text-to-speech with a Gaussian-mixture codec matches or beats a discrete-codec VALL-E baseline with a fraction of the language model parameters.
-
Inclusion 2024 Global Multimedia Deepfake Detection Challenge: Towards Multi-dimensional Face Forgery Detection
The Inclusion 2024 challenge introduces MultiFF, a large and diverse face forgery benchmark, and reports that top solutions achieve high AUC but weak true-positive rates at low false-positive rates on unseen forgery types.
-
Bridging the Data Provenance Gap Across Text, Speech and Video
A manual audit of nearly 4,000 text, speech, and video datasets finds AI training data increasingly comes from web and social media sources, carries hidden non-commercial restrictions, and remains Western-centric with...
-
GLCF: A Global-Local Multimodal Coherence Analysis Framework for Talking Face Generation Detection
MSTF, a 143k-video talking face dataset built from 22 forgery techniques, and a global-local audio-visual coherence detector, outperform prior deepfake detectors on this benchmark.
-
CA-SSLR: Condition-Aware Self-Supervised Learning Representation for Generalized Speech Processing
CA-SSLR injects condition-aware language and speaker embeddings into a frozen SSL encoder via lightweight FiLM-style adapters, improving ASR, LID, and SV performance and transfer.
-
Synergizing Motion and Appearance: Multi-Scale Compensatory Codebooks for Talking Head Video Generation
A talking-head video generation method that jointly learns multi-scale motion and appearance codebooks and compensates warped features with transformer-based code retrieval, showing improved reconstruction quality on ...
-
BackdoorMBTI: A Backdoor Learning Multimodal Benchmark Tool Kit for Backdoor Defense Evaluation
BackdoorMBTI is the first backdoor security benchmark and toolkit that covers image, text, and audio modalities with a unified evaluation pipeline.
-
Personal VAD: Speaker-Conditioned Voice Activity Detection
A 130K-parameter LSTM network conditioned on an enrolled speaker's embedding detects that speaker's voice activity at frame level, outperforming a combined VAD plus speaker verification baseline.
-
Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks
A system-first taxonomy and literature synthesis of multimodal unlearning across vision, language, video, and audio, with datasets, benchmarks, metrics, applications, and open challenges.
-
SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings
A parameter-efficient adapter bridging Whisper and TinyLlama reports relative improvements in speech recognition, named entity recognition, and sentiment analysis on low-resource benchmarks.
-
Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody
Maestro-EVC independently controls content, speaker, and emotion in voice conversion using separate references and explicit prosody modeling, outperforming StyleVC and ZEST on emotion similarity and prosody.
-
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs
OmniEval releases a Chinese-English, audio-visual-text benchmark with fine-grained temporal grounding questions, and reports that today's omni-modal models score low and depend mainly on textual cues.
-
Few-Shot Speech Deepfake Detection Adaptation with Gaussian Processes
ADD-GP, a Gaussian Process classifier with XLS-R speech embeddings, adapts to unseen TTS models with as few as 5 samples and achieves state-of-the-art low error rates on the new LibriFake benchmark.
-
Multiplicity is an Inevitable and Inherent Challenge in Multimodal Learning
The authors argue that many-to-many cross-modal correspondences, termed 'multiplicity', are inevitable and require rethinking multimodal learning, training, evaluation, and dataset construction.
-
Introducing voice timbre attribute detection
This paper defines voice timbre attribute detection as pairwise voice comparison, and shows a FACodec-based classifier generalizes better to unseen speakers than an ECAPA-TDNN baseline.
-
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training
A masked generative model pre-trained on unlabeled speech then fine-tuned per task matches or beats task-specific systems across TTS, voice conversion, speaker extraction, enhancement, and lip-to-speech.
-
Dementia classification from spontaneous speech using wrapper-based feature selection
Using full-recording acoustic features and wrapper selection, the Extreme Minimal Learning Machine provides competitive dementia classification accuracy at lower computational cost on ADReSS and Pitt datasets.
-
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis
VoiceDiT generates speech and matching environmental sounds from text, audio, or image prompts, and reports better speech intelligibility than the VoiceLDM baseline.
-
A Study on Angular Based Embedding Learning for Text-independent Speaker Verification
Angular margin losses with an inter-class regularization reduce speaker verification EER from 5.33% to 4.45% on a VoxCeleb test set compared with a softmax baseline.
-
Towards Robust Uncertainty-Aware Speaker Modeling
Inter- and intra-speaker hardness in an uncertainty-aware softmax plus source-prior uncertainty calibration improves speaker verification reliability under domain shift.
-
FreeTalk:A plug-and-play and black-box defense against speech synthesis attacks
FreeTalk adds masked, smoothed frequency-domain noise, optimized against a speaker-embedding model, to keep voice-cloning models from reproducing a victim's voice, while preserving speech-to-text accuracy.
-
A Summer Meridional Subsurface Temperature Dipole Mode in the South China Sea
The manuscript body does not match the abstract, so the ocean dipole claim is unsupported by any presented evidence.
-
PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association
PAEFF aligns face and voice embeddings in hyperbolic space before gated feature fusion and reports improved face-voice verification and matching on VoxCeleb1.
-
The Voice Timbre Attribute Detection 2025 Challenge Evaluation Plan
It is an evaluation plan for a competition where systems compare two voices and decide which one is stronger on a given timbre descriptor.
-
DiffAttack: Diffusion-based Timbre-reserved Adversarial Attack in Speaker Identification
DiffAttack injects adversarial constraints into the reverse diffusion process of DiffVC, boosting targeted speaker-identification attack success from 28.4% to 65.8% on LibriTTS while retaining speech quality.
-
An Overview and Discussion of the Suitability of Existing Speech Datasets to Train Machine Learning Models for Collective Problem Solving
Existing SLU speech datasets are insufficient for training ML models on collaborative problem solving because they lack multimodal, longitudinal, ambiguous, and team-dynamics data.
-
VQTalker: Towards Multilingual Talking Avatars through Facial Motion Tokenization
VQTalker learns a discrete facial-motion codebook via GRFSQ and generates talking heads from speech tokens, reporting improved lip sync on non-Indo-European languages at lower bitrate.
-
Developing an Effective Training Dataset to Enhance the Performance of AI-based Speaker Separation Systems
A playback-and-record method creates a realistic two-speaker training set that yields up to 1.65 dB SI-SDR improvement over synthetic training.
-
VAE-based Domain Adaptation for Speaker Verification
Adapting a VAE normalization model on out-of-domain x-vectors improved speaker verification EER from 18.51% to 12.73% on a small proprietary test set.
-
Face Deepfakes -- A Comprehensive Review
A review of face deepfake generation and detection finds that off-the-shelf deepfake tools such as Wav2Lip and SimSwap achieve high attack success rates against lightweight face recognition models.
-
Survey on Deep Neural Networks in Speech and Vision Systems
A broad survey of deep learning architectures and systems for vision and speech, with an emphasis on mobile deployment and emerging applications, containing no new results.
Discussion (0). Continue with ORCID to comment.