Multimodal foundation models outperform speech and music models for closed-set source attribution of singing voice deepfakes on CtrSVDD, with Chernoff-distance fusion of LanguageBind and ImageBind reaching 91.2% accuracy.
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
In this work, we introduce the task of singing voice deepfake source attribution (SVDSA). We hypothesize that multimodal foundation models (MMFMs) such as ImageBind, LanguageBind will be most effective for SVDSA as they are better equipped for capturing subtle source-specific characteristics-such as unique timbre, pitch manipulation, or synthesis artifacts of each singing voice deepfake source due to their cross-modality pre-training. Our experiments with MMFMs, speech foundation models and music foundation models verify the hypothesis that MMFMs are the most effective for SVDSA. Furthermore, inspired from related research, we also explore fusion of foundation models (FMs) for improved SVDSA. To this end, we propose a novel framework, COFFE which employs Chernoff Distance as novel loss function for effective fusion of FMs. Through COFFE with the symphony of MMFMs, we attain the topmost performance in comparison to all the individual FMs and baseline fusion methods.
citation-role summary
citation-polarity summary
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models
Multimodal foundation models outperform speech and music models for closed-set source attribution of singing voice deepfakes on CtrSVDD, with Chernoff-distance fusion of LanguageBind and ImageBind reaching 91.2% accuracy.