Multimodal foundation models outperform speech and music models for closed-set source attribution of singing voice deepfakes on CtrSVDD, with Chernoff-distance fusion of LanguageBind and ImageBind reaching 91.2% accuracy.
Table 2 presents the evaluation scores for modeling with various combinations of SFMs
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models
Multimodal foundation models outperform speech and music models for closed-set source attribution of singing voice deepfakes on CtrSVDD, with Chernoff-distance fusion of LanguageBind and ImageBind reaching 91.2% accuracy.