REVIEW 5 cited by
From Audio Deepfake Detection to AI-Generated Music Detection -- A Pathway and Overview
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
As Artificial Intelligence (AI) technologies continue to evolve, their use in generating realistic, contextually appropriate content has expanded into various domains. Music, an art form and medium for entertainment, deeply rooted into human culture, is seeing an increased involvement of AI into its production. However, despite the effective application of AI music generation (AIGM) tools, the unregulated use of them raises concerns about potential negative impacts on the music industry, copyright and artistic integrity, underscoring the importance of effective AIGM detection. This paper provides an overview of existing AIGM detection methods. To lay a foundation to the general workings and challenges of AIGM detection, we first review general principles of AIGM, including recent advancements in deepfake audios, as well as multimodal detection techniques. We further propose a potential pathway for leveraging foundation models from audio deepfake detection to AIGM detection. Additionally, we discuss implications of these tools and propose directions for future research to address ongoing challenges in the field.
Forward citations
Cited by 5 Pith papers
-
MixAssist: An Audio-Language Dataset for Co-Creative AI Assistance in Music Mixing
MixAssist is the first audio-grounded, multi-turn conversational dataset for co-creative music mixing instruction, and fine-tuning Qwen-Audio on it yields human-comparable mixing advice.
-
Finding the noise: Zero-shot AI Music Detection
A zero-shot method based on fakeprints, NMF and a blur-based reconstruction error detects unknown AI-music generators in one-class and clustering setups, working for most services but missing Mubert and pre-v9 Mureka.
-
Echoes: A semantically-aligned music deepfake detection dataset
A semantically aligned, multi-provider music deepfake dataset is harder for detectors and trains models that transfer better than prior AI-music datasets.
-
Double Entendre: Robust Audio-Based AI-Generated Lyrics Detection via Multi-View Fusion
A late-fusion model that combines ASR-transcribed lyrics and speech embeddings detects AI-written lyrics from audio alone, achieving 94.9% recall in-domain and staying robust to attacks.
-
Improved Robustness in AI-Generated Music Detection
Log-frequency remapping plus a single cross-correlation filter makes AI-music artifact detection invariant to speed change by design, matching clean-audio SOTA while recovering the speed factor.
Discussion (0). Sign in to comment.