HAIM is a new labeled dataset for granular tracking of AI interventions across music production stages, enabling evaluation beyond binary AI-or-human classification.
Detecting music deepfakes is easy but actually hard
5 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.SD 5verdicts
UNVERDICTED 5representative citing papers
ArtifactNet extracts codec residuals from spectrograms with a 4M-parameter network to detect AI music at F1=0.9829 and 1.49% FPR on unseen tracks from 22 generators, outperforming larger baselines.
Introduces AWM adaptive attack using two-stage optimization and distribution estimation to bypass audio watermark detectors with low detection rates on voice datasets.
Experiments on an open dataset show X-Codec tokens perform best under Udio shift while MERT tokens perform best under Suno-v3.5 shift, indicating token space choice is a key variable for generator-robust detection.
The authors provide the first systematic benchmark of traditional ML, DNN, Transformer, state-space, and multimodal models for machine-generated music detection, augmented with XAI analysis, and report ResNet18 as the strongest performer on in-domain and out-of-domain tests.
citing papers explorer
-
HAIM: Human-AI Music Datasets for AI Music Production Tracking Benchmark
HAIM is a new labeled dataset for granular tracking of AI interventions across music production stages, enabling evaluation beyond binary AI-or-human classification.
-
ArtifactNet: Detecting AI-Generated Music via Forensic Residual Physics
ArtifactNet extracts codec residuals from spectrograms with a 4M-parameter network to detect AI music at F1=0.9829 and 1.49% FPR on unseen tracks from 22 generators, outperforming larger baselines.
-
Learning to Evade: Adaptive Attacks on Audio Watermarking
Introduces AWM adaptive attack using two-stage optimization and distribution estimation to bypass audio watermark detectors with low detection rates on voice datasets.
-
Probing Token Spaces under Generator Shift in AI-Generated Music Detection
Experiments on an open dataset show X-Codec tokens perform best under Udio shift while MERT tokens perform best under Suno-v3.5 shift, indicating token space choice is a key variable for generator-robust detection.
-
Explainable Detection of Machine Generated Music and Early Systematic Evaluation
The authors provide the first systematic benchmark of traditional ML, DNN, Transformer, state-space, and multimodal models for machine-generated music detection, augmented with XAI analysis, and report ResNet18 as the strongest performer on in-domain and out-of-domain tests.