Two CNN detectors that are near-perfect on clean AI music drop to F1 0.19 to 0.47 on real TV broadcast recordings from the new BAMM dataset.
Assessing AI-generated music detection in real-world broadcast monitoring
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
The proliferation of AI-generated music in broadcast media raises concerns about transparency and fair compensation, but reliable detection under real broadcast conditions remains unresolved. Existing studies report substantial performance degradation in this domain, yet their evaluations are limited to synthetic broadcast data. To address this gap, we introduce BAMM (Broadcast AI-Music Monitoring), a 40-hour dataset of real-world television recordings containing AI-generated and human-made music. We compare clean-trained and broadcast-trained CNN variants across three progressively more challenging scenarios: Clean Foreground Music (CFM), Synthetic TV Broadcast (STB), and Real TV Broadcast (RTB). Both models achieve near-perfect performance on CFM but degrade substantially under synthetic broadcast conditions. Broadcast-oriented training improves robustness compared with clean training, although performance remains limited. On RTB, evaluated using BAMM, both models degrade further and show substantial score overlap between AI-generated and human-made music. These results expose a critical domain gap and show that current training approaches on CNN-based detectors remain insufficient for reliable AI-generated music detection in broadcast monitoring.
fields
eess.AS 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Assessing AI-generated music detection in real-world broadcast monitoring
Two CNN detectors that are near-perfect on clean AI music drop to F1 0.19 to 0.47 on real TV broadcast recordings from the new BAMM dataset.