{"id":"0c3e04d4-6f7a-4108-9c1b-80430123c1ce","arxiv_id":"2412.00571","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A review of AI-generated music detection that proposes intrinsic music features and multimodal fusion as the basis for adapting audio deepfake detection methods.","lead":"This paper reviews detection methods for AI-generated music, a younger field than audio deepfake detection. It argues that music-specific features like melody, harmony, and lyrics should guide detection, and sketches a path to adapt speech deepfake detectors to music.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The review's scoping assumption that composition/arrangement carry the detectable essence, while mixing/mastering are irrelevant, is load-bearing and unsupported; if later-stage artifacts carry signal, the 'comprehensive' map misses a key axis.","rationale":"The reader's weakest assumption identifies exactly the load-bearing concern: the review assumes composition and arrangement are the only detection-relevant stages. The paper provides no evidence for this, and the survey's own discussion of artifact-based deepfake detection suggests the opposite. Since the reader's conditional verdict already flags this assumption and requests it be addressed, my read does not move the verdict. The citation defects and '[?]' placeholder are real but secondary; they do not threaten the central claim. The concern about later-stage artifacts, however, is substantive and testable. If the proposed concrete test shows detectors key on production-stage differences, the review's claim to be comprehensive and its recommendation to focus on intrinsic music features would both be weakened. Until that test is run, conditional acceptance with a request to either defend or soften the scoping claim is the appropriate posture.","tokens_in":22676,"tokens_out":1942,"duration_ms":21953,"concrete_test":"Run a controlled detection experiment on the FakeMusicCaps benchmark: take the human MusicCaps recordings and pass them through the same vocoder/mastering/encoding chain used by a current AIGM generator (e.g., the output stage of Suno or MusicLM), and conversely render AI-composed music through a professional human mastering chain. Then train or evaluate a state-of-the-art detector (e.g., SpecTTTra or Afchar's model) on the original AI-vs-human split and test on the cross-rendered sets. If detection accuracy remains high when composition is held constant and only the production chain is swapped, the detector is keying on production-stage artifacts, directly contradicting the review's scoping assumption.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central contribution is framed as the first comprehensive review of AIGM detection, organized around the claim in Section I that composition and arrangement are the creative, detectable stages, while sound design, mixing, and mastering 'primarily serve as aesthetic enhancements rather than altering the essence of the music.' This scoping decision is load-bearing: it determines which detection methods and features the review highlights and which it deprioritizes. However, the claim is asserted without evidence or citation, and it conflicts with evidence cited elsewhere in the paper. In Section III, the authors note that audio deepfake detectors often rely on surface-level artifacts (e.g., Shih et al.'s BEAR framework shows detectors key on noise-like artifacts), and they criticize AIGM detectors for 'dependence on surface-level features.' Such artifacts typically live in the rendering, synthesis, or mastering stages, not in the composition or arrangement. Moreover, modern AIGM systems such as MusicLM or Suno generate audio end-to-end, so the composition, synthesis, and production are entangled; there is no clean separation in which the later stages can be ignored a priori. If detectable traces of AI generation reside primarily in spectral texture, compression, or mastering characteristics, then the review's proposed pathway—anchoring detection on intrinsic music features like melody, harmony, and rhythm—would miss the most reliable signal. This does not invalidate the survey's descriptive content, but it does undercut the 'comprehensive' framing and the prescriptive future direction.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey/overview of AI-generated music (AIGM) detection. It begins with an introduction to music features and their role in generation, then describes existing AIGM detection datasets and detectors, reviews audio deepfake detection methods and multimodal fusion techniques, and proposes a pathway for adapting deepfake audio foundation models to AIGM detection. The paper argues that intrinsic musicological features (melody, harmony, rhythm, lyrics, timbre, etc.) should be the core signals for detection, and it closes with challenges and future directions, including the need for benchmarks, domain-specific models, and explainability.","tokens_in":23017,"tokens_out":4082,"duration_ms":42254,"significance":"If the paper's claims hold, it would provide a useful early map of a small but emerging field, and its proposed focus on musicological features is a distinctive position that could guide subsequent research. The paper also gathers the two known AIGM-specific datasets (FakeMusicCaps and SONICS) and connects them to the broader audio deepfake literature, which is valuable for researchers entering the area. However, the paper's scholarly value depends on accurate citations and on the defensibility of its scoping decision to prioritize composition and arrangement while downplaying later production stages; both of these currently need attention.","major_comments":[{"comment":"The claim that composition and arrangement carry 'the most creativity and foundational stages' while sound design, mixing, and mastering 'primarily serve as aesthetic enhancements rather than altering the essence of the music' is load-bearing for the review's scope and for the Section V recommendation that 'intrinsic features unique to music are essential and should be prioritised as core detection features.' No evidence or citation is provided for this empirical claim, and it stands in tension with the paper's own discussion in Section III.B: the BEAR framework of Shih et al. shows that audio deepfake detectors can key on surface-level, noise-like artifacts, and the paper itself criticizes AIGM detectors for 'dependence on surface-level features.' Such artifacts can plausibly be introduced in rendering, synthesis, or mastering stages, and modern end-to-end generators such as MusicLM or Suno do not separate composition from production. The authors should either support the scoping claim with evidence or explicitly expand the review to include artifact-based and production-stage detection signals; without this, the proposed pathway rests on an unsupported assumption.","section":"Section I"},{"comment":"The reference list contains several defects that undermine the verifiability expected of a review: Suno AI is cited with a placeholder '[?]' in the 'Commercial tools' paragraph of Section II; reference [73] is labeled 'Image transformers' with an author list that does not match arXiv:2010.11929, which is in fact the same paper as reference [104] ('An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale'); and the SongComposer entry appears twice as references [40] and [59]. In addition, references [110] and [114] duplicate the same AlexNet paper. For a survey whose contribution is explicitly to provide a reliable map of the literature, these errors are not merely cosmetic; they need to be corrected before publication.","section":"Section II; References"}],"minor_comments":[{"comment":"There is a typo in 'catagorised' (should be 'categorised') and the phrase 'thought it is a bit out-dated' is grammatically awkward and should be revised.","section":"Section II"},{"comment":"The figure caption contains scattered music symbols and a somewhat unstructured sentence; a cleaner description of the five production steps would improve readability.","section":"Section I, Figure 1"},{"comment":"The paper uses 'Afchar [60]' both for the proposed dataset and for the detector described in the same section; please disambiguate the dataset reference from the detection method reference.","section":"Section III.A and Table I"},{"comment":"The dataset name 'sound8k' appears with inconsistent capitalization (elsewhere as 'Sound8K'); please standardize the naming.","section":"Section III.B"},{"comment":"The column heading 'Compared Baseline' is unclear: some entries list model names (e.g., RawNet2, ResNet) rather than a baseline comparison, and the column should be more precisely titled or supplemented with a description in the table notes.","section":"Table II"}],"recommendation":"major_revision","confidential_remarks":"The manuscript addresses a timely and underserved topic, and the authors' broad knowledge of the audio deepfake literature is evident. However, the citation defects (placeholder reference, garbled reference [73], duplicated references) are surprising for a survey and should be corrected in a careful revision. In addition, the scoping assumption about which production stages carry detectable information is asserted rather than argued; I recommend the authors either supply supporting evidence or broaden the review's stated scope. The 'first comprehensive review' claim should also be rechecked against a systematic literature search before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, this is genuinely the first survey that centers AI-generated music detection as its own task, and it does a decent job of organizing what little exists: two dedicated datasets, a handful of detectors, and a transfer map from audio deepfake detection. Second, the paper's central bet—that detection should anchor on intrinsic music features (melody, harmony, rhythm, lyrics) because mixing and mastering are 'aesthetic enhancements'—is asserted rather than argued, and it sits in tension with its own review of deepfake detectors that key on surface artifacts. Keep that in mind before adopting the proposed pathway.\n\nWhat it does well: the survey is timely and useful for anyone entering the field. The tables comparing datasets and deepfake detection models are practical. The section on music features is a reasonable primer. The distinction between audio deepfake detection and AIGM detection—especially the multimodality introduced by lyrics—is worth making. The proposed transfer learning direction from deepfake audio foundation models is plausible and not overcooked.\n\nSoft spots. The citation quality is below what a review needs. There's a literal '[?]' for Suno AI in Section II. Reference [73] lists a phantom author list for 'Image transformers' and does not match the linked arXiv paper; the correct ViT citation appears later as [104]. Reference [106] for wav2vec2 has a garbled author string. These are fixable, but they erode confidence in the reference list as an entry point. The bigger conceptual issue is the load-bearing assumption about composition/arrangement being the 'essence' while production stages are mere polish. Modern AIGM like MusicLM or Suno generates audio end-to-end; artifacts may live in spectral texture, mastering, or synthesis. The paper even notes that deepfake detectors often exploit noise-like artifacts (BEAR) and criticizes 'surface-level features'—so the line between intrinsic and surface is not clean. This doesn't sink the survey's descriptive content, but the 'comprehensive' framing and the prescriptive pathway lean on a claim that needs evidence or at least explicit acknowledgment as a working assumption.\n\nVerdict: conditional acceptance. Fix the references, soften the claim in Section I, and either justify the composition/arrangement focus or present it as a testable hypothesis rather than a settled fact. For whom: graduate students and researchers looking for a map of AIGM detection. It deserves a serious referee.\n\nRecommendation: send to peer review with a request for revision.","headline":"Genuinely the first map of AI-generated music detection, useful for newcomers, but the composition/arrangement 'essence' assumption is unproven and the reference list needs cleanup before it can serve as a trusted entry point.","tokens_in":23451,"tokens_out":2590,"would_cite":true,"duration_ms":25042,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This review claims to be the first comprehensive overview of AI-generated music detection, arguing that detectors should be built on intrinsic musical features—melody, harmony, rhythm, and lyrics—rather than surface artifacts like…","keywords":["AI-generated music detection","audio deepfake detection","intrinsic music features","musicology","multimodal detection","music datasets","foundation model transfer","music production pipeline"],"falsifier":"Take songs composed and arranged by humans, pass only the synthesis, mixing, and mastering stages through AI tools, then test whether detectors built on intrinsic musical features can still separate them from fully human productions; if artifact-focused detectors succeed while intrinsic-feature detectors fail, the paper's central recommendation would be refuted.","tokens_in":22441,"feed_emoji":"🎵","tokens_out":8609,"duration_ms":78404,"temperature":0.7,"pith_summary":"As AI music tools spread, the music industry, copyright holders, and audiences need ways to tell whether a song was made by a human; this paper argues that such detection is a distinct, underexplored problem that deserves its own methods. The authors claim no previous review specifically covers AI-generated music (AIGM) detection and present what they call the first comprehensive review of existing approaches. Their central recommendation is that detectors should be anchored in intrinsic music features—melody, harmony, rhythm, lyrics, timbre, genre, and emotion—because surface cues such as watermarks can be bypassed by generators. The paper also proposes a pathway for adapting speech-oriented audio deepfake detectors to music, while cautioning that music's structure and subjectivity require domain-specific design. If this framing is right, the field's priorities should shift toward musicological feature understanding, multimodal audio-and-lyrics modeling, and benchmarks built on dedicated AIGM datasets.","feed_headline":"AI-generated music detection gets its first comprehensive review","feed_subtitle":"Review argues melody, harmony, and rhythm should anchor detectors instead of surface artifacts.","key_machinery":"The load-bearing object is the five-stage music production pipeline—composition, arrangement, sound design, mixing, and mastering—with the first two stages treated as the creative core that defines a song's essence. Around this the review builds a feature taxonomy: content-based features (melody, harmony, rhythm, lyrics) and decoration-based features (timbre, instrument, emotion tone, genre), matched against a detection taxonomy of end-to-end versus feature-based classifiers. This machinery organizes the survey: generation models condition on these musical features, detectors are judged by whether they capture them, and datasets are evaluated by whether they include them, making musicological feature understanding the proposed anchor for transferring foundation models from audio deepfake detection to AIGM detection.","core_discovery":"The paper's central claim is that AIGM detection is an emerging but neglected field, and that its first comprehensive review shows a workable direction: intrinsic musical features carry the detectable signature of AI authorship, while surface-level artifacts such as watermarks do not. It surveys the two dedicated AIGM datasets (FakeMusicCaps, an audio-only collection of human and AI-generated clips, and SONICS, a multimodal audio-and-lyrics set), along with the few existing detectors such as SpecTTTra, a transformer that tokenizes spectro-temporal features, and a convolutional music deepfake detector. The authors ground this in a five-stage music production model in which composition and arrangement establish the creative core, so detectors should classify on content-based features and decoration-based features rather than on post-production polish. They conclude that direct transfer from speech deepfake detection is insufficient and that future detectors need music-specific, multimodal, explainable designs.","pith_inferences":["A testable extension is the review's composition-and-arrangement focus: if AI artifacts concentrate in mixing and mastering, detectors built on intrinsic features will miss them, and artifact-focused methods would win in that scenario.","The paper leaves open whether self-supervised speech representations or music-specific pretraining should anchor the transfer pathway; a likely next experiment is to compare both on the same AIGM benchmark.","The adversarial relationship between generation and detection implies that once detectors learn intrinsic-feature signatures, generators will be optimized against those signatures, so benchmarks will need periodic refreshment to stay meaningful.","A concrete editorial extension is that the survey's gap analysis suggests a shared out-of-domain evaluation protocol, where detectors are tested on unseen generators and unseen production styles, would serve the field better than any single accuracy number."],"forward_implications":["If intrinsic musical features are the right signal, future AIGM detectors should be benchmarked on whether they capture melody, harmony, rhythm, and lyrics, not just on in-domain accuracy against a single generator.","Audio deepfake detection models can serve as starting points for AIGM detection, but the paper argues direct transfer or fine-tuning is insufficient because music carries musicological structure and subjective qualities that speech audio lacks.","Because lyrics are a separate text modality that shapes a song's meaning and emotion, the paper concludes that multimodal audio-and-lyrics detectors are necessary for accurate AIGM detection.","The scarcity of dedicated datasets—only FakeMusicCaps and SONICS exist—means that building comprehensive, accessible benchmarks is a prerequisite for progress, the paper states.","Surface-level cues such as watermarking are unreliable because generators can avoid them, so the paper prioritizes intrinsic features as the core detection signal."],"supporting_citations":[{"why":"It supplies the five-step music production pipeline that grounds the paper's focus on composition and arrangement.","marker":"[19]"},{"why":"It provides the musicological feature extraction example that anchors the proposed intrinsic-feature direction.","marker":"[22]"},{"why":"It describes a music deepfake detector showing high in-domain accuracy but degraded out-of-domain performance, a central baseline for AIGM detection.","marker":"[60]"},{"why":"It introduces FakeMusicCaps, one of the two dedicated AIGM detection datasets the review is built around.","marker":"[62]"},{"why":"It introduces SONICS and the SpecTTTra detector, the multimodal AIGM resource the review analyzes.","marker":"[63]"},{"why":"It is an existing audio deepfake detection survey the paper positions as leaving AIGM detection underexplored.","marker":"[8]"},{"why":"It documents the generalization failures of deepfake detectors that motivate the paper's call for robust, feature-based detection.","marker":"[117]"},{"why":"It demonstrates transfer learning with audio spectrogram transformers, cited as evidence that deepfake detection techniques can move toward music.","marker":"[137]"}],"fun_headline_variants":["Music AI detection needs intrinsic features, not watermarks","First comprehensive review of AI-generated music detection","Detecting AI music: focus on melody, harmony, rhythm","From deepfake audio to AI music: a detection roadmap","AI music detection: why surface artifacts fail"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The review assumes that what makes music creative and detectable resides in the composition and arrangement stages, with sound design, mixing, and mastering acting only as aesthetic polish, so if AI artifacts enter mainly during those later production stages the proposed intrinsic-feature focus would miss the signal.","fun_headline_variants_meta":{"raw":{"variants":["Music AI detection needs intrinsic features, not watermarks","First comprehensive review of AI-generated music detection","Detecting AI music: focus on melody, harmony, rhythm","From deepfake audio to AI music: a detection roadmap","AI music detection: why surface artifacts fail"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001185,"raw_usage":{"total_tokens":4866,"prompt_tokens":894,"completion_tokens":3972,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":510,"completion_tokens_details":{"reasoning_tokens":3897}},"tokens_in":510,"tokens_out":3972,"duration_ms":29791,"temperature":1.0,"reasoning_tokens":3897,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T05:11:41.888710+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take songs composed and arranged by humans, pass only the synthesis, mixing, and mastering stages through AI tools, then test whether detectors built on intrinsic musical features can still separate them from fully human productions; if artifact-focused detectors succeed while intrinsic-feature detectors fail, the paper's central recommendation would be refuted.","supporting_citations":[{"cited_title":"Parameter-Efficient Transfer Learning of Audio Spectrogram Transformers","cited_arxiv_id":"2312.03694","evidence_quote":"It demonstrates transfer learning with audio spectrogram transformers, cited as evidence that deepfake detection techniques can move toward music."}],"review_version":1}