{"id":"7a3bf3a3-c3bd-4299-aa51-386adacb53a7","arxiv_id":"2507.10447","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"SONICS, a recent fake-music detector, suffers large accuracy drops under light audio augmentations and fails to generalize to unseen generative models.","lead":"Researchers tested a state-of-the-art fake-music detector called SONICS by applying common audio edits such as pitch shifts, compression, and noise. The detector's accuracy dropped sharply even with small edits that humans barely notice, and it mislabeled songs from generators it had not been trained on.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Claimed -2-semitone Suno flip is not numerically supported: no per-condition statistics, selective Figure 1, n=20 per generator.","rationale":"The reader's weakest_assumption is essentially the same load-bearing concern I identify: the generality and size of the degradation are unestablished because the paper relies on a selective figure and lacks statistical detail. I sharpen this by pointing to the specific missing quantity: the post-augmentation probability for the -2 semitone condition on Suno, which is never reported numerically. Because that one example carries the abstract's central claim, the paper should not be read as having demonstrated the effect with the current text alone. However, the paper is an explicitly preliminary extended abstract, the repository is public, and the direction of the effect is plausible; the issue is verifiability and overstatement, not internal inconsistency. I therefore do not change the reader's CONDITIONAL verdict: the empirical claim is likely right in direction but needs the missing statistical support before it should be accepted at face value. The proposed test is directly executable from the released code and would settle whether the -2 semitone flip is a robust average effect or a selection artifact.","tokens_in":2694,"tokens_out":5860,"duration_ms":75049,"concrete_test":"Use the released GitHub repository to rerun the exact §3.2 augmentation pipeline on the same 20 Suno tracks for pitch shift -2 semitones (plus -1 and -3 as controls). Compute per-track fakeness probabilities, report the mean and 95% bootstrap CI, and run a paired Wilcoxon signed-rank test against the unmodified baseline. If the mean stays near 96% or the CI overlaps the baseline, the central claim fails; if the mean drops below 50% with the CI excluding baseline and the effect is consistent across tracks, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim ('a shift down of two semitones already fools the model into classifying suno as highly real', §3.2) is the quantitative anchor of the paper's headline. The only reported numbers are the baseline probabilities in Table 1 (Suno 96.24 ± 5%). The augmentation result itself is presented as a qualitative description of Figure 1, and the authors state that Figure 1 shows only 'augmentations with the most significant impact.' No post-augmentation mean, per-track scores, confidence interval, or paired test is given for the -2 semitone condition; it is not even stated whether 'classifying suno as highly real' refers to an average over the 20 Suno tracks or to one selected example. Since the figure was chosen after searching a large augmentation/parameter grid, the displayed flip could be a selection artifact rather than a representative light augmentation. With n=20 and a baseline SD of 5 percentage points, a per-condition confidence interval is needed; without it, the central robustness failure is not established. The paper's other qualitative observations (silence pushing predictions toward fake, high-frequency corruption correlating with fake) are similarly unsupported by numerical detail, but the -2 semitone Suno result is the load-bearing example for the abstract's 'light augmentations' claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This late-breaking extended abstract evaluates the robustness of the SONICS fake-music detector under a range of audio augmentations. The authors construct a dataset of 20 real songs and 20 songs from each of Suno, Udio, YuE, and MusicGen, using deepseek-r1 to generate independent prompts. They first measure baseline fake-probability scores (Table 1), then apply augmentations such as pitch shifting, filtering, compression, time stretching, silence insertion, and noise, reporting that the model's predictions degrade under supposedly light modifications. The central claims are that the model generalizes poorly to unseen generators, that a two-semitone downward pitch shift can make Suno tracks appear highly real, that silence pushes predictions toward fake, and that high-frequency corruption correlates with fake predictions. The paper concludes that SONICS is not robust and suggests augmentations during training and explainability as future work.","tokens_in":2956,"tokens_out":2881,"duration_ms":35548,"significance":"If the central claims hold, this is a timely and useful external evaluation of a state-of-the-art music deepfake detector, highlighting concrete failure modes that are perceptually minor for humans. The authors provide a public dataset and code, use independent prompt generation to avoid contamination, and test several generators beyond the detector's training distribution. However, the significance is limited by the small sample size (20 songs per generator), the use of a single detector model, the absence of statistical tests or confidence intervals on the augmentation results, and the selective presentation of only the most impactful augmentations in Figure 1. These issues leave the magnitude and generality of the reported degradation unestablished, despite the rhetorical strength of the abstract's claim that performance 'decreases significantly even with the introduction of light augmentations.'","major_comments":[{"comment":"The central claim that a two-semitone downward pitch shift 'already fools the model into classifying suno as highly real' is not numerically supported. The text refers to Figure 1, but the figure is explicitly said to show only 'augmentations with the most significant impact,' and no per-condition mean, per-track scores, confidence interval, or paired significance test is provided for the -2 semitone condition. With n=20 and a baseline standard deviation of 5 percentage points for Suno, the reported flip could be a selection artifact or an outlier rather than a representative light augmentation. This claim is load-bearing for the abstract's 'light augmentations' statement and should be backed by summary statistics or a scatter plot with paired comparisons.","section":"§3.2, Figure 1"},{"comment":"The generalization results are reported only as means with parenthetical '±' values, with no statement of whether these are standard deviations, standard errors, or ranges, and no statistical tests against chance. For Udio, the mean is 50.51% with a spread of ±43 percentage points, making the baseline statistically indistinguishable from random guessing on 20 tracks. The conclusion that the model 'doesn't generalize well' is therefore not formally established, nor is the claim that only real and Suno songs are classified correctly. Confidence intervals or a paired test (e.g., Wilcoxon against 50%) should be reported for each generator.","section":"§3.1, Table 1"},{"comment":"The parameter ranges that define 'light' augmentations are not stated in the paper; the authors only say the ranges are 'reasonable' and refer the reader to a GitHub repository. Because the central claim concerns the effect of 'light' modifications, the actual values (e.g., pitch-shift step sizes, filter cutoff frequencies, noise levels, bit-crush amounts) must be specified in the text or an appendix. Without this, the reader cannot assess whether the tested augmentations are indeed light or whether the reported effects are driven by severe settings that would be obvious to a human listener.","section":"§3.2"}],"minor_comments":[{"comment":"The caption reads 'Mean Probability of fakeness' and parenthetical values like '6.1 (±4)%' are ambiguous; state explicitly whether the parenthetical is a standard deviation, standard error, or interquartile range.","section":"Table 1"},{"comment":"Several references have formatting issues, such as 'R. Y . et al.' and 'D.-A. et al.'; these should be expanded or corrected for completeness.","section":"References"},{"comment":"The statement that 'an empty file is classified as fake' is intriguing, but it is not explained how a silent file is presented to a model that expects 120-second audio; clarify the preprocessing.","section":"§3.2"},{"comment":"The title and author names contain odd spacing and encoding artifacts; these should be cleaned up for publication.","section":"Title and author block"}],"recommendation":"major_revision","confidential_remarks":"This is an extended abstract rather than a full journal paper, so the bar for completeness is arguably lower. However, the paper's headline quantitative claim (the -2 semitone Suno flip) is currently supported only by a deliberately selective figure and lacks the numerical evidence needed for a reader to assess its reliability. I believe the authors can address this with a small amount of additional analysis, so major_revision rather than reject seems appropriate. The paper would also benefit from at least one significance test or confidence interval for the main augmentation effect. The fit with the journal is reasonable given the ISMIR LBD provenance, but the depth of analysis is closer to a workshop contribution than a full research article."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The useful thing here is the artifact: the authors built a small but real evaluation set across four generators, ran a wide battery of augmentations against a published detector, and released the code and data. That is the right way to probe a state-of-the-art model. The direction of the findings is credible: SONICS generalizes poorly to MusicGen and YuE, Udio sits at chance despite being in the training set, and several mild edits shift predictions. The silence-as-fake and high-frequency-correlation observations are suggestive and worth following up.\n\nThe soft spots are real and mostly about reporting. The central example in the abstract and §3.2—two-semitone pitch shift fooling the model into calling Suno \"highly real\"—is not backed by a number. It is a description of Figure 1, and the figure deliberately shows only the most impactful augmentations. With n=20 per generator and a baseline SD of 5 points for Suno, you need the post-augmentation mean, the spread, and ideally per-track values to know whether that flip is a robust effect or a selected outlier. The same goes for the silence claim: does the probability rise on average, or just in one example? The stress-test note is right on this.\n\nThat said, this is an extended abstract, so the lack of statistical machinery is partly a venue constraint. The authors do not overclaim much in the body: they call it a starting point and acknowledge that some augmentations are already perceptible. The abstract's \"light augmentations\" phrase is a bit strong given that bit crush and white noise at high levels are not light, but they hedge that in the text.\n\nIf this becomes a full paper, the fixes are straightforward: report per-condition means with confidence intervals, show the full augmentation grid rather than a cherry-picked figure, and run a paired test on the pitch-shift condition. I would also like to see more than one model, though that is clearly future work.\n\nOverall, this deserves a serious referee rather than a desk reject: it is a legitimate external evaluation with code and data, and the negative result is important if it holds. My recommendation: take the robustness finding as provisional but worth citing as unpublished work; ask the authors to publish the per-condition numbers before treating the two-semitone flip as established.","headline":"A useful, honest early probe of SONICS robustness with released code and data, but the headline pitch-shift claim is a qualitative reading of a selected figure, not a reported number.","tokens_in":3441,"tokens_out":1714,"would_cite":false,"duration_ms":24345,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A two-semitone pitch shift is enough to make a fake-music detector call Suno real.","keywords":["fake music detection","audio deepfake detection","SONICS","audio augmentations","robustness","generalization","pitch shift","synthetic music"],"falsifier":"Re-run the same SONICS configuration on a larger held-out sample (say, hundreds of songs per generator) with pitch shifts of −2 to +2 semitones and blind human perception ratings for the same shifts. If the model's accuracy drop vanishes or falls within the baseline variance, or if human listeners reliably notice the two-semitone change, the central claim is not supported.","tokens_in":2559,"feed_emoji":"🎵","tokens_out":3969,"duration_ms":44264,"temperature":0.7,"pith_summary":"This paper asks whether a recently proposed fake-music detector, SONICS, still recognizes synthetic music after ordinary audio edits. Testing 20 songs each from real music and four generators, the authors find that light augmentations already cut performance: shifting a Suno track down two semitones makes the model rate it as highly real, and adding silence pushes predictions toward fake. They conclude that the model reacts consistently to modifications, relies on high-frequency spectral cues, and needs more diverse augmentation training. The result matters because deployed detectors will meet edited audio in the wild, and current models may not be fit for that.","feed_headline":"Two-semitone pitch shift fools fake-music detector","feed_subtitle":"Light edits like pitch shifts and silence make the SONICS detector call Suno real, so deployed checks need hardening.","key_machinery":"The load-bearing object is the SONICS SpecTTTra-α detector, the configuration its creators reported as most accurate, applied to 120-second audio samples. The evaluation machinery is a purpose-built dataset of 20 songs per source (real, Suno, Udio, YuE, MusicGen), generated from LLM-written prompts, plus a battery of audio augmentations (aliasing, bit crush, equalization, filtering, masking, mp3/ogg compression, pitch shift, speed changes, silence, reverb, vibrato, white noise) applied across parameter ranges kept within a reasonable range. The model's output is the mean probability of fakeness, and the paper's figure isolates augmentations with the most significant impact. This setup is what turns a robustness suspicion into the claim that light edits fool the detector.","core_discovery":"The central discovery is that SONICS (the SpecTTTra-α variant with 120-second samples), despite high accuracy on unmodified Suno tracks (96.24±5% fake probability), degrades sharply under light audio transformations. A pitch shift down by two semitones is enough for the model to label Suno as highly real, even though a human listener would barely perceive it. Silencing fragments raises the fake probability monotonically, with an empty file classified as fake. The paper also reports weak generalization to unseen generators: MusicGen tracks are mostly classified as real (34.83±31% fake probability) and Udio hovers near chance (50.51±43%). The authors infer that the model latches onto spectral artifacts in high frequencies rather than holistic musical content.","pith_inferences":["The silence finding implies a practical gating mechanism, such as detecting non-silent segments before classification, which the paper does not spell out.","If the model relies on absolute pitch, a simple fix—training with pitch-transposed copies of real music—could be tested; the paper does not propose this.","The lack of confidence intervals means the headline drop may shrink or grow with more data; a larger benchmark would be needed before trusting the effect size.","The same augmentation battery could be applied to other detectors, and comparing their response patterns would show whether pitch fragility is specific to SONICS or common to the approach."],"forward_implications":["SONICS cannot be assumed safe for deployment on user-uploaded or streaming audio, where pitch shifts and silent gaps are common.","Training detectors with a diverse augmentation set, including pitch shifts and pauses, should harden them against this failure mode.","Even an empty audio file is classified as fake, so silence handling (e.g., an ambiguous class or a loss term for silence) is needed before the detector can be used in realistic pipelines.","The model's consistent reaction across sources suggests a shared spectral artifact cue, so explainability work on frequency bands may identify the root cause."],"supporting_citations":[{"why":"Introduces the SONICS detector and its training data (Suno and Udio), the model whose robustness is tested.","marker":"[3]"},{"why":"Sets the research agenda by listing robustness and generalization as key fake-music-detection challenges; this paper targets those two.","marker":"[4]"},{"why":"Supplies YuE-generated test tracks, one of the unseen-generator generalization cases.","marker":"[5]"},{"why":"Supplies MusicGen-generated test tracks, the generator the model misclassifies most.","marker":"[6]"},{"why":"Provides the LLM-based prompt generation used to construct the dataset across all generators.","marker":"[7]"}],"fun_headline_variants":["Pitch shift flips fake-music detector verdict","Two semitones down: AI music detector fooled","Light edits break fake-music AI","Suno slips past detector with subtle pitch shift","Detector blind to slight pitch changes in music"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that 'light augmentations' fool the detector assumes the chosen augmentations and their parameter ranges fairly represent real-world light edits, and that 20 songs per generator and single-run probabilities without error bars are enough to estimate the effect.","fun_headline_variants_meta":{"raw":{"variants":["Pitch shift flips fake-music detector verdict","Two semitones down: AI music detector fooled","Light edits break fake-music AI","Suno slips past detector with subtle pitch shift","Detector blind to slight pitch changes in music"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000358,"raw_usage":{"total_tokens":1866,"prompt_tokens":796,"completion_tokens":1070,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":412,"completion_tokens_details":{"reasoning_tokens":998}},"tokens_in":412,"tokens_out":1070,"duration_ms":8396,"temperature":1.0,"reasoning_tokens":998,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T19:29:56.980704+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same SONICS configuration on a larger held-out sample (say, hundreds of songs per generator) with pitch shifts of −2 to +2 semitones and blind human perception ratings for the same shifts. If the model's accuracy drop vanishes or falls within the baseline variance, or if human listeners reliably notice the two-semitone change, the central claim is not supported.","supporting_citations":[{"cited_title":"Evaluating Fake Music Detection Performance Under Audio Augmentations","cited_arxiv_id":"2507.10447","evidence_quote":"Introduces the SONICS detector and its training data (Suno and Udio), the model whose robustness is tested."},{"cited_title":"This is to be ex- pected, since machine learning models are known for be- ing unable to handle distribution shift between training and test data","cited_arxiv_id":null,"evidence_quote":"Sets the research agenda by listing robustness and generalization as key fake-music-detection challenges; this paper targets those two."}],"review_version":1}