{"id":"e59a3413-bece-4e9d-a09d-7a5776db1b20","arxiv_id":"2608.06106","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"An audit of Suno and Lyria 3 shows Lyria compresses music within genres while Suno blurs boundaries between genres, and both systems remain easily distinguishable from human-made music.","lead":"This paper tests whether AI music generators make music that is measurably more uniform than human music, using two commercial systems and four genres. It finds two distinct flattening patterns and argues these patterns raise cultural and economic justice concerns as AI music spreads.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central variance/separation ratios compare instrumental AI tracks to human tracks that contain vocals in three of four genres; the paper's instrumental-only control is applied only to the classifier (RQ4), not to RQ1/RQ2, so the claimed Lyria compression and Suno genre-boundary collapse may be…","rationale":"The reader's named weakest assumption is human-corpus representativeness (playlist curation, 30-second previews). My concern is adjacent but distinct: even if the human corpora are perfectly representative, the AI and human conditions differ systematically in vocal content because the audit forces AI outputs to be instrumental while human tracks are not. This is an experimental-design mismatch rather than a sampling issue. The paper already has the tool to address it (instrumental Afrobeats) but uses it only for RQ4. Because the central claim is structural distinctness, the variance/separation ratios must be controlled for vocal content. The Afrobeats number offers partial reassurance, so I do not reject the paper; I keep the reader's conditional verdict and specify the concrete matched-instrumental test that would upgrade or overturn it.","tokens_in":17131,"tokens_out":6723,"duration_ms":50784,"concrete_test":"Build matched instrumental-only human baselines for all four genres (e.g., filter the existing Spotify/Deezer candidate pools with a vocal-activity detector, or collect instrumental versions of the same tracks/artists), then recompute the within-genre variance ratios (Table 4) and genre separation ratios (Table 5) against these baselines. If the Lyria aggregate ratio (0.839) moves materially toward 1.0 and/or the Suno separation ratio (0.442) rises toward the human 0.662, the structural-distinctness claim is not robust to vocal content; if the ratios persist, the concern is resolved. A quick partial check is to report the already-available Afrobeats-only variance ratio and an Afrobeats-only genre-separation diagnostic as matched controls before extending to the other genres.","verdict_should_be":"UNCHANGED","load_bearing_attack":"All AI generations use prompts requesting instrumental, no-vocals audio (Methods, Experiment 1 and Experiment 2), while the human reference corpora are intended to be representative of commercial genres; the paper states that Afrobeats is the only genre where human tracks are also instrumental. For K-pop, Dance Pop, and Heavy Metal, the comparison is therefore instrumental-AI versus vocal-inclusive-human. Vocals contribute large acoustic variance (MFCC envelope, spectral shape, dynamics) and genre-differentiating cues. This can mechanically produce (a) lower within-genre variance for Lyria's instrumental outputs, and (b) reduced genre-centroid separation for Suno's instrumental outputs, independent of any learned homogenization. The paper's only matched control restricts the AI-vs-human classifier to Afrobeats (RQ4); it does not recompute the RQ1 variance ratios or RQ2 separation ratios on a matched instrumental baseline. The Limitations paragraph acknowledges the vocal/instrumental dimension as a confound for 'AI–human separation on timbre-linked features,' but stops short of applying the same control to the homogenization diagnostics that carry the central claim. Since the Afrobeats-only Lyria ratio (0.832) suggests compression can persist in the matched genre, the confound does not necessarily falsify the claim, but its magnitude and the Suno separation-ratio gap are currently unquantified under matched conditions.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper audits two commercial text-to-music systems (Suno v5.5 and Lyria 3) across four genres (Afrobeats, K-pop, Dance Pop, Heavy Metal), generating 100 tracks per system per genre under MIR-derived and minimal prompts, and compares their acoustic properties to 100 human tracks per genre using 72 MIR features and five homogenization diagnostics. The paper reports that Lyria compresses within-genre acoustic variance (aggregate AI/human variance ratio 0.839) while Suno expands it (1.667) but collapses genre-centroid separation (separation ratio 0.429–0.442 vs human 0.662), that the two systems do not converge (RQ3 ratios 2.70–4.64), and that a classifier separates AI from human tracks almost perfectly (AUC 0.991). It then develops a justice-centered interpretation of these findings in terms of recognition, redistribution, and epistemic justice.","tokens_in":17407,"tokens_out":8740,"duration_ms":67798,"significance":"If the empirical claims are supported, this is a valuable black-box audit: it compares two deployed systems against matched human baselines, introduces a minimal-prompt condition to separate system priors from prompt effects, includes an Afrobeats instrumental-only check for the classifier result, and clearly labels the normative conclusions as interpretive. The paper also makes good-faith efforts to audit the representativeness of the human corpora (geographic audit for Afrobeats, Billboard overlap for K-pop). The principal weaknesses are that the central homogenization ratios for RQ1/RQ2 are not tested under matched vocal/instrumental conditions and are reported without uncertainty quantification, which the revision should address.","major_comments":[{"comment":"The AI tracks in both experiments are generated with prompts that explicitly request instrumental, no-vocals audio (Methods, Experiment 1 and 2), whereas the human reference corpora contain vocals in three of four genres (Limitations states this explicitly). The paper's only matched instrumental control is the Afrobeats-only classifier run in RQ4; it does not recompute the within-genre variance ratios (RQ1) or genre-separation ratios (RQ2) on a matched instrumental human baseline. Since vocal content contributes substantially to acoustic variance and to genre-discriminative timbral features, the aggregate Lyria ratio (0.839) and the Suno separation ratio (0.429/0.442 vs 0.662) could partly reflect instrumentality rather than learned homogenization. The Afrobeats results (Lyria ratio 0.832) suggest compression survives matching in that genre, but the magnitude across genres and the Suno separation gap remain unquantified under matched conditions. Please rerun RQ1/RQ2 on instrumental human corpora (or otherwise equate vocal status) and report the resulting ratios.","section":"Methods (Experimental Designs, Human Reference Corpora); Findings (RQ1, RQ2); Limitations"},{"comment":"The central claim that Suno collapses genre boundaries rests on a single separation ratio (0.429 in text vs 0.442 in Table 5; human 0.662) with no confidence interval, bootstrap, permutation test, or any inferential statistic. Likewise, the RQ3 convergence ratios (2.70–4.64) are reported as point estimates without uncertainty quantification. The paper reports permutation/bootstrap procedures for D1 in Table 3, but these are not applied to the RQ2/RQ3 ratios that carry the paper's structural claims. Please provide CIs and significance tests for the separation and convergence ratios, and reconcile the text/table discrepancies.","section":"Findings (RQ2, RQ3) and Appendix Table 5"},{"comment":"The paper states \"All analysis used 30-second excerpts from the middle of the track\" for comparability, but the human reference corpora consist of 30-second Deezer previews (Human Reference Corpora). If the Deezer previews are not the middle 30 seconds, then all AI/human comparisons mix segment-location effects with system effects. Please clarify whether the human previews were aligned to the same segment location, and if not, quantify the sensitivity of the main ratios to preview position.","section":"Methods, Audio standardization; Human Reference Corpora"}],"minor_comments":[{"comment":"The \"null-prompt\" condition is not actually null because the prompts still contain \"Instrumental ... No vocals, no lyrics\"; consider renaming it \"minimal prompt\" or \"genre-only prompt\" to avoid confusion.","section":"Methods, Experimental Designs"},{"comment":"The mixed-effects model that yields Cohen's d = −0.265 is not described; please specify the model formula, random effects structure, and how the variance ratio was computed (e.g., mean over all feature×genre cells).","section":"Findings, RQ1"},{"comment":"The k-means cluster count k=10 is introduced without justification or sensitivity analysis; please report whether the main ratios are robust to reasonable choices of k.","section":"Methods, Human Reference Corpora"},{"comment":"The abstract's claim that Lyria reduces within-genre acoustic diversity is stronger than the data in Table 4, where Dance Pop shows a ratio of 1.119 (Lyria more diverse); please add a qualifier such as \"in three of four genres\" or \"on aggregate.\"","section":"Abstract and Findings, RQ1"},{"comment":"No data or code availability statement is included; making the feature extraction pipeline and the anonymized feature matrices available would substantially strengthen reproducibility.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a timely and important topic and the overall design is sound in many respects. The main obstacle is the vocal/instrumental confound for RQ1/RQ2: the paper's own Limitations section acknowledges the confound for the classifier but does not apply the same logic to the homogenization ratios that carry the central claims. I recommend major revision rather than rejection because the Afrobeats-only results suggest the directional findings may survive matching, but the magnitudes and the Suno separation-ratio gap need to be re-estimated on a matched baseline. The authors should also add inferential uncertainty to the RQ2/RQ3 ratios and fix the internal numerical inconsistencies."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nWhat's new: this is the first comparative audit I've seen that measures acoustic homogenization in two deployed text-to-music systems (Suno v5.5 and Lyria 3) against matched human corpora, and the null-prompt baseline is a genuinely good design choice for isolating system priors from prompt effects. The authors also check prompt fidelity (both systems ignore most instructions), sample human tracks with k-means to avoid playlist bias, and use a separate title-diversity signal. The non-convergence result (AI systems are more distant from each other than random human splits) is likely robust because it compares like with like.\n\nWhat it does well: the operationalization of homogenization is narrow and measurable, the five diagnostics are sensible, and the limitations section is honest about black-box constraints. The Afrobeats-only classifier control was a good instinct.\n\nThe soft spot is the main one. All AI generations are instrumental, while human tracks in three of four genres contain vocals. Vocals inflate both within-genre variance and between-genre separation in the human baseline, so Lyria's compression and Suno's genre-boundary collapse could be mechanical artifacts rather than learned homogenization. The paper acknowledges this confound only for the classifier (RQ4), not for the RQ1/RQ2 ratios that carry the central claim. The Afrobeats data — the one genre with instrumental human tracks — show Lyria's within-genre compression persists (0.832), so that part isn't purely artifactual. But the Suno separation collapse cannot be tested under matched conditions at all, because that requires multiple genres with instrumental baselines. That is a genuine gap in the evidence for the paper's headline distinction. Also, key ratios are reported without confidence intervals, and no code or data are released for reproduction.\n\nNet: this is a serious paper with a serious confound. The central claim is plausible but conditional on a matched instrumental analysis across all four genres, CI reporting, and artifact release. I'd bring it to a reading group — it's a good example of an audit that is careful in most places and shows exactly where the field needs to push.\n\nRecommendation: it deserves peer review, but the referees should treat the vocal/instrumental imbalance as a load-bearing issue and require the matched analysis before the homogenization ratios are accepted.","headline":"A careful audit of two commercial music generators with a genuinely new null-prompt design, but the central homogenization ratios rest on an instrumental-vs-vocal confound that the paper only partially controls.","tokens_in":17923,"tokens_out":4281,"would_cite":true,"duration_ms":33306,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An audit of 1,600 tracks shows Suno and Lyria 3 homogenize music in opposite but measurable ways.","keywords":["AI music homogenization","text-to-music generation","music information retrieval","algorithmic auditing","cultural justice","genre boundaries","Suno","Lyria 3"],"falsifier":"A re-audit using full-length human tracks (or a differently sampled human corpus) that found Lyria's variance ratio near or above 1.0 and Suno's genre separation ratio at or above the human 0.662 would directly contradict the claimed homogenization patterns; likewise, a classifier trained on matched instrumental human tracks across all four genres that dropped below roughly 0.9 AUC would weaken the claim that AI outputs are near-perfectly discriminable from human music on acoustic features alone.","tokens_in":16919,"feed_emoji":"🎵","tokens_out":7574,"duration_ms":72123,"temperature":0.7,"pith_summary":"The paper tries to establish that commercial text-to-music systems produce measurably homogeneous music, and that the homogenization is system-specific rather than a single 'AI sound.' Auditing Suno and Lyria 3 across Afrobeats, K-pop, Dance Pop, and Heavy Metal with 72 audio features, it finds Lyria shrinks within-genre acoustic variation (variance ratio 0.839) while Suno collapses acoustic distance between genres (separation ratio 0.442 vs. human 0.662) without narrowing within-genre spread. A standard classifier separates AI from human tracks almost perfectly (mean AUC 0.991) on these features alone, and the two systems are more distant from each other than random human subsamples (convergence ratios 2.70–4.64). The authors argue these patterns matter for cultural, economic, and epistemic justice because generated music increasingly flows through platforms that reward legible, predictable audio.","feed_headline":"AI music flattens sound two ways: Lyria thins genres, Suno blurs them","feed_subtitle":"An audit of 1,600 tracks finds AI-generated music is acoustically distinct from human music—and homogenizes differently per system.","key_machinery":"The argument is carried by a 72-dimensional MIR feature space spanning rhythm and timing, spectral shape, MFCC timbral envelope, timbre and texture, harmonic content, structure and repetition, and dynamics, plus five complementary diagnostics: global dispersion (track-to-centroid distances), feature variance ratios, entropy ratios, PCA geometric coverage, and separability classification. The load-bearing quantities are three ratios—the aggregate variance ratio (AI/human within-genre variance), the genre separation ratio (between-genre centroid distance divided by within-genre spread), and the system convergence ratio (AI-system centroid distance divided by expected distance between random human splits)—together with the classifier's cross-validated AUC. These operationalize homogenization as reduced acoustic variation in a standardized space rather than as a subjective aesthetic judgment.","core_discovery":"On the paper's own terms, the discovery is that homogenization in AI music is not a single 'AI sound' but two structurally distinct tendencies that can be separated and measured. Using 72 music information retrieval features across four genres, the audit finds that Lyria 3 compresses the acoustic spread within each genre—aggregate AI/human variance ratio of 0.839, with 58% of feature-by-genre cells showing reduced variance—while Suno leaves within-genre spread intact or even enlarged (variance ratio 1.667) but collapses the acoustic distance between genres, lowering the genre separation ratio to 0.442 from the human 0.662. In addition, a random-forest classifier distinguishes AI from human tracks nearly perfectly on the same features (mean AUC 0.991 ± 0.003), with timbral dynamics and rhythmic regularity as the dominant cues, and this separability survives an instrumental-only Afrobeats control, indicating vocals are not the cause. The two systems are also more acoustically distant from each other than two random human subsamples (convergence ratios 2.70–4.64 across genres), so the findings point to learned, system-specific priors rather than a convergent 'AI default' or a prompt artifact.","pith_inferences":["A testable extension would be a longitudinal audit: if human producers begin mimicking AI-typical timbral and rhythmic signatures to stay visible on platforms, the AI–human separation measured here should shrink over time, weakening detector-style watermarking that depends on a stable acoustic gap.","The near-perfect discriminability is measured on features that encode Western production norms; a listener study or a feature set built around microtiming and groove might find that the homogenization is partly an artifact of the measurement space, especially for Afrobeats and K-pop.","The authors' framework implies that the most consequential genre drift will occur in underrepresented genres, since those are learned from smaller and more Western-skewed samples; this could be tested by auditing additional non-Western genres such as Amapiano, reggaeton, or regional rap scenes.","If Deezer's 28% AI-upload figure is representative, the platform-feedback loop described here can be checked by measuring whether recommendation systems disproportionately surface AI tracks and whether that exposure shifts listener genre expectations in controlled experiments."],"forward_implications":["If Lyria's within-genre compression generalizes, listeners streaming AI-heavy playlists of a genre will be exposed to a narrower acoustic band of that genre than human catalogs offer.","If Suno's genre-boundary collapse generalizes, the categorical identities that organize music libraries and recommendation systems—the separations that make Afrobeats and Heavy Metal distinct—become harder to maintain as AI tracks accumulate.","Because both systems' outputs are near-perfectly separable from human tracks on MIR features, acoustic fingerprinting of AI-generated music is feasible in principle, supporting disclosure, watermarking, or filtering—but also adversarial evasion if producers adapt.","The low prompt fidelity observed (Suno max |r|=0.26) implies that a user's attempt to steer generation toward a specific acoustic target will largely fail by default, so homogenization patterns reflect the systems' priors, not user choices.","If AI-generated music is especially legible to the classifiers and playlist algorithms that organize streaming platforms, generated tracks may be preferentially surfaced, creating a feedback loop that shifts genre reference distributions over time."],"supporting_citations":[{"why":"Supplies the feature-based operationalization of musical homogenization as reduced variation in audio descriptors rather than genre labels.","marker":"Bourreau, Moreau, and Wikström 2022"},{"why":"Demonstrates with large Western corpora that pitch and timbral diversity can be measured over time, establishing that acoustic homogenization is empirically testable.","marker":"Serra et al. 2012"},{"why":"Shows stylistic convergence in popular music via harmonic and timbral network features, a second precedent for the measurement approach.","marker":"Mauch et al. 2015"},{"why":"Provides the training-data imbalance statistics (roughly 94% Western, 0.3% African) that motivate the genre choices and the justice stakes.","marker":"Mehta, Chauhan, and Choudhury 2024"},{"why":"Supplies the black-box audit methodology used to evaluate Suno and Lyria 3 without access to model internals.","marker":"Metaxa et al. 2021"},{"why":"Grounds the mechanism: generative models sample from learned distributions optimized to reproduce dominant training patterns, making homogenization a plausible output-level outcome.","marker":"Dhariwal et al. 2020"},{"why":"Justifies the 72-feature MIR set as established extraction methods and flags their Western-norm limitations.","marker":"Peeters et al. 2025"}],"fun_headline_variants":["Two ways AI music flattens: Lyria thins, Suno blurs","AI's two homogenization modes: Lyria thins within, Suno blurs across","Not one AI sound but two: Lyria narrows genres, Suno blends them","AI music has two distinct flattenings: Lyria compresses, Suno merges"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The human reference corpora—100 tracks per genre sampled from Spotify playlists via k-means and analyzed from 30-second Deezer previews—are representative of each genre's true acoustic range; if those playlists or previews omit genre-defining variation, every AI/human comparison is measured against a distorted baseline.","fun_headline_variants_meta":{"raw":{"variants":["Two ways AI music flattens: Lyria thins, Suno blurs","AI's two homogenization modes: Lyria thins within, Suno blurs across","Not one AI sound but two: Lyria narrows genres, Suno blends them","AI music has two distinct flattenings: Lyria compresses, Suno merges"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.002166,"raw_usage":{"total_tokens":8480,"prompt_tokens":1110,"completion_tokens":7370,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":726,"completion_tokens_details":{"reasoning_tokens":7278}},"tokens_in":726,"tokens_out":7370,"duration_ms":37868,"temperature":1.0,"reasoning_tokens":7278,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T14:44:31.353325+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A re-audit using full-length human tracks (or a differently sampled human corpus) that found Lyria's variance ratio near or above 1.0 and Suno's genre separation ratio at or above the human 0.662 would directly contradict the claimed homogenization patterns; likewise, a classifier trained on matched instrumental human tracks across all four genres that dropped below roughly 0.9 AUC would weaken the claim that AI outputs are near-perfectly discriminable from human music on acoustic features alone.","supporting_citations":[{"cited_title":"The Evolution of Popular Music: USA 1960-2010","cited_arxiv_id":"1502.05417","evidence_quote":"Shows stylistic convergence in popular music via harmonic and timbral network features, a second precedent for the measurement approach."},{"cited_title":"S.; Robertson, R","cited_arxiv_id":null,"evidence_quote":"Supplies the black-box audit methodology used to evaluate Suno and Lyria 3 without access to model internals."},{"cited_title":"2025 , eprint =","cited_arxiv_id":null,"evidence_quote":"Justifies the 72-feature MIR set as established extraction methods and flags their Western-norm limitations."}],"review_version":1}