{"id":"ba8443d3-97ac-46d1-96aa-fffc47fdf4da","arxiv_id":"2507.03599","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"MusGO is a community-refined framework with 13 openness categories, applied to 16 music-generative models to produce a public openness leaderboard.","lead":"This preprint introduces MusGO, a 13-category framework for scoring how open music-generative AI models are. It evaluates 16 popular models and publishes a leaderboard meant to expose 'open-washing' and guide responsible development.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The leaderboard ordering is not shown to be stable under alternative weighting or essential/desirable splits, so the central empirical claim lacks sensitivity support.","rationale":"The reader identified survey representativeness as the weakest assumption. I agree that the biased sample is a real limitation, but the more load-bearing issue is that the paper never tests whether its weighting and essential/desirable decisions matter. The essential/desirable split is not uniquely determined by the survey data: Model Card and Datasheet have median relevance 4, the same as categories that became essential, and Evaluation Procedure became essential without any survey relevance measurement. The leaderboard is explicitly used for ordering, so if the ordering changes under reasonable alternative choices, the central empirical claim about significant variation and about training data being most closed becomes an artifact of modeling choices rather than a robust finding. The paper deserves credit for a public repository, transparent criteria, and a community-contribution mechanism, and the claims are framed provisionally in places. However, a sensitivity analysis would settle whether the concern actually lands. This reinforces the reader's CONDITIONAL verdict without moving it further; the contribution is still valuable, but the empirical ordering should be presented as provisional until robustness is demonstrated.","tokens_in":13402,"tokens_out":5474,"duration_ms":68345,"concrete_test":"Using the per-model category scores from the public GitHub repository, recompute O scores under at least four alternative schemes: (1) equal weights across all eight essential categories; (2) weights derived from reported medians with M=5→2, M=4→1, while promoting Model Card and Datasheet to essential; (3) bootstrap resampling of the 110 survey responses (if raw responses are available in the repository) to obtain a 95% confidence interval for weights; and (4) including desirable categories as scored 0/1 in O. Compare each alternative ranking with the published leaderboard using Spearman rank correlation and agreement on the top/bottom quartiles. If ρ < 0.9 under any plausible scheme, the ordering and category-level findings should be reported as provisional pending a stated sensitivity analysis.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's empirical headline—that training data is the most closed category and that the 16 models show significant variation in openness—depends on the O-score construction in §4.3, where E1–E3 are weighted 2 and the other five essential categories 1, and on the essential/desirable classification in §3.4. Neither is demonstrated to be robust. The split is not a direct read-off from Table 1: Model Card and Datasheet both have median relevance 4, the same as Research Paper and Training Procedure, yet the former are classified desirable and the latter essential. Evaluation Procedure, which was added after the survey (§3.3), is classified essential without any survey relevance score. Moreover, the survey sample is acknowledged in §3.2.1 and §5.3 to be biased toward male academics in Europe and North America, and the final framework was further refined through internal MTG discussions (§3.3). If equally defensible weightings or classifications change the ordering, the paper's central claims about variation and the leaderboard are not yet evidence-based. The missing piece is not just a caveat about survey bias, but a quantitative sensitivity analysis showing that the reported ordering is stable under reasonable alternatives.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces MusGO, a community-driven framework for assessing the openness of music-generative AI models. The authors adapt the LLM openness framework of Liesenfeld and Dingemanse (2024) to the music domain, refine it using a survey of 110 MIR community members, and arrive at 13 categories (8 essential, 5 desirable). They apply the framework to 16 music-generative models, compute a weighted openness score (O-score) from the essential categories, and publish a leaderboard plus an open repository with per-model evidence. The central empirical claims are that openness varies significantly across models, that Training procedure is the most open category (11/16 fully open), and that Training data is the most closed category (only 1/16 fully open).","tokens_in":13687,"tokens_out":4544,"duration_ms":54298,"significance":"If the framework and leaderboard are accepted, MusGO would be a useful, reproducible, and publicly inspectable tool for identifying 'open-washing' and for tracking openness in a domain where copyright and IP constraints make openness particularly contested. Strengths of the paper include the open repository, the documented consensus-review process for model assessments, the transparent presentation of survey results, and the explicit adaptation of an existing evidence-based framework rather than inventing categories from scratch. The main weakness is that the quantitative leaderboard ordering rests on weighting and classification decisions that are not shown to be robust; the paper's headline findings therefore need additional sensitivity support before they can be treated as fully evidence-based.","major_comments":[{"comment":"The tie-breaking rule stated in §4.3 (when O-scores are equal, the model with more fulfilled desirable categories is ranked higher) is another arbitrary component of the ordering. The sensitivity analysis requested above should also vary this tie-breaking rule, for example by breaking ties in favor of the model with higher scores in particular essential categories, to confirm that the reported ordering does not hinge on this convention.","section":"§4.3"},{"comment":"The paper acknowledges that the survey sample is biased toward male academics in Europe and North America and notes that this matches typical ISMIR demographics. However, because the category weights and the essential/desirable classification are derived from this sample, the external validity of the leaderboard depends on whether these preferences are representative of the broader MIR community and other stakeholders such as artists and developers. The paper should either perform a subsample robustness check (e.g., recomputing the category relevance and classification after excluding or reweighting regions/genders, if the anonymized response data permit) or explicitly discuss which pairwise orderings in the leaderboard are most fragile under plausible weight shifts. The current discussion in §5.3 treats the bias as a limitation but does not assess its consequences for the ranking.","section":"§3.2.1 and §5.3"},{"comment":"The operationalization of 'fully open' for Training data is relaxed in the framework: a model qualifies as fully open when direct access to training data is restricted by legal concerns, provided that detailed information about all sources is disclosed. This deviates from the survey statement, which was presented as reflecting the fully open level. The relaxation is acknowledged in §5.3, but the paper does not quantify how this choice affects the leaderboard; if a stricter criterion (e.g., requiring actual data access or a closed audit process) were applied, the set of models achieving full openness in Training data—and potentially the overall ordering—could change. The paper should discuss this sensitivity or justify the relaxation more concretely.","section":"§5.1"}],"minor_comments":[{"comment":"The paper states that the final framework was refined through both survey feedback and internal MTG discussions, but it does not itemize which changes came from which source. A short attribution list would strengthen the 'community-driven' claim and make the refinement process more transparent.","section":"§3.3"},{"comment":"Figure 1 (the leaderboard) is referenced in §4.3 but is not reproduced in the text provided; since the leaderboard is a central output, the paper should either include the figure with at least an abbreviated example row or explicitly direct readers to the online leaderboard with a description of the columns and symbols.","section":"Figure 1"},{"comment":"The sentence 'we do not intend to reduce openness to a single value' is somewhat in tension with the use of the O-score to order the leaderboard. Clarifying that the score is an ordering heuristic rather than a measurement would help readers interpret the leaderboard.","section":"§4.3"},{"comment":"When discussing the Foundation Model Transparency Index, the paper criticizes it for not allowing individual data points to be scrutinized. Since MusGO makes its evidence public, it may be worth noting explicitly that the FMTI has since released its data or that the criticism is specifically about the version cited; otherwise the contrast is slightly out of date.","section":"§2.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid, well-documented contribution, and the open repository is a genuine asset. The main gap is the absence of a sensitivity analysis for the weighting and essential/desirable classification, which is load-bearing for the empirical conclusions. I would encourage the editor to ask the authors to add such an analysis and to release the anonymized survey response data (if ethics conditions permit) to enable independent robustness checks. The survey-bias caveat alone is not sufficient; the paper needs to show that the leaderboard ordering is not an artifact of the sample or the chosen aggregation rules."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear [Colleague],\n\nThe one thing you should know: MusGO is a genuinely useful, carefully documented adaptation of Liesenfeld and Dingemanse's openness framework to music-generative AI, with a 110-person community survey and a public, appealable leaderboard of 16 models. It does what it says. But the specific leaderboard ordering is underdetermined by the survey data: the weights and the essential/desirable split involve judgment calls that are plausible but not tested for robustness.\n\nWhat's actually new: the music-specific categories (sonified examples, user-oriented applications, supplementary material page) are sensible and clearly explained. The survey methodology is described well, and the open repository with per-model evidence is a real step up from a one-off scoring table. The paper is honest about the survey's demographic bias and about the relaxation of the training-data criterion for legal reasons. That transparency counts.\n\nThe main soft spots: first, the decisive design choices—double-weighting the three top categories and classifying Model Card and Datasheet as desirable while Research Paper and Training Procedure are essential—are not directly forced by the survey. Evaluation Procedure was added after the survey, and is essential without any relevance score. A simple sensitivity analysis showing that the ordering is stable under plausible alternative weights or splits would address this. Second, the model scoring is consensus-based but not backed by inter-rater reliability measures; it's a qualitative assessment, which is fine, but the exact scores should be taken as provisional.\n\nThat said, the core empirical findings are robust to these concerns. Training data being the most closed category (only 1/16 fully open) and training procedure the most open (11/16 fully open) are stark patterns that won't flip with a different weighting scheme. So the paper's central message stands.\n\nWho should read it: anyone in MIR or AI governance working on openness assessment, open-washing, or the EU AI Act's open-model provisions. It deserves a serious referee. I'd suggest asking the authors to add a robustness check and to justify the essential/desirable split more explicitly, but this is a strong contribution as is.\n\nVerdict: accept for review without hesitation, and I'd bring it to our reading group.","headline":"A well-documented, community-driven openness framework for music-gen AI with a public leaderboard; the ordering is only as solid as its weighting choices, but the main findings are robust.","tokens_in":14114,"tokens_out":2700,"would_cite":true,"duration_ms":32337,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"MusGO is a 13-category, evidence-based framework that scores how open a music-generating AI model really is, and applying it to 16 models shows training data is the most closed component.","keywords":["music-generative AI","openness","transparency","open-washing","community-driven framework","leaderboard","training data","responsible AI"],"falsifier":"Recompute the MusGO leaderboard with equal weights across the eight essential categories, or with weights taken from a new survey of practicing musicians and non-academic developers, and compare the ordering; if models shift positions substantially, the published ranking depends on the double-weighting of source code, training data, and model weights rather than on stable properties of the models.","tokens_in":13190,"feed_emoji":"🎵","tokens_out":8833,"duration_ms":92349,"temperature":0.7,"pith_summary":"The paper sets out to make 'open' a checkable, quantitative property of music-generating AI rather than a label companies attach to their models. It introduces MusGO, a framework that scores a model across 13 categories: 8 essential components such as source code, training data, weights, and licensing, plus 5 desirable extras such as model cards and datasheets. The authors surveyed 110 members of the music information retrieval community to decide which categories matter and how much, then applied the framework to 16 well-known music-generation models and published a publicly updatable leaderboard. The results show training data is the least open category, with only one model fully open, while training procedure is the most open, giving regulators and artists a shared yardstick for spotting 'open-washing'.","feed_headline":"Most 'open' music AI keeps its training data closed","feed_subtitle":"Community-built scorecard rates 16 music generators on 13 categories, making 'open-source' claims checkable.","key_machinery":"The carrying object is the MusGO framework: a checklist of 13 openness categories, split into 8 essential components (each graded closed, partial, or fully open) and 5 desirable components (each binary, present or absent). Two design choices carry the argument: the essential/desirable split, decided from survey relevance scores, and the weighted openness score that doubles the three most relevant categories (source code, training data, model weights) and normalises the result to a 100-point scale for ordering. The framework is operationalised as an evidence-based protocol: each model is scored by one author with written justification, reviewed by two others following consensual qualitative research, and the full evidence set is committed to a public repository so that any score can be inspected and contested. Distinctive to MusGO is the rule that training data counts as fully open when legal restrictions prevent direct release but detailed source information is disclosed.","core_discovery":"The paper's central claim is that openness in music-generative AI is not a binary status but a composite, graded property that can be assessed evidence by evidence through a domain-specific framework. To build that framework, the authors adapt a recently proposed openness methodology for large language models, refine it with feedback from a 110-person survey of the music information retrieval community, and produce MusGO: 13 categories, of which 8 are essential (scored closed, partial, or fully open) and 5 are desirable (binary present or absent). The essential categories are weighted — source code, training data, and model weights count double — and normalised into a 100-point openness score that orders the leaderboard. Applying the framework to 16 state-of-the-art models shows that training data is the most closed category, with only Stable Audio Open fully open, while training procedure is the most open, with 11 of 16 models fully open; it also shows that models releasing model weights tend to provide code documentation and are typically released under open-source or responsible-AI licenses. These assessments, along with their written justifications, are released in a public repository so scores can be scrutinised and appealed.","pith_inferences":["We infer that the essential/desirable split carries a normative claim about what openness should mean in music: reproducibility components are necessary, while documentation extras are optional; testing that claim would require surveying the artists and independent developers the current sample under-represents.","The treatment of IP-restricted training data — rating a model fully open when detailed sources are disclosed but the data itself is not released — creates a possible loophole where a model could score fully open on training data while providing no access to the data at all.","The framework's categories are currently static; as the paper notes, controllability, real-time use, and hardware requirements are emerging concerns it does not yet operationalize, and a natural extension is a hardware-efficiency category whose weight increases for low-resource settings.","The observed correlation between open weights, open code, documentation, and licensing suggests a cluster behaviour: groups that release one key component tend to release several, making targeted pressure to open training data a potentially high-leverage policy point."],"forward_implications":["If MusGO is used as intended, a music-AI release labelled 'open' can be checked against 13 concrete criteria, so incomplete claims such as weights without training-data details become visible and contestable.","The leaderboard can track how individual models change over time, as maintainers add code, datasheets, or licenses in response to community requests.","Because training data and model weights carry double weight, the framework encodes the position that these two components are the core of openness, and that documentation alone cannot compensate for their absence.","The survey-grounded refinement shows that domain-specific adaptation is workable, and the same adaptation template could be applied to other AI domains beyond music."],"supporting_citations":[{"why":"Supplies the base evidence-based, graded framework for assessing openness in large language models that MusGO adapts to music.","marker":"[11]"},{"why":"Establishes the earlier composite graded openness approach and the practice of an open, updatable leaderboard for text generators.","marker":"[29]"},{"why":"The standard community definition of open AI models, used as the reference point for debates about training-data requirements.","marker":"[20]"},{"why":"Defines the datasheet documentation standard that MusGO's Datasheet desirable category checks.","marker":"[16]"},{"why":"Defines the model card documentation standard that MusGO's Model card desirable category checks.","marker":"[18]"},{"why":"Demonstrates that generative models can leak training data, grounding the argument that training-data openness is critical to transparency.","marker":"[7]"},{"why":"A tiered model openness framework whose dimensions align with MusGO, serving as a comparison point for the field.","marker":"[35]"},{"why":"The one evaluated model that achieves fully open status for training data, anchoring the finding that training data is the most closed category.","marker":"[53]"},{"why":"The consensual qualitative research method used to reconcile the three reviewing authors' assessments of each model.","marker":"[59]"}],"fun_headline_variants":["Music AI openness scorecard: training data is the weak link","Training data stays closed even in 'open' music AI models","Community scorecard rates 16 music generators on 13 openness categories","16 music AI models assessed: training data most closed category","MusGO: 13-category framework clarifies openness in music AI"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the survey of 110 participants represents the community's priorities; the paper itself acknowledges the sample skewed male, academic, and European/North American, so if artists, developers, or non-Western stakeholders valued the categories differently, the weights and the leaderboard order built on them would change.","fun_headline_variants_meta":{"raw":{"variants":["Music AI openness scorecard: training data is the weak link","Training data stays closed even in 'open' music AI models","Community scorecard rates 16 music generators on 13 openness categories","16 music AI models assessed: training data most closed category","MusGO: 13-category framework clarifies openness in music AI"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000754,"raw_usage":{"total_tokens":3370,"prompt_tokens":975,"completion_tokens":2395,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":591,"completion_tokens_details":{"reasoning_tokens":2308}},"tokens_in":591,"tokens_out":2395,"duration_ms":19059,"temperature":1.0,"reasoning_tokens":2308,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T20:04:45.961053+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the MusGO leaderboard with equal weights across the eight essential categories, or with weights taken from a new survey of practicing musicians and non-academic developers, and compare the ordering; if models shift positions substantially, the published ranking depends on the double-weighting of source code, training data, and model weights rather than on stable properties of the models.","supporting_citations":[{"cited_title":"Artificial intelligence and music: Open questions of copyright law and engineering praxis,","cited_arxiv_id":null,"evidence_quote":"Supplies the base evidence-based, graded framework for assessing openness in large language models that MusGO adapts to music."},{"cited_title":"Open-source AI must reveal its training data, per new OSI definition,","cited_arxiv_id":null,"evidence_quote":"Establishes the earlier composite graded openness approach and the practice of an open, updatable leaderboard for text generators."},{"cited_title":"Why open-source generative ai models are an ethical way forward for science,","cited_arxiv_id":null,"evidence_quote":"The standard community definition of open AI models, used as the reference point for debates about training-data requirements."},{"cited_title":"Par- ticipants were informed about the scope and purpose of the study, as well as the intended use of the collected data","cited_arxiv_id":null,"evidence_quote":"Demonstrates that generative models can leak training data, grounding the argument that training-data openness is critical to transparency."},{"cited_title":"General-Purpose AI Models in the AI Act - Questions & Answers,","cited_arxiv_id":null,"evidence_quote":"A tiered model openness framework whose dimensions align with MusGO, serving as a comparison point for the field."},{"cited_title":"DITTO-2: Distilled diffusion inference-time t- optimization for music generation,","cited_arxiv_id":null,"evidence_quote":"The consensual qualitative research method used to reconcile the three reviewing authors' assessments of each model."}],"review_version":1}