{"id":"e2d6b6bf-407e-446f-8dd4-a1f56b43dd21","arxiv_id":"2412.04100","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A survey of over one million hours of music datasets and 244 papers finds that Global South music accounts for only 14.6% of training data and is nearly absent from research authorship.","lead":"This paper finds that AI music generation datasets and papers overwhelmingly focus on music from the Global North, with only 14.6% of dataset hours covering Global South genres. It matters because this imbalance shapes what AI can generate, which music gets preserved, and who participates in research.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's 86% dataset-hours figure is computed from a region-tagged subset of about 11.5k hours, not the >1M hours it claims to analyze, so the headline percentage is not supported as a total-hours statistic.","rationale":"The reader's weakest_assumption was annotation quality. I agree that manual labeling is a risk, but the sharper problem is the denominator: the region-level analysis rests on a tiny, likely non-random subset of hours, not the 'over one million hours' claimed. The numbers in Table 1 show this directly: the region rows sum to 11.5k hours while the paper claims >1M total. Even granting perfect labels, a statistic computed on 11.5k hours cannot be reported as 'percent of total dataset hours' unless that subset is shown to be representative or the missing hours are included in the denominator. The paper excludes only 7.9% of datasets (5,772 hours) for missing labels, which is far too small to explain the gap; most hours are simply not region-coded. This makes the headline 86%/14.6% figures fragile. Because the central claim is about magnitude, not just direction, the paper needs to either recompute from a complete per-dataset release or explicitly qualify the percentages as applying to the region-tagged subset. The direction of the finding is probably correct (authorship data are complete and show 229/244 first authors from GN institutions), so a revise-and-resubmit condition is appropriate rather than rejection. This is consistent with the reader's CONDITIONAL verdict; I would keep it unchanged, with the condition sharpened to require denominator transparency and a sensitivity analysis.","tokens_in":12587,"tokens_out":9335,"duration_ms":90567,"concrete_test":"Release the per-dataset annotation table (dataset, total hours, region/genre tags, source of tag) and recompute Table 1b from it. Then perform a two-bound sensitivity analysis: assign all 152 datasets with missing or mixed region tags first entirely to Global North and then entirely to Global South, recalculating the 86% and 14.6% shares. If the region-annotated subset sums to ~11.5k hours while the full corpus is >1M hours, and if the shares move by more than 5 percentage points across the two bounds, the headline should be reworded as conditional on the tagged subset.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central dataset statistic is not robust to the paper's own missing-data rule. The abstract and §2.2 state the survey covers 'over one million hours' from 152 datasets. But the region-wise rows of Table 1 sum to only 11,485 hours (Europe 6,127.92 + East Asia 2,817.73 + America 921.84 + South Asia 588.78 + Central Asia 57.01 + Latin America 332.86 + Oceania 41.99 + Africa 27.50 + Middle East 569.86), and §3.1.2 interprets these as hours ('only 28 hours ... African'). The headline 86% is therefore the share of this 11.5k-hour subset, not of the 'total hours in available datasets.' The 7.9% exclusion of 5,772 hours does not account for the gap; if region tags were only available for a small, non-random slice of the corpus (e.g., region-specific datasets), the 86%/14.6% split can change arbitrarily depending on how the >1M untagged hours are distributed. Annotator noise (no inter-annotator agreement) is secondary; even perfect labels on the current subset cannot support a 'total dataset hours' claim. Additionally, the 93% 'papers focus on Global North music' is actually measured as first-author affiliation (§2.2), conflating authorship geography with research content.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a manual audit of 152 music datasets and 244 papers from eleven AI/music venues, claiming that roughly 86% of dataset hours and 93% of papers focus on Global North music, with Global South genres receiving only 14.6% of data. It also reports correlations between regional representation and digital/population proxies, and discusses implications such as biased evaluation, symbolic-music dominance, and cultural erosion, with mitigation recommendations.","tokens_in":12845,"tokens_out":4095,"duration_ms":41118,"significance":"The topic is timely and important: quantifying geographic and cultural imbalances in AI music generation can inform dataset construction, evaluation, and policy. The authors have assembled a large corpus of papers and datasets, and they make their surveyed paper lists publicly available. If the headline statistics were properly supported, the paper would be a valuable empirical contribution to the discourse on fairness and inclusion in generative AI. However, the central quantitative claims are currently undermined by a mismatch between the stated corpus size and the hours actually analyzed, internal inconsistencies, and lack of validation of the manual annotations.","major_comments":[{"comment":"The headline claim that 'approximately 86% of the total dataset hours' come from the Global North is not supported by the data as presented. The region-wise hours in Table 1(b) sum to 11,485 hours, not the 'over one million hours' attributed to the 152 datasets in §2.2. The paper must either provide region annotations for the full corpus or explicitly rephrase the statistic as applying only to the region-annotated subset, and then justify why that subset is representative of the whole. Without this, the 86% figure cannot be read as a statement about total dataset hours.","section":"§2.2, Table 1(b)"},{"comment":"The paper states that only 7.9% of datasets (5,772 hours) were excluded for lacking explicit genre or region information, yet the annotated regional hours in Table 1(b) amount to only about 1% of the claimed one million hours. This is a large unexplained discrepancy. The authors should clarify how many hours were actually annotated by region and how the remaining hours were treated, since the regional distribution is the empirical basis for RQ1.","section":"§2.2, §3.1.2"},{"comment":"The abstract reports 'over 93% of researchers' and '6.1% of papers' from the Global South, while §3.2 states that the Global North represents 'approximately 88%' and the Global South 'about 12%'. These numbers are mutually inconsistent and must be reconciled; in particular, the 93% figure appears to come from first-author affiliations in Table 1(b), while the 88% figure appears in the prose of §3.2, so the discrepancy is not merely a typographical variation.","section":"Abstract and §3.2"},{"comment":"The claim that '93% of papers focus primarily on music from the Global North' is derived from the first author's institutional affiliation, not from the musical content of the paper. This conflates authorship geography with research focus. The paper should either rename the metric (e.g., 'first-author affiliation share') or provide explicit evidence about the geographic focus of the papers' content; as written, the conclusion about research focus is not supported.","section":"§2.2 and §3.2"},{"comment":"The correlations in §3.3 are computed over only nine regions (n=9) and are reported without significance tests or confidence intervals. For instance, Corr(π_da, α_r)=0.65 has a critical value of approximately 0.666 at the 0.05 level for n=9, so the reported 'strong correlation' is not statistically distinguishable from noise. The authors should report p-values or bootstrap intervals and temper the interpretation of these correlations accordingly.","section":"§3.3"}],"minor_comments":[{"comment":"The manual annotation of datasets and papers has no inter-annotator agreement or validation. Given that the entire regional and genre analysis rests on these labels, the paper should at least report a detailed annotation protocol and ideally a random-sample double-annotation study.","section":"§2.2"},{"comment":"The text contains typos such as 'A vant-garde & Experimental' and 'Highlights significant disparities'; please proofread the manuscript for these and similar errors.","section":"§3.2"},{"comment":"The units for the Duration column are inconsistent with the text: the table lists hours in 10^3, but the text says Pop music forms '200K+ hours' while the table shows 228.26, which would be 228,260 hours. Please align the units and the textual descriptions.","section":"Table 1(a)"},{"comment":"The equation P_s = Σ P_s is tautological and appears to contain a typo; please correct the notation or clarify the intended aggregation.","section":"§2.1"},{"comment":"Several citations are incomplete, including 'See ? ]' in §2.1, '? ]' in §4 regarding fairness in text embeddings, and '? ]' in §4 for feedback loops. Please fill in these references.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a genuinely important issue and has assembled a useful corpus of papers and datasets. However, the central quantitative claim is currently unsupported because the reported 86% figure is computed from a small annotated subset rather than the full corpus, and the paper contains internal numerical inconsistencies. These are fixable in principle—by re-analyzing the full corpus, or by substantially revising the claims to explicitly scope them to the annotated subset—so major revision rather than rejection seems appropriate. I would also encourage the editor to ask for the full annotation data to be released alongside the revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take. The paper identifies a real problem, and the direction of the finding is almost certainly right. But the headline statistic is not supported by the data as presented. The 86% figure for Global North dataset hours is computed from a region-tagged subset that sums to roughly 11.5k hours in Table 1, not from the 'over one million hours' the paper claims to survey. Even after excluding 5,772 hours for missing tags, nearly all of the corpus is unaccounted for in the regional analysis. Without knowing how the tagged subset was selected, the 86% number is an artifact of metadata availability, not a total-hours statistic.\n\nWhat's genuinely useful here is the scale of the audit: 152 dataset papers and 244 generation papers across eleven venues, with genre and region annotations, plus the instrument analysis showing sitar and tabla are under 1% in VGGish and PANN training data. Applying Joshi et al.'s representation metrics to music is a sensible extension, and the GitHub list of surveyed datasets is a practical resource. The recommendations—explicit genre disclosure, avoiding generation when uncertain, investing in inclusive datasets—are reasonable and measured.\n\nThe soft spots are in the measurement, not the motivation. First, the subset problem above is load-bearing; the paper needs to either report percentages only over the annotated subset with a clear coverage statement, or impute missing regional data with a documented procedure. Second, the '93%' claim conflates first-author affiliation with research content. The abstract says 'researchers focus primarily on music from the Global North,' but the actual variable is the first author's institution. Section 3.2 even reports 88% for the same concept, so there's an internal inconsistency. Third, the manual annotation has no inter-annotator agreement, and the East-Asia-as-Global-North taxonomy is contestable; these affect exact magnitudes, though not the overall conclusion. Correlations over nine regions have no confidence intervals, which is minor given the small sample size.\n\nWho is this for? Anyone working on AI music generation or dataset documentation; it's a useful call to action with a curated resource. Does it deserve peer review? Yes, but with the expectation of major revision. If the authors fix the subset statistic, the 88/93 inconsistency, and the authorship/content conflation, this becomes a citable measurement paper. As is, it's a promising draft with a headline number that doesn't yet mean what it claims.","headline":"The paper's central claim is almost certainly right, but its headline 86% number is computed from a ~11.5k-hour subset of a >1M-hour corpus, so the paper as written overstates what the data show.","tokens_in":13407,"tokens_out":5372,"would_cite":false,"duration_ms":47013,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AI music generation is dominated by the Global North in both its training data and its research workforce.","keywords":["AI music generation","Global South","dataset bias","music genre representation","research authorship","generative music models","cultural diversity"],"falsifier":"Take a stratified random sample of the 152 annotated datasets and have two fresh annotators re-label genre and region from the same sources using a published codebook; if inter-annotator agreement is low, say Cohen's kappa below 0.6 on region, then the exact 86%/14.6% split is not stable and should be reported as a range. A second check: recompute all percentages under an alternative taxonomy that moves East Asia into the Global South or weights hours per-file rather than per-dataset; if the Global South share rises above roughly 25%, the phrase 'nearly complete omission' would overstate the case.","tokens_in":12355,"feed_emoji":"🎵","tokens_out":8687,"duration_ms":80706,"temperature":0.7,"pith_summary":"This paper tries to quantify who is represented in AI music generation, both in the data used to train models and in the people who publish the research. After manually annotating 152 datasets totaling more than a million hours of audio and reviewing 244 papers from eleven AI and music venues, the authors report that roughly 86% of dataset hours and 93% of surveyed papers center on music of the Global North, while genres from South Asia, the Middle East, Africa, Latin America, Oceania, and Central Asia together account for only 14.6% of dataset hours. Around 40% of the datasets do contain some non-Western music, but as a share of hours it is tiny, and 51% of papers use symbolic music generation, a paradigm the authors say is ill-suited to microtonal and ornamented traditions. A sympathetic reader would care because if this imbalance is real, AI music systems are being built to serve a narrow slice of the world's music, with consequences for evaluation accuracy, cultural diversity, and the economic survival of Global South traditions.","feed_headline":"Global South gets only 14.6% of AI music training data","feed_subtitle":"A million-hour survey of datasets and 244 papers finds Western dominance in data and authorship, threatening musical diversity.","key_machinery":"The measurement machinery is a two-level annotation and aggregation pipeline. Region and genre representation are defined as fractions of dataset hours ($\\delta_r$, $\\delta_g$) and fractions of papers by style and first-author affiliation ($\\rho_s$, $\\alpha_r$), with a musical style defined as a genre-region pair. Load-bearing is the manual annotation step: for datasets over 10,000 hours the authors mine per-file metadata for genre and region tags, for smaller datasets they rely on the paper's own description, and 7.9% of datasets totaling 5,772 hours are excluded when no tags are available. The paper also constructs comparison proxies—digital availability and listenership from MusicBrainz and SoundCharts, and regional population as an 'ideal' proxy—and correlates these with the representation fractions. The fixed region taxonomy, which places Europe, East Asia, and America in the Global North and everything else in the Global South, determines all headline percentages.","core_discovery":"The paper's central claim is that AI music generation is overwhelmingly a Global North enterprise in both its training data and its research community. Using dataset-hours as the unit of representation, the authors find that European, East Asian, and American music together constitute 85.9% of dataset hours; South Asian and Middle Eastern music sit near 5% each, and African and Central Asian music below 1%. On the publication side, 93% of surveyed papers have first authors affiliated with Global North institutions, and only 6.1% of papers have first authors from South Asia, the Middle East, Oceania, Central Asia, Latin America, or Africa. The authors also document an instrument-level imbalance in the audio embedding backbones used for evaluation: guitar, piano, and drums appear in over half of training clips, while sitar, tabla, accordion, and bagpipes together appear in under 3%. They argue this skew is not neutral—it biases automatic metrics, pushes research toward notation-friendly Western genres, and risks eroding the very traditions it excludes.","pith_inferences":["As an editorial extension, the exact percentages should be treated as order-of-magnitude estimates rather than precise measurements, given the annotation pipeline and the lack of reported coder agreement; the direction of the imbalance would likely survive re-annotation even if the magnitude shifts.","The paper's placement of East Asia in the Global North makes 'Global South' a geopolitical category rather than a purely geographic one; a reader interested in, say, Japanese or Korean traditional music should check whether those are included in the 86% rather than the 14.6%.","A testable extension suggested by the paper's logic is a systematic prompt-based audit: feed text-to-music models a matched set of Global North and Global South genre prompts and measure timbral, microtonal, and rhythmic fidelity against human recordings; the paper offers anecdotal evidence but no such benchmark.","If the correlation results generalize, digital availability as measured by MusicBrainz tracks research publication representation more closely than dataset representation, implying that dataset curation practices, rather than listener demand alone, are the key bottleneck for inclusion."],"forward_implications":["Models trained on the current corpus will tend to produce Western tonal and rhythmic defaults when asked for non-Western styles, since the training data contains so few hours of those genres.","Automatic evaluation metrics built on audio embedding backbones inherit the dataset skew, so reported quality scores for Global South genres will be unreliable and can mislead comparative judgments.","The dominance of symbolic music generation in 51% of surveyed papers steers research infrastructure toward notation-friendly Western genres, leaving microtonal and ornamented traditions without equivalent tools.","Synthetic data augmentation, as illustrated by a dataset that expands MusicCaps from 5k to 37k samples while preserving its skewed genre distribution, can amplify the imbalance instead of correcting it.","The mitigation steps proposed—explicit genre disclosure, refusing to generate for unrepresented genres, community-built inclusive datasets, transfer learning, and genre-specific evaluation—would each require deliberate investment to counteract the documented skew."],"supporting_citations":[{"why":"Central text-to-music system and MusicCaps dataset whose genre distribution anchors the skew analysis.","marker":"[2]"},{"why":"Model cited as defaulting to Western tonal structures on non-Western prompts.","marker":"[10]"},{"why":"Generative system whose outputs the authors test for Global South genre fidelity.","marker":"[12]"},{"why":"Efficient text-to-music system used as an example of training-data bias.","marker":"[30]"},{"why":"MusicBench synthetic dataset that scales up MusicCaps' skewed genre distribution; used to illustrate feedback loops.","marker":"[26]"},{"why":"Methodological template for quantifying global-majority underrepresentation in a research field.","marker":"[19]"},{"why":"Data-statement practice motivating the paper's call for explicit genre disclosure.","marker":"[4]"},{"why":"Fréchet Audio Distance metric whose embedding backbone inherits dataset skew, supporting the evaluation-bias argument.","marker":"[21]"},{"why":"AudioSet-style backbone training corpus used in the instrument-composition analysis.","marker":"[17]"}],"fun_headline_variants":["AI music data: Global South just 14.6%, researchers 93% North","Global South gets 14.6% of AI music data, 7% of papers","AI music's blind spot: 14.6% data, 93% researchers from North","Missing melodies: Global South only 14.6% of AI music data","AI music training data: Global South under 15%, researchers under 7%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole argument depends on the manual genre and region labels assigned to 152 datasets: for large datasets the labels come from metadata mining, for small ones from the papers' own descriptions, and no inter-annotator agreement is reported, so noisy labels or a contested Global North/South taxonomy could change the headline percentages.","fun_headline_variants_meta":{"raw":{"variants":["AI music data: Global South just 14.6%, researchers 93% North","Global South gets 14.6% of AI music data, 7% of papers","AI music's blind spot: 14.6% data, 93% researchers from North","Missing melodies: Global South only 14.6% of AI music data","AI music training data: Global South under 15%, researchers under 7%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000356,"raw_usage":{"total_tokens":1976,"prompt_tokens":1034,"completion_tokens":942,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":650,"completion_tokens_details":{"reasoning_tokens":831}},"tokens_in":650,"tokens_out":942,"duration_ms":8330,"temperature":1.0,"reasoning_tokens":831,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T21:44:13.860794+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a stratified random sample of the 152 annotated datasets and have two fresh annotators re-label genre and region from the same sources using a published codebook; if inter-annotator agreement is low, say Cohen's kappa below 0.6 on region, then the exact 86%/14.6% split is not stable and should be reported as a range. A second check: recompute all percentages under an alternative taxonomy that moves East Asia into the Global South or weights hours per-file rather than per-dataset; if the Global South share rises above roughly 25%, the phrase 'nearly complete omission' would overstate the case.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Efficient text-to-music system used as an example of training-data bias."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Data-statement practice motivating the paper's call for explicit genre disclosure."}],"review_version":1}