{"id":"5284d46e-46e9-451d-a86b-dc8c79ca1e4c","arxiv_id":"2501.13720","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"LLMs like ChatGPT and Mixtral show a strong Western bias when asked to name top musical contributors and to rate the musical cultures of countries.","lead":"This paper measures whether large language models favor Western music when asked to list top artists or to rate countries' music. It finds a consistent Western skew in both ChatGPT and Mixtral, across several languages.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Rating experiment may measure format compliance rather than stable musical judgment; consistency check needed.","rationale":"The reader's weakest_assumption correctly identified the rating experiment's ill-posedness as the central vulnerability. The paper's abstract makes a two-pronged claim ('strong preference ... in both experiments'), and the rating experiment is the less secure prong. The manuscript itself flags that the rating task is not well-posed, yet no evidence is provided that the numeric outputs have stable, meaningful semantics. The proposed concrete check—either a factual control attribute or a pairwise consistency test—would directly resolve whether the ratings reflect learned musical judgments or are merely format-compliant numbers. Until such validation is provided, a conditional verdict is appropriate; the descriptive Top-100 results are plausible but the combined claim should not be accepted as strongly as stated.","tokens_in":5485,"tokens_out":5118,"duration_ms":47976,"concrete_test":"Run the rating experiment with the same models but replace the music characteristics with a factual, geographically varying attribute (e.g., 'average annual temperature' or 'population density') using the identical prompt structure from Listing 1. Compute Spearman correlation between the model's numeric ratings and real-world values. If the correlation is near zero (ρ < 0.3) for a factual attribute that the model certainly knows, then the numeric rating format does not elicit meaningful relative judgments, and the musical ratings cannot be interpreted as stable learned opinions. Alternatively, for a random sample of 50 country pairs, ask 'Which country has the more complex music?' in a separate forced-choice prompt and measure agreement with the numeric ratings; agreement at chance levels would invalidate the ordinal interpretation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Experiment 2's inference that LLMs rate Western musical cultures higher rests on the assumption that the numeric ratings elicited by Listing 1 are meaningful ordinal judgments about musical culture. The paper itself concedes (Section 6) that 'the whole task of rating musical cultures is not well-posed.' If the model merely produces plausible-sounding numbers to satisfy the format, the observed Western preference in Experiment 2 could be an artifact of prompt compliance (e.g., assigning higher numbers to more familiar countries) rather than a learned cultural judgment. The paper reports no consistency checks: no variance across the three runs, no comparison against pairwise judgments, and no validation that the rating scale has stable semantics. Since the central claim explicitly covers 'both experiments,' the rating experiment's validity is load-bearing. A failure here would reduce the claim to the Top-100 result only.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates whether LLMs exhibit a Western (especially U.S.) bias in music-related judgments. In Experiment 1, the authors prompt ChatGPT-4 and Mixtral-8x7B, through their online interfaces, to generate 'Top 100' lists of musical contributors (bands, singers, instrumentalists, composers, solo artists) in several languages, then map the countries of origin. In Experiment 2, the same models are asked to rate countries on six musical attributes (agreeableness, successfulness, creativity, global influence, tradition, complexity) using a prompt adapted from Manvi et al. (2024). The results are presented as choropleth maps and informal comparisons. The paper concludes that LLMs show a strong preference for Western music cultures in both experiments, with the U.S. dominating, and that this bias persists across models and languages.","tokens_in":5495,"tokens_out":2222,"duration_ms":22090,"significance":"If the central claim is established, this is a useful contribution to the emerging literature on geocultural bias in LLMs, extending prior geographic-bias findings (Manvi et al., 2024) to a previously understudied domain, music. The paper also ships a public repository with data and analysis notebooks, which supports reproducibility and follow-up work. However, the current evidence is descriptive: no statistical tests, confidence intervals, or variance measures are reported, and the rating experiment rests on an assumption that the authors themselves partially disavow in Section 6. The strength of the claim ('strong preference in both experiments') is therefore not yet commensurate with the analysis.","major_comments":[{"comment":"The above is a single complete comment.","section":"Section 4.2, Listing 1, Section 6"},{"comment":"The above is a single complete comment.","section":"Section 4"},{"comment":"The above is a single complete comment.","section":"Section 3.1, Section 6"}],"minor_comments":[{"comment":"The above is a single complete comment.","section":"Section 3.2"},{"comment":"The above is a single complete comment.","section":"Section 3.3"},{"comment":"The above is a single complete comment.","section":"Section 4.2"},{"comment":"The above is a single complete comment.","section":"Appendix Figures 3 and 4"}],"recommendation":"major_revision","confidential_remarks":"The paper is suitable in scope for a CS/CL venue, but the current evidence is too thin for the strength of the claim. The author's own Section 6 admission about the rating task is the key risk; the revision should either add the requested validation or soften the claim to the Top-100 result only. The single-author nature and the self-citation to Kruspe (2024) are not problems, but the author should be encouraged to make use of anonymous links for the repository in a double-blind review process."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a small, honest exploratory paper that gives the music domain a first direct measurement of geocultural bias in LLMs. The Top-100 experiment is its solid core: asking ChatGPT and Mixtral for top bands, singers, composers, and mapping countries of origin produces a consistent, visually striking Western/U.S. skew across models and languages. That is real descriptive evidence, and the paper does not oversell it much.\n\nThe rating experiment is the weak spot, and the stress-test note lands. Asking an LLM to rate \"musical complexity\" of Denmark on a numeric scale relative to all countries is not a well-posed task—the authors admit it in Section 6. Without variance across runs, consistency checks, or any validation that the numbers have stable semantics, the ratings could partly be prompt compliance: plausible numbers assigned to familiar countries. The paper calls the Western preference \"strong in both experiments,\" but the rating experiment alone would not support that. The Top-100 result carries the claim.\n\nWhat the paper does well: it adapts a known methodology (Manvi et al.) to a genuinely new domain, tests multiple languages and two models from different regions, releases code and data, and is unusually explicit about limitations. The citation pattern looks fine; the relevant cultural-bias and music-LLM benchmarks are there.\n\nThe biggest omission is statistical: no error bars, no tests, three runs averaged. For a descriptive pilot that is not disqualifying, but it limits how strongly the results can be stated. The authors should either add basic quantitative support or soften the \"strong preference\" framing and mark the rating experiment as exploratory.\n\nWho gets value: anyone working on AI fairness, particularly geocultural or music-related bias, and people building recommendation or cultural-heritage applications. It deserves a serious referee; the domain is novel enough and the topic important enough that a venue should not desk-reject it, though it needs revision.","headline":"A modest but honest first measurement of geocultural bias in LLM music output; the Top-100 evidence is solid, the rating experiment needs validation before its results carry weight.","tokens_in":6089,"tokens_out":1585,"would_cite":true,"duration_ms":15248,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Large language models show a strong, consistent preference for Western music cultures in both top-artist lists and country-level rating tasks.","keywords":["large language models","cultural bias","geocultural bias","musical ethnocentrism","ChatGPT","Mixtral","prompt evaluation","music culture"],"falsifier":"Run the same two prompts on a model trained predominantly on non-Western web text, for example a Chinese LLM, with the country list translated into Chinese. If the \"Top 100\" lists and country ratings still place U.S. and European music on top and Asia and Africa at the bottom, then the bias is not explained by the regional composition of training data; if the model favors its home region, the paper's training-data explanation gains support.","tokens_in":5191,"feed_emoji":"🎵","tokens_out":5309,"duration_ms":44502,"temperature":0.7,"pith_summary":"This paper asks whether large language models, which are already known to mirror geographic biases, also carry a musical version of that bias. The author prompts two models, ChatGPT-4 and Mixtral-8x7B, to produce lists of the \"Top 100\" performers in several categories and to rate the music of every country on six attributes, with prompts in English, Spanish, Chinese, and French. The reported result is a strong preference for Western music cultures, above all the U.S., in both tasks and across nearly all prompt versions. A sympathetic reader would care because the same models are used to summarize, recommend, and write about music, so an unnoticed Western skew would quietly shape which music cultures users encounter.","feed_headline":"LLMs rank Western music above all others","feed_subtitle":"Top-artist lists and country ratings both skew toward the US and Europe even when prompts are translated.","key_machinery":"The machinery is a pair of prompt-based probes plus a postprocessing pipeline. The \"Top 100\" probe asks the model to enumerate performers in five categories and then attaches countries of origin; the rating probe adapts the prompt protocol of Manvi et al. (2024), asking the model to rate a sampled country relative to all populated locations on Earth on six named musical attributes. Mention frequencies and normalized, run-averaged ratings are then mapped to world maps. The paired probes are meant to catch two different things: open-ended generation reveals which cultures come to mind, while numeric ratings are meant to reveal implicit judgments about them.","core_discovery":"The paper's central claim is that current LLMs display musical ethnocentrism in a specific, measurable way. In the first experiment, asking for top bands, singers, solo artists, instrumentalists, and composers produces country-of-origin distributions concentrated in Western countries, especially the U.S., with Asia and Africa almost absent and South America in between. In the second, asking for numeric ratings of agreeableness, successfulness, musical creativity, global influence, musical tradition, and musical complexity yields the same Western-leaning ordering, with India as the main exception under \"Tradition.\" The author interprets the consistency across the two models and four languages as evidence that the bias comes from the composition and value judgments of shared training data rather than from any single prompt.","pith_inferences":["An implication the paper leaves implicit is that the \"Top 100\" probe likely conflates commercial success or streaming volume with musical importance; a testable extension would compare the model's country distribution against global recorded-music market shares to separate learned consumption statistics from cultural valuation.","The rating probe could be validated by asking models to justify a few ratings in free text before giving numbers, or by rating fictional countries; if scores do not track the justification or track a random-seeming baseline, the numeric scale is measuring prompt compliance rather than a stable belief about music culture.","The same two-prompt design could be transferred to other cultural domains such as cuisine, literature, or cinema; if the Western skew reproduces there, the paper's result would generalize from music to a broader cultural ethnocentrism in LLMs.","Because the author tested only Western-trained models, the decisive comparison would come from a large model trained primarily on Chinese, Arabic, or Hindi web text; the paper's training-data explanation predicts that model would favor its own cultural region, while a \"global elite\" variant of the bias would look different."],"forward_implications":["LLM-generated music rankings, recommendations, and cultural overviews will systematically underrepresent Asian and African music while overrepresenting U.S. and European music.","Changing the prompt language to Spanish, Chinese, or French does not fix the skew, so mitigation cannot rely on localization of prompts alone.","The skew is not specific to one provider: two models trained on different continents still land on the same Western-heavy distribution, suggesting a shared training-data cause.","Because the bias is implicit, it can leak into downstream tasks such as writing assistance and recommendation pipelines, where it is harder to notice than an explicitly biased filter.","For tasks like \"Top 100,\" users may expect the biased answer, which poses a design choice about whether models should mirror human majority taste or offer broader diversity."],"supporting_citations":[{"why":"Supplies the rating-prompt methodology and the prior evidence that LLMs are geographically biased, which this paper extends to music.","marker":"[Manvi et al., 2024]"},{"why":"Provides prior evidence that LLMs favor Western, English-speaking norms, the pattern this paper tests in the music domain.","marker":"[Tao et al., 2024]"},{"why":"Supplies earlier evidence of cultural bias in LLMs, framing the underrepresentation of non-Western cultures.","marker":"[Naous et al., 2023]"},{"why":"Establishes that music-specific bias exists in LLMs and argues for domain-specific evaluation, motivating the experiments.","marker":"[Li et al., 2024c]"},{"why":"Provides the survey framing of how culture is modeled and measured in LLMs, positioning the contribution.","marker":"[Adilazuarda et al., 2024]"}],"fun_headline_variants":["LLMs favor Western music in lists and ratings","ChatGPT and Mixtral bias toward Western music","Musical ethnocentrism: LLMs rank West first","LLM musical bias: West tops all categories"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The rating experiment assumes that asking an LLM to put a number on a country's musical agreeableness, success, creativity, influence, tradition, or complexity extracts a learned cultural judgment rather than a plausible-sounding number with no stable semantics; the paper itself admits the task is not well-posed.","fun_headline_variants_meta":{"raw":{"variants":["LLMs favor Western music in lists and ratings","ChatGPT and Mixtral bias toward Western music","Musical ethnocentrism: LLMs rank West first","LLM musical bias: West tops all categories"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000391,"raw_usage":{"total_tokens":2007,"prompt_tokens":843,"completion_tokens":1164,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":459,"completion_tokens_details":{"reasoning_tokens":1102}},"tokens_in":459,"tokens_out":1164,"duration_ms":8484,"temperature":1.0,"reasoning_tokens":1102,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T15:39:04.704285+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same two prompts on a model trained predominantly on non-Western web text, for example a Chinese LLM, with the country list translated into Chinese. If the \"Top 100\" lists and country ratings still place U.S. and European music on top and Asia and Africa at the bottom, then the bias is not explained by the regional composition of training data; if the model favors its home region, the paper's training-data explanation gains support.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies earlier evidence of cultural bias in LLMs, framing the underrepresentation of non-Western cultures."}],"review_version":1}