{"id":"0685918c-32c1-4dee-8b95-4af520871a98","arxiv_id":"2608.08885","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"An LLM-based nine-dimension scoring protocol applied to 1,259 reggaeton songs shows explicit sexual content rising over 2002-2025 and Spotify's flag catching only about 27% of model-flagged explicit songs.","lead":"This paper uses a large language model to score 1,259 reggaeton songs on nine themes, including separate measures of suggestive and explicit sexual content. It reports that explicit sexual content has risen since 2002 and that Spotify's explicit flag misses most songs the model scores as sexually explicit.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"LLM-score validity rests on a ~10-song calibration and 5-song holdout with no agreement metric; every Section 3 result inherits that assumption.","rationale":"The reader's weakest_assumption identifies the same load-bearing concern: the LLM scores are treated as valid measurements after a very small qualitative calibration. This is indeed the single most load-bearing point because every descriptive and comparative result in Section 3 is a function of those scores. I do not see an internal inconsistency in the deduplication logic, the correlation structure, or the regression calculations; the problem is the external validity of the measurement. The paper is transparent, releases code, prompt, and corpus, and explicitly frames the findings as exploratory, so rejection would be too harsh. However, the evidence for measurement validity is too thin for the numerical claims to be treated as established. The appropriate response is to keep the reader's CONDITIONAL verdict: the paper is a credible proof-of-concept, but its central empirical findings require independent reliability assessment and pinned model versions before they can be taken as measuring the intended constructs. Hence no verdict change is needed.","tokens_in":9668,"tokens_out":4410,"duration_ms":40815,"concrete_test":"Independently annotate a stratified random sample of 100–150 songs from the released corpus with two or more human raters using the published rubric and anchor examples; compute weighted Cohen's kappa or ICC between raters and between each rater and the LLM scores for the two sexual dimensions (and ideally all nine). Also rerun the released prompt on the same corpus with the current gpt-5.1 alias and with a second independent LLM, and recompute the Spotify 26.8% figure and the explicitness trend slope. If human–LLM weighted kappa on sexual explicitness/suggestiveness is below 0.6, or if the 26.8% headline shifts by more than about 5 percentage points across raters or models, the central measurements are not yet established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing assumption is that gpt-5.1's nine-dimension scores are valid measurements of the intended constructs. Section 2.2 supports this with a calibration sample of 'about ten songs' and a held-out test of five songs, with no agreement statistic, no inter-rater reliability, no confidence interval, and no per-dimension error analysis. The prompt is iterated until scores are 'consistently aligned' by the authors' own qualitative judgment; the paper does not report how many iterations, what discrepancies remained, or how close 'aligned' is. All Section 3 results—artist rankings, the rise in sexual explicitness, the flat suggestiveness trend, and the Spotify comparison—inherit this. The Spotify result is especially sensitive: a song counts as explicit whenever explicit_norm > 0, a one-point threshold on a 0–4 ordinal scale, so a systematic tendency of the LLM to output 1 instead of 0 on ambiguous songs would directly manufacture the 456-song disagreement quadrant behind the 26.8% claim. The artist-demeaned trend (0.064 per decade; CI [0.028, 0.130], bootstrapped over 12 artists) also depends on per-song score accuracy. Finally, gpt-5.1 is an unpinned alias (Section 2.4), so even the calibrated prompt cannot be reproduced exactly after a model repoint.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a three-step, open-source method for scoring thematic content in song lyrics using an LLM: (1) collect lyrics and metadata from Spotify and Genius with a manual review pass, (2) score each song on nine ordinal dimensions (0-4) using gpt-5.1 with a prompt calibrated against hand-labeled songs, and (3) normalize, deduplicate, and assemble a composite sexual score. The method is applied to 1,259 reggaeton songs by 12 artists spanning 2002-2025. The empirical analyses describe the corpus, compare artists, trace longitudinal trends, and compare the method's sexual-explicitness score against Spotify's explicit flag. The main findings are that sexual suggestiveness is about twice as prevalent as sexual explicitness, artists differ markedly (composite scores 0.14 to 0.56), explicit sexual content rises over time while suggestiveness stays flat, and Spotify's flag marks only 26.8% of songs the method scores as sexually explicit. The paper releases the data collection code, the scoring prompt, and the corpus.","tokens_in":10029,"tokens_out":4771,"duration_ms":44640,"significance":"If the scoring method is valid, the paper makes a valuable methodological contribution: a reusable, auditable, and adaptable measurement layer for content analysis of lyrics, complementing existing hand-coding and word-frequency approaches. The explicit/suggestive distinction, the artist-demeaned trend analysis, and the external comparison with Spotify's flag are useful and go beyond prior work. The transparency in releasing code, prompt, and corpus, as well as the candid discussion of limitations (model specificity, language bias, unpinned alias), are strengths. However, the significance depends heavily on whether the LLM scores can be trusted as measurements, which the current validation evidence does not adequately establish; the empirical claims, while plausible, inherit this foundational uncertainty.","major_comments":[{"comment":"","section":"Section 2.2"},{"comment":"","section":"Section 3.4 / Figure 10"},{"comment":"","section":"Section 2.4"},{"comment":"","section":"Section 3.3"}],"minor_comments":[{"comment":"","section":"Abstract and Section 3.4"},{"comment":"","section":"Section 2.2"},{"comment":"","section":"Section 3.2"},{"comment":"","section":"Section 3.3"},{"comment":"","section":"Section 3.4"},{"comment":"","section":"Section 2.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is in scope for a computational social science audience, though the physics.soc-ph category is a slight stretch. The central methodological premise is promising, but the validation evidence is currently too weak to support the empirical conclusions. The author should be encouraged to strengthen the validation section and to release the raw model outputs and exact model version. I also note that several cited 2026 preprints may not have undergone peer review; the author should check their current status."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a transparent, well-scoped proof-of-concept for using an LLM as a thematic rater for lyrics, and the release of code, prompt, and corpus makes it a genuinely useful contribution despite a thin validation core.\n\nWhat's new: the nine-dimension protocol with the deliberate split between suggestiveness and explicitness, the small calibration routine, and the 1,259-song reggaeton corpus with a Spotify-flag comparison. The paper doesn't oversell—it repeatedly calls findings exploratory, and it gives a clear account of the method's limits, including model drift and the unpinned alias. That's more honest than most work in this area.\n\nWhere it's soft: the validity of every Section 3 number hangs on a calibration set of about ten songs and a five-song holdout, with no agreement statistic or confidence interval. The prompt is iterated until scores are \"consistently aligned\" by qualitative judgment, which is not a measurement. The stress-test note is right that the Spotify comparison is especially sensitive: a song counts as explicit whenever explicit_norm > 0, so a systematic tendency to return 1 rather than 0 on ambiguous songs would inflate the 26.8% claim. The unpinned gpt-5.1 alias means exact replication isn't guaranteed. These are real problems, but they are problems with a proof-of-concept, not with a finished measurement instrument. The paper says as much in Section 2.4.\n\nThe empirical findings—stability of suggestiveness vs. rise in explicitness, artist heterogeneity, Spotify's low flag rate—are plausible and interesting, but I'd treat them as hypotheses until the scoring layer is validated against a properly sized human-labeled set with inter-rater metrics. The released artifacts make that validation feasible.\n\nBottom line: this deserves a serious referee. It's a well-written, reproducible method paper with honest limitations, and peer review should push the authors to add agreement statistics, a larger holdout, and a pinned model version. I'd bring it to a reading group if the group cares about computational content analysis; otherwise it's a useful cite for anyone building lyric-scoring tools.","headline":"A transparent, well-scoped proof-of-concept for LLM-based lyric scoring with released code, prompt, and corpus; the main soft spot is the thin human-validation layer underneath every score.","tokens_in":10432,"tokens_out":1501,"would_cite":true,"duration_ms":14941,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A reproducible LLM scoring procedure for lyrics, calibrated against human labels, measures nine thematic dimensions at scale; applied to 1,259 reggaeton songs it finds explicit sexual content roughly doubled from 2002 to 2025 while…","keywords":["LLM scoring","song lyrics","reggaeton","sexual explicitness","thematic content analysis","Spotify explicit flag","longitudinal lyrics analysis","measurement method"],"falsifier":"Take a random sample of, say, 100 songs from the released corpus, have two or more independent human raters score them under the paper's dimension definitions, and compare human scores to the model's scores with an agreement measure such as weighted kappa or intraclass correlation. If human–model agreement is no better than chance, or if independent human coding finds that most of the 456 songs Spotify leaves unflagged are not sexually explicit, the central measurement claim fails.","tokens_in":9436,"feed_emoji":"🎤","tokens_out":10844,"duration_ms":94736,"temperature":0.7,"pith_summary":"This paper is trying to turn the familiar claim that reggaeton lyrics are saturated with sexual content into something measurable. It proposes a reproducible method in which a large language model scores each song on nine thematic dimensions on a 5-point ordinal scale, using a scoring prompt that is calibrated against hand-labeled songs before being applied. Applied to 1,259 songs by 12 reggaeton artists released between 2002 and 2025, the method produces three headline findings: sexually suggestive content is nearly twice as prevalent as sexually explicit content; explicit content rises over the two decades while suggestiveness stays flat; and Spotify's explicit flag marks only 26.8% of the songs the method scores as sexually explicit. If the scores are valid measurements, the paper provides a reusable, auditable measurement layer for content analysis of song lyrics, not limited to sexual content.","feed_headline":"Spotify flags only 1 in 4 sexually explicit reggaeton songs","feed_subtitle":"A calibrated LLM scorer applied to 1,259 songs also finds explicit lyrics doubled from 2002 to 2025 while suggestiveness stayed flat.","key_machinery":"The load-bearing mechanism is the scoring prompt, a fixed instruction text given to gpt-5.1 for each song, which defines nine thematic dimensions and anchors each point of a 0–4 ordinal scale with examples. The key conceptual distinction is between sexual suggestiveness (evoking sexuality through metaphor or connotation) and sexual explicitness (naming sexual acts, body parts, or behavior directly); for these two dimensions the model must supply a written justification that is stored for auditing. The prompt is calibrated by hand-labeling about ten songs, comparing the model's scores to those labels, revising the prompt until aligned, and then checking alignment on five held-out songs. Per-song raw scores are rescaled to a 0–1 range, a composite sexual score is defined as the average of the two normalized sexual dimensions, and duplicate songs are deduplicated both within and across artists. This prompt-plus-calibration loop is what carries the claim that the resulting numbers are measurements rather than arbitrary model outputs.","core_discovery":"The paper's central discovery, stated on its own terms, is that a single LLM scoring protocol, calibrated on a small hand-labeled sample, can assign usable ordinal scores for nine thematic dimensions of song lyrics, and that those scores support quantitative answers to questions about magnitude that qualitative studies could only assert. The method deliberately separates sexual suggestiveness from sexual explicitness and finds the two are related but not interchangeable ($r=0.47$): suggestiveness has the highest mean of the nine dimensions (0.45) while explicitness is about half as prevalent (0.27). Over 2002–2025 the fitted sexual-explicitness score roughly doubles, from 0.15 to 0.31 ($r=0.39$, $p=0.065$), and the rise survives restricting the variation to each artist's own catalogue (0.064 per decade; 95% CI $[0.028, 0.130]$), while suggestiveness shows no detectable trend ($r=0.18$, $p=0.40$). The method also finds that of 623 songs scored as containing explicit sexual content, Spotify's flag marks only 167 (26.8%), leaving 456 songs, or 36.2% of the corpus, unflagged.","pith_inferences":["Because Spotify's flag is supplied by rights-holders with no published criteria, the large unflagged gap suggests that content-moderation or age-rating systems built on such flags will under-rate sexual explicitness in Spanish-language catalogs; the paper does not itself test platform-level consequences.","The calibration sample of about ten songs and a five-song held-out check is too small to certify the scores as ground truth; a formal inter-rater study on a larger random sample would be needed before reusing the released scores as a benchmark.","The suggestiveness/explicitness distinction could be imported into longitudinal studies of other lyric corpora, such as English-language pop or hip-hop, to ask whether the documented rise in explicit content comes from direct naming or from innuendo.","Because the paper notes that gpt-5.1 is a model alias that the provider can repoint, exact numerical replication of the released scores requires pinning a model snapshot; the paper's own numbers should be treated as version-specific."],"forward_implications":["If the scores are valid, the rise in reggaeton's sexual content from 2002 to 2025 is concentrated in explicitness, not suggestiveness: the fitted explicit score roughly doubles while suggestiveness remains statistically flat.","Spotify's explicit flag is not a reliable proxy for sexual explicitness in this corpus, since it misses about 73% of the songs the method scores as sexually explicit.","Artists within the genre differ in kind, not just degree: Plan B's profile is dominated by both sexual dimensions, Daddy Yankee's by party/nightlife with low explicitness, and Camilo's by romantic emotion with explicitness nearly absent.","The same calibrated prompt-and-scoring procedure can be redirected to other coding schemes and other lyric corpora or languages, with a new calibration cycle for each context.","Some longitudinal trends are compositional rather than behavioral: romantic emotion and street crime trends do not survive artist-demeaned analysis, while explicitness and substance use do."],"supporting_citations":[{"why":"Supplies gpt-5.1, the model that actually performs the per-song scoring.","marker":"OpenAI, 2025"},{"why":"Precedent for using an LLM to rate music content on multiple ordinal aspects, which the method extends to a calibrated, reusable scoring protocol.","marker":"Zhang et al., 2024"},{"why":"Benchmark showing LLMs can be evaluated zero-shot against human-annotated lyric emotion, supporting the feasibility of LLM-as-rater.","marker":"Dahary et al., 2025"},{"why":"Applies language models to longitudinal lyrics analysis across decades, the precedent for the paper's trend analysis.","marker":"Chandra et al., 2025"},{"why":"Earlier fine-tuned LLM detection of sexually explicit Spanish lyrics, the detection-side counterpart and motivation for streaming rating schemes.","marker":"Zamacola Sánchez de Lamadrid and Garrido-Merchán, 2026"},{"why":"Manual coding of 70 reggaeton songs over time, the small-scale quantitative baseline that the present corpus scales up.","marker":"Arévalo et al., 2018"}],"fun_headline_variants":["LLM: reggaeton explicitness doubled, suggestiveness flat","Spotify misses 73% of explicit reggaeton songs","Explicit reggaeton up 2x since 2002; suggestiveness flat","LLM finds 73% of explicit reggaeton lyrics unflagged by Spotify"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire analysis rests on the assumption that the LLM's numeric scores really measure the intended themes after a calibration step that uses about ten hand-labeled songs and a check on five, with no reported agreement statistic or confidence interval.","fun_headline_variants_meta":{"raw":{"variants":["LLM: reggaeton explicitness doubled, suggestiveness flat","Spotify misses 73% of explicit reggaeton songs","Explicit reggaeton up 2x since 2002; suggestiveness flat","LLM finds 73% of explicit reggaeton lyrics unflagged by Spotify"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001396,"raw_usage":{"total_tokens":5660,"prompt_tokens":972,"completion_tokens":4688,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":588,"completion_tokens_details":{"reasoning_tokens":4608}},"tokens_in":588,"tokens_out":4688,"duration_ms":31337,"temperature":1.0,"reasoning_tokens":4608,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:20:45.726892+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of, say, 100 songs from the released corpus, have two or more independent human raters score them under the paper's dimension definitions, and compare human scores to the model's scores with an agreement measure such as weighted kappa or intraclass correlation. If human–model agreement is no better than chance, or if independent human coding finds that most of the 456 songs Spotify leaves unflagged are not sexually explicit, the central measurement claim fails.","supporting_citations":[{"cited_title":"González, and Thamar Solorio","cited_arxiv_id":null,"evidence_quote":"Precedent for using an LLM to rate music content on multiple ordinal aspects, which the method extends to a calibrated, reusable scoring protocol."}],"review_version":1}