{"id":"00fe1d4b-a032-4ea8-b507-b3ced9b96a44","arxiv_id":"2411.11692","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Traditional n-gram captioning metrics are oversensitive to syntax and correlate weakly with human judgments on music captions.","lead":"This paper tests whether standard text metrics like BLEU and ROUGE agree with human opinions of music captions, and finds they do not. The authors suggest the field needs new ways to judge music captions that capture meaning rather than wording.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The correlation analysis in Table 2 cannot support the central claim that metrics 'lack alignment' with human judgments: only 20 entries, no confidence intervals or significance tests, and the observed r≈0.08 is statistically indistinguishable from zero or from a moderate correlation.","rationale":"The reader's weakest_assumption correctly flags the 30-vs-20 discrepancy and the absence of significance tests; my concern sharpens this into the single most load-bearing issue: the statistical sufficiency of the correlation evidence. If the correlation analysis is unreliable, the headline quantitative conclusion collapses, because the paper's illustrative examples of syntactic brittleness (Figure 1, Table 1) are only anecdotal and the abstract's claim about 'correlat[ing] well with human judgments' depends on Table 2. However, this is a fixable empirical weakness rather than a fundamental flaw in the central thesis that n-gram metrics are brittle to paraphrase—that thesis is plausible and partially supported by the examples. Thus the appropriate verdict remains CONDITIONAL, not REJECT. I kept the verdict unchanged because the concern, while real, does not move the verdict away from what the reader already concluded; it reinforces the need for the stated revisions.","tokens_in":3245,"tokens_out":4103,"duration_ms":41034,"concrete_test":"Ask the authors to release per-clip MOS and per-metric scores for all 30 sampled clips and all three caption types, then recompute: (a) Pearson r with 95% bootstrap confidence intervals using all 30 clips rather than 20; (b) a mixed-effects model with random intercepts for audio clip to account for the three non-independent captions; and (c) an equivalence test (e.g., TOST with bounds ±0.2 on Fisher-transformed r) to determine whether the true correlation is practically zero. If the upper confidence bound exceeds 0.2 or the equivalence test fails, the claim that metrics 'lack alignment' with human evaluations would need to be softened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's quantitative claim that BLEU, METEOR, ROUGE, SPICE, and FENSE do not correlate with human MOS rests entirely on Table 2, which reports Pearson r values near 0.08 for '20 entries' even though Section 2 says 30 audio clips were sampled. No confidence intervals, p-values, or power analysis are provided. With n=20, the 95% confidence interval for r=0.08 spans roughly [−0.38, 0.51], so the data are consistent with correlations ranging from moderately negative to moderately positive. The conclusion 'lack alignment' requires an equivalence test or an explicit upper bound on the plausible correlation, not merely a point estimate near zero. Additionally, the correlation is likely computed across non-independent observations (three captions per audio clip), and the set includes paraphrased captions deliberately constructed to have low n-gram overlap while receiving similar MOS; this range restriction and clustering can artificially deflate Pearson r. Without a per-clip analysis or a mixed-effects model, the central numerical claim is not established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"Do Captioning Metrics Reflect Music Semantic Alignment? (Lee and Lee, ISMIR 2024 late-breaking abstract) investigates whether standard captioning metrics (BLEU, METEOR, ROUGE, SPICE, FENSE) reflect human semantic judgments of music captions. The authors sampled 30 clips from the MusicCaps evaluation set, generated three caption types per clip (original ground truth, LP-MusicCaps inference, and author-authored paraphrases), and collected MOS ratings from 50 MTurk participants. They report metric scores for original, paraphrased, and distorted captions (Table 1), a bar plot of average scores (Figure 1), and Pearson correlations between each metric and MOS for 20 entries (Table 2). The paper concludes that n-gram metrics are overly sensitive to syntactic variation and lack alignment with human evaluations.","tokens_in":3574,"tokens_out":4435,"duration_ms":38358,"significance":"The question is timely: music captioning is growing, and the field needs to know whether borrowed NLP metrics are valid. The paper provides concrete examples (Table 1) that nicely illustrate how paraphrases are penalized by n-gram overlap, and it attempts a human-subject correlation study. If the correlation results were statistically well-grounded, the paper would make a valuable contribution to the evaluation methodology discussion. As it stands, the evidentiary basis for the central quantitative claim is too weak, so the paper functions more as a motivating demonstration than as a demonstrated finding.","major_comments":[{"comment":"The correlation analysis reports Pearson r values for '20 entries' although Section 2 states that 30 audio clips were sampled. The manuscript does not explain the discrepancy (e.g., exclusions, missing ratings, or data filtering). With n=20, the 95% confidence interval for r=0.08 spans roughly [-0.38, 0.51], so the data are statistically indistinguishable from zero and also from moderately positive or negative correlations. The statement that metrics 'do not show significant correlations' is not supported without p-values, confidence intervals, or an equivalence test that bounds the plausible correlation. This is the central numerical evidence for the 'lack alignment' claim, so it is load-bearing.","section":"Section 3.2, Table 2"},{"comment":"The distorted captions are labeled 'Semantic X, Syntactic O' in Table 1, but no human MOS values are provided for them, and no validation is reported that the distortions are semantically incorrect while the paraphrases are semantically equivalent. The claim that 'most n-gram-based metrics tend to favor distorted captions over paraphrased captions' depends on these unverified labels; without a human ground-truth check on the distorted captions, the comparison cannot be interpreted as evidence about semantic alignment. This also weakens the syntactic/semantic distinction that motivates the paper.","section":"Section 3.1, Table 1"},{"comment":"The text says that evaluation metrics 'show a significant decrease' when comparing original to paraphrased captions, but no statistical test, confidence interval, or effect size is reported. Figure 1 appears to show only mean values, with no error bars or per-clip data. A paired significance test (e.g., Wilcoxon signed-rank test on per-caption metric differences) is needed before concluding that the metrics are systematically sensitive to paraphrasing.","section":"Section 3.1, Figure 1"},{"comment":"The correlation is computed over pooled observations that are not independent: each audio clip contributes multiple caption types (original, inference, paraphrased), and the paraphrased captions were deliberately constructed to have low n-gram overlap while receiving high human ratings. This clustering and range restriction can attenuate Pearson r. The authors should report per-clip correlations, a mixed-effects model, or at least a correlation that accounts for the repeated-measures structure. Without this, the numerical r-values do not establish that the metrics 'lack alignment' with human judgments.","section":"Section 3.2, Table 2"}],"minor_comments":[{"comment":"Report p-values or confidence intervals, and specify the exact number of observations and why it differs from the 30 clips in Section 2.","section":"Section 3.2, Table 2"},{"comment":"Add error bars and indicate the number of captions per bar; also annotate which differences are statistically significant.","section":"Figure 1"},{"comment":"Provide details on participant recruitment and filtering, how many ratings each caption received, and the aggregation method for MOS (e.g., mean vs. median).","section":"Section 2"},{"comment":"The conclusion 'we demonstrate existing metrics are overly sensitive to syntactic variations' should be softened to 'provide evidence consistent with' given the small sample and the lack of statistical testing.","section":"Conclusion"},{"comment":"Reference [10] is the MusicLM paper; consider citing the MusicCaps dataset (or its specific description) directly for the dataset details.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is an extended abstract, so I have not held it to the standards of a full paper. However, the central quantitative claim is presented in a way that is not yet substantiated; the authors should either add the missing statistical analysis or reframe the contribution as a position/demonstration. The overlap between the paper's authors and the LP-MusicCaps model is not itself a concern, but the inference captions' provenance should be transparent. I would be willing to consider a revised version that addresses the statistical and validation issues."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is an honest little extended abstract that points at a real problem: n-gram overlap metrics are brittle to paraphrasing, and music captioning researchers should not pick models on BLEU alone. The illustrative examples in Table 1 and Figure 1 make that point clearly, and the side observation that FENSE is less sensitive to paraphrase than BLEU, METEOR, or ROUGE is worth taking seriously.\n\nThe soft spot is the quantitative core. The paper claims these metrics 'lack alignment' with human judgments, but the support is one correlation table with 20 entries while the method says 30 clips were sampled. The r values hover near 0.08, with no confidence intervals, no p-values, and no power analysis. With n=20, r=0.08 is consistent with anything from moderately negative to moderately positive correlation, so the data do not show 'lack alignment'—they just do not show much of anything. To make the claim you need either an equivalence test or a pre-specified upper bound on the plausible correlation. There is also a clustering problem: three captions per clip, including paraphrased captions deliberately built to have low n-gram overlap, which can artificially deflate Pearson r. A per-clip analysis or mixed-effects model would be the minimum.\n\nI also want to flag the distorted captions. They get no human MOS, yet the text says the metrics 'tend to favor distorted captions over paraphrased captions.' Without human ratings for the distorted items, the semantic labels are just the authors' assertion. That may be right, but it is not evidence as reported. And there is no released code or data, so the numbers are not independently checkable.\n\nFor an ISMIR late-breaking demo, this is a reasonable discussion piece. As a citable empirical study, it is not there yet. I would not send this version to peer review; I would ask the authors to fix the sample discrepancy, add proper inference in the correlation analysis, collect human ratings for the distorted captions, and release the data. If they do that, the result could be worth revisiting.","headline":"A plausible warning about n-gram metrics for music captioning, but the correlation evidence is too under-powered and inconsistent to carry the central claim.","tokens_in":3969,"tokens_out":2583,"would_cite":false,"duration_ms":26584,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper shows that BLEU, METEOR, ROUGE, SPICE, and FENSE barely track listener judgments of music captions, scoring paraphrases lower than distorted rewrites.","keywords":["music captioning","evaluation metrics","n-gram overlap","semantic alignment","human evaluation","MusicCaps","FENSE","mean opinion score"],"falsifier":"Run a larger human-rating study, for example 100 or more clips with all caption types rated by fresh listeners, and check whether n-gram metrics correlate positively with human MOS; if BLEU, METEOR, or ROUGE show robust correlation, or if humans rate the distorted captions as good as the paraphrases, the paper's conclusion would not hold.","tokens_in":3015,"feed_emoji":"🎵","tokens_out":5867,"duration_ms":51381,"temperature":0.7,"pith_summary":"This paper argues that the text-overlap metrics commonly used to evaluate music captions do not measure whether a caption is semantically aligned with the audio. In a human listening study on the MusicCaps evaluation set, the authors find that paraphrased captions with similar human quality ratings receive much lower BLEU, METEOR, ROUGE, and SPICE scores than the originals, while semantically distorted captions can score higher. They also report that all five tested metrics, including the audio-specific FENSE, have near-zero Pearson correlation with human mean opinion scores (around 0.08). The point is that current automated scores can mislead researchers about which captioning models actually produce descriptions people agree with.","feed_headline":"Music caption scores barely match human listening judgments","feed_subtitle":"BLEU, METEOR, ROUGE, and even FENSE correlate near zero with human ratings on MusicCaps.","key_machinery":"The load-bearing mechanism is the contrast between three caption types, original, paraphrased, and distorted, measured against human mean opinion scores. Paraphrases hold semantics roughly constant while changing surface wording; distorted captions hold syntax roughly constant while changing musical facts. Running standard n-gram and embedding-based metrics over these pairs exposes which scores track wording rather than meaning, and correlating metric scores with human MOS across MusicCaps quantifies how far the metrics depart from listener judgment.","core_discovery":"The central claim is that existing captioning metrics are overly sensitive to syntactic variation and lack alignment with actual human evaluations when applied to music captions. Using 50 listeners and 30 sampled clips from MusicCaps, the authors compare human mean opinion scores with metric scores for original, model-generated, and paraphrased captions. N-gram metrics such as BLEU, METEOR, and ROUGE fall sharply for paraphrases that human raters judge about as good as the originals, and distorted captions that preserve syntax but change musical content receive higher overlap scores than meaning-preserving paraphrases. The correlation table (20 entries) shows all metrics, including FENSE, sit near zero, leading the authors to conclude that the field needs a reevaluation of how music captions are scored.","pith_inferences":["A natural next step, not run in the paper, is to test whether a semantic metric built from music-tag or metadata categories (genre, instrument, mood) correlates with human MOS more strongly than FENSE does.","If the near-zero correlations generalize beyond MusicCaps, then other audio-captioning benchmarks may need similar human-validation studies before their automatic scores are trusted.","Because the paper's correlation analysis uses 20 entries rather than the 30 sampled clips, a larger replication study would give a sharper estimate of the true correlation values and test whether the near-zero result is stable."],"forward_implications":["If these results hold, published comparisons of music-captioning models that rely on BLEU or ROUGE may rank models by wording similarity rather than by perceived caption quality.","A metric that rewards synonyms and restructured sentences while checking musical content, such as genre, instruments, and mood, would be needed to replace n-gram overlap.","Distorted captions scoring above paraphrases means n-gram metrics can be gamed by outputting fluent-looking text that repeats reference syntax, so automated leaderboards should not be treated as quality measures.","Even FENSE, which was designed with audio captions in mind, does not align with human judgments on music captions, so the problem is not solved simply by borrowing a newer NLP metric."],"supporting_citations":[{"why":"Supplies the MusicCaps evaluation set and original human-written captions used for all metric comparisons.","marker":"[10]"},{"why":"Defines BLEU, the n-gram overlap metric that the paper shows drops sharply for paraphrases.","marker":"[6]"},{"why":"Defines METEOR, another n-gram metric whose scores favor distorted captions over paraphrases.","marker":"[7]"},{"why":"Defines ROUGE, the third overlap-based metric under scrutiny in the correlation study.","marker":"[8]"},{"why":"Introduces FENSE, an audio-caption metric that the paper tests and finds also uncorrelated with human MOS.","marker":"[11]"},{"why":"Provides LP-MusicCaps, the model whose output supplies the 'inference' caption type in the human evaluation.","marker":"[2]"}],"fun_headline_variants":["Captions scored by BLEU? Humans disagree","Music caption metrics ignore human taste","Old metrics fail music captions","BLEU, METEOR misjudge music captions","Human vs metric: music caption gap"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire argument depends on the authorship judgment that the paraphrased captions truly mean the same as the originals and the distorted captions do not, a judgment that is not independently verified by human ratings for the distorted captions.","fun_headline_variants_meta":{"raw":{"variants":["Captions scored by BLEU? Humans disagree","Music caption metrics ignore human taste","Old metrics fail music captions","BLEU, METEOR misjudge music captions","Human vs metric: music caption gap"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00013,"raw_usage":{"total_tokens":1043,"prompt_tokens":781,"completion_tokens":262,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":397,"completion_tokens_details":{"reasoning_tokens":195}},"tokens_in":397,"tokens_out":262,"duration_ms":3224,"temperature":1.0,"reasoning_tokens":195,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T18:13:17.354680+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a larger human-rating study, for example 100 or more clips with all caption types rated by fresh listeners, and check whether n-gram metrics correlate positively with human MOS; if BLEU, METEOR, or ROUGE show robust correlation, or if humans rate the distorted captions as good as the paraphrases, the paper's conclusion would not hold.","supporting_citations":[{"cited_title":"Bleu: a method for automatic evaluation of machine transla- tion,","cited_arxiv_id":null,"evidence_quote":"Supplies the MusicCaps evaluation set and original human-written captions used for all metric comparisons."},{"cited_title":"Llark: A multimodal instruction-following language model for music,","cited_arxiv_id":null,"evidence_quote":"Defines METEOR, another n-gram metric whose scores favor distorted captions over paraphrases."},{"cited_title":"Meteor: An automatic met- ric for mt evaluation with improved correlation with human judgments,","cited_arxiv_id":null,"evidence_quote":"Introduces FENSE, an audio-caption metric that the paper tests and finds also uncorrelated with human MOS."},{"cited_title":"We use Amazon Mechanical Turk [9] to recruit 50 par- ticipants for a listening test","cited_arxiv_id":null,"evidence_quote":"Provides LP-MusicCaps, the model whose output supplies the 'inference' caption type in the human evaluation."}],"review_version":1}