{"id":"7f4278b0-54d0-45e9-b9bd-2e8b017f44c4","arxiv_id":"2506.23582","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A new open dataset of human relevance scores for text-audio pairs, plus a supervised predictor that outperforms CLAPScore.","lead":"RELATE is an open dataset of human relevance ratings for text-audio pairs, covering original and machine-synthesized sounds. The paper also trains a model that predicts those ratings and finds it beats the CLAPScore baseline.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Dataset screening may systematically exclude low-anchor listeners, and the reported SRCC gap of 0.032 over LAION-CLAP lacks uncertainty quantification; a bootstrap/SRCC CI check would settle whether the headline advantage is robust.","rationale":"The reader's weakest_assumption is the anchor-screening premise in Section 3.3, and my analysis agrees that this is the most load-bearing unverified assumption in the dataset construction. I also identify a closely related second issue: the headline model comparison in Table 6 lacks uncertainty quantification, so even if the labels are good, the claimed advantage over LAION-CLAP (SRCC 0.383 vs. 0.351) is not statistically established. Both concerns are empirical and can be settled by bootstrap CIs and screening-sensitivity reruns, which is why I keep the verdict CONDITIONAL rather than REJECT or UNCHANGED. The contribution is genuinely useful if the labels hold up: the dataset is open-sourced, the screening design is disclosed, the analysis is thoughtful, and the comparison to CLAPScore is a reasonable baseline. A conditional accept is appropriate until the uncertainty analysis and at least one screening-sensitivity check are provided. I do not see an internal inconsistency in the loss function, model architecture, or reported statistics; the weak point is purely the empirical robustness of the dataset and the headline comparison. I therefore agree with the reader's verdict and recommend no change in the verdict category, though the conditions should be made explicit as described.","tokens_in":8920,"tokens_out":1699,"duration_ms":15588,"concrete_test":"Bootstrap the test set (e.g., 10,000 resamples over the 1,311 audio-text pairs) and compute percentile CIs for the SRCC of Ours vs. LAION-CLAP and MS-CLAP from Table 6; report the fraction of resamples where Ours wins. Then rerun the REL model training/evaluation with two alternative screenings (e.g., no listener screening, and the stricter test threshold applied to training) and check whether the SRCC ordering and its CIs change. If the Ours-vs-LAION-CLAP SRCC difference has a CI crossing zero or reverses under a reasonable screening variant, the headline 'outperforms CLAPScore' is not established.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central contribution is RELATE as a reusable benchmark plus a supervised predictor that outperforms CLAPScore. The load-bearing condition is that the human REL labels are trustworthy enough to both train the model and serve as ground truth for the comparison. That condition is secured only by the anchor-screening premise in Section 3.3: mismatched AudioCaps label-audio pairs are assumed to be genuinely low-relevance, and thresholds (training average anchor < 2, test < 1) are assumed to remove only unreliable raters. If anchors are not reliably low for some text-audio pair types, or if the thresholds mainly remove listeners with particular rating styles or particular audio/text categories, then the labels, the factor analysis, and the benchmark numbers all inherit that bias. No inter-annotator agreement, label-noise analysis, or distribution of anchor scores is reported, so the effect of the screening on label quality is unquantified. Additionally, the headline claim itself is thin: Table 6 reports SRCC 0.383 vs. 0.351 for LAION-CLAP, but there are no confidence intervals, bootstrap resamples, or significance tests on the difference, and the test set is split without error bars. The per-category results in Table 7 also largely lack uncertainty quantification. The most load-bearing single check is therefore to quantify the uncertainty and label sensitivity of the central comparison: if the SRCC gap vanishes or reverses under bootstrap or under re-estimation with relaxed/stricter anchor thresholds, the paper's central claim does not hold. If the gap persists with CIs excluding zero and the labels are robust to the screening thresholds, the claim is credible.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces RELATE, an open-source dataset of subjective relevance scores between text and audio. The dataset contains original AudioCaps audio-text pairs plus audio synthesized by four TTA models, annotated by many listeners on 11-point REL, IS, and OS scales, together with listener attributes. The authors analyze the REL scores with respect to audio category and text complexity, using nonparametric tests and ART ANOVA. They then train a supervised relevance prediction model combining BYOL-A audio features, RoBERTa text features, listener embeddings, and a BLSTM with a class-balanced loss, and compare it against zero-shot CLAPScore baselines on a held-out split. The reported results show the proposed model achieving SRCC 0.383 versus 0.351 for LAION-CLAP and 0.181 for MS-CLAP on the test set, with per-category advantages in most sound categories.","tokens_in":9149,"tokens_out":5551,"duration_ms":58450,"significance":"If the dataset is reliable and the evaluation split is sound, RELATE is a potentially useful public benchmark for text-to-audio relevance evaluation, and the proposed model could serve as a practical screening tool. The paper's strengths include the open release of the dataset and code, the use of multiple TTA systems and a large listener pool, and the comparison against a reasonable zero-shot baseline. The main caveat is that the headline performance advantage is numerically small and is reported without uncertainty quantification, and the dataset quality rests on an anchor-screening procedure whose premises are not validated. The attribute analysis is also informative but is performed on a further-filtered subset that is not the released dataset.","major_comments":[{"comment":"The central claim that the proposed model outperforms CLAPScore is supported only by point estimates on a single split. The SRCC gap over LAION-CLAP is 0.032 (0.383 vs 0.351), which could be within sampling noise; the paper reports no confidence intervals, bootstrap resamples, or significance tests for any metric in Tables 6 and 7. I request bootstrap confidence intervals for all metrics, ideally with a paired test for the SRCC difference, and corresponding confidence intervals for the per-category SRCCs in Table 7.","section":"Sec. 5.4, Table 6"},{"comment":"The anchor-screening procedure is the only quality-control step that supports treating the subjective REL labels as ground truth, yet its two central premises are unverified. It is asserted that intentionally mismatched AudioCaps label-audio pairs are low-relevance, but no distribution of anchor scores is shown; and the thresholds (training average score below 2, test below 1) are presented without justification, sensitivity analysis, or a report of how many listeners were excluded. Because these labels both train the benchmark model and serve as evaluation ground truth, a failure of either premise would bias all reported results. Please provide anchor-score distributions, exclusion counts, an inter-annotator agreement metric (e.g., ICC or Krippendorff's alpha), and a sensitivity analysis with alternative thresholds.","section":"Sec. 3.3"},{"comment":"The sentence describing the data split, \"the test data of REL scores was divided into two subsets so there was no overlap between audio samples and texts,\" is ambiguous. Please state explicitly whether the training, validation, and test sets are pairwise disjoint in both audio and text, and report the numbers of unique audio items and unique text prompts in each split. If any text prompt appears in both training and test, the benchmark could be affected by text memorization.","section":"Sec. 5.3, Table 5"}],"minor_comments":[{"comment":"The full-text title reads \"RELA TE\" instead of \"RELATE.\"","section":"Title"},{"comment":"There is a missing space in the \"Audio duration [s]\" row, and the last entry is formatted inconsistently (\"11901\" instead of \"11,901\").","section":"Table 3"},{"comment":"The text says that CBL is effective, but Table 6 shows that \"Ours\" has higher MSE (0.073) than \"Ours w/o CBL\" (0.069), even though its SRCC is higher. Please acknowledge this trade-off explicitly or clarify that the effectiveness claim refers only to rank correlation metrics.","section":"Sec. 5.4, Table 6"},{"comment":"The stricter screening used for the attribute analysis (excluding listeners with average anchor score at least 2, average original-audio score at most 6, and the lowest 5% entropy) is not part of the released dataset. Please state clearly that Section 4 analyzes a further-filtered subset, and define how the rating entropy was computed.","section":"Sec. 4"},{"comment":"The paper says \"For each text, two synthesis models are selected and synthesized\" but does not describe how the two models are selected, whether the selection is random, or which seed is used. Please report this for reproducibility.","section":"Sec. 3.2"},{"comment":"The hyperparameters are described only as \"empirically chosen.\" Please report the validation metric used for model selection and the range of values that were explored, at least briefly.","section":"Sec. 5.3"}],"recommendation":"major_revision","confidential_remarks":"The dataset contribution is potentially valuable and within the scope of a datasets-and-benchmarks paper. The main risks are the unvalidated anchor screening and the lack of uncertainty quantification around a small performance gap. The proposed changes---bootstrap confidence intervals, inter-annotator agreement, anchor-score analysis, and a clearer description of the split---are feasible and would substantially strengthen the claims. I do not see a fundamental flaw that would require rejection if these points are adequately addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe one thing you should know: this is a dataset paper, and the dataset is the contribution. RELATE gives subjective relevance, inclusion, and order scores for text-audio pairs, with listener attributes, for AudioCaps originals plus outputs from four TTA models. That is genuinely new and needed; nothing in the cited prior work gives you subjective text-audio relevance with listener and item attributes. The benchmark model is a standard VoiceMOS-style regression, but it does the job.\n\nWhat the paper does well: the dataset construction is described in enough detail to trust the basic labels. The held-out split has no audio or text overlap, which is the right call. The analysis of factors (animal sounds score lower, longer texts and temporal prepositions hurt synthesized audio) is plausible and matches what I would expect from TTA models. The predictor beats CLAPScore on the main test set, and the trend is consistent across most categories.\n\nWhere the soft spots are. The headline SRCC gap is thin: 0.383 vs 0.351 for LAION-CLAP, on a single split, with no confidence intervals or significance tests. That gap could vanish with a different split or a bootstrap. The screening procedure in Section 3.3 is the load-bearing assumption: anchors are mismatched AudioCaps pairs, assumed to be low-relevance, and listeners with average anchor >=2 (train) or >=1 (test) are excluded. Those thresholds are disclosed but their effect on label quality is not quantified. No inter-annotator agreement is reported. The factor analysis adds more post-hoc exclusions (original-audio rating <=6, lowest 5% entropy) that are reasonable but could bias the trends. The paper also does not ship training code, which makes it harder to verify the model numbers.\n\nNone of this is disqualifying. The dataset is reusable and the central comparison is run on a held-out set. The worry is that the reported margins are fragile. I want to see error bars on the SRCC difference and a sensitivity check with relaxed or stricter anchor thresholds. If the dataset and code are released, that check is easy.\n\nWho it is for: anyone building or evaluating text-to-audio models who wants a cheap relevance check. It deserves a serious referee. I would send it out, with the expectation that the authors add uncertainty quantification and label-noise analysis.\n\nRecommendation: peer review, conditional on those revisions.","headline":"RELATE is a genuinely useful dataset for text-to-audio relevance evaluation, and the benchmark model is reasonable, but the headline advantage over CLAPScore needs uncertainty quantification before it can be trusted.","tokens_in":9767,"tokens_out":2225,"would_cite":true,"duration_ms":23685,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper builds a public dataset of human text–audio relevance scores and shows that a supervised predictor trained on it beats the standard CLAPScore baseline.","keywords":["text-to-audio","subjective evaluation","relevance score","CLAPScore","sound event synthesis","human evaluation","automatic evaluation metric","AudioCaps"],"falsifier":"Have a fresh group of listeners rate a random sample of RELATE pairs without anchor-based screening, then retrain the predictor on those unscreened labels; if the advantage over CLAPScore shrinks substantially, or if the new mean scores shift by more than a point on the 11-point scale, the benchmark result depends on the screening assumption.","tokens_in":8696,"feed_emoji":"🔊","tokens_out":8361,"duration_ms":74113,"temperature":0.7,"pith_summary":"This paper introduces RELATE, an open dataset of human ratings of how well a text description matches an audio clip, collected for original AudioCaps recordings and for audio synthesized by four text-to-audio generators. It then trains a supervised model on these ratings to predict the average human relevance score for a new text–audio pair, and reports that this model outperforms the standard unsupervised CLAPScore baseline on every correlation metric used. If that result holds, text-to-audio developers can screen large batches of synthesized audio for content mismatch without running costly listening tests, and different studies can compare against one fixed human ground truth. The dataset analysis also identifies where current synthesis tends to fail: animal sounds, longer texts, and texts with temporal ordering all receive lower relevance scores.","feed_headline":"New benchmark model beats CLAPScore on text-audio relevance","feed_subtitle":"RELATE adds 5,400 human-rated text-audio pairs and beats CLAPScore on most sound categories","key_machinery":"The load-bearing machinery is the dataset plus the predictor trained on it. RELATE is built from AudioCaps captions paired with original audio and with outputs of AudioLDM, AudioLDM2, Tango, and Tango2; each pair is rated on an 11-point scale for overall relevance, inclusion of sound events, and order of sound events. To keep labels reliable, each listening batch contains intentionally mismatched anchor pairs, and listeners whose average anchor rating is too high are excluded from the released labels. The predictor concatenates frozen BYOL-A audio features, RoBERTa text features, and a listener-embedding vector, passes them through a bidirectional LSTM and linear layers, and is trained with a class-balanced weighted sum of clipped mean-squared-error and contrastive losses.","core_discovery":"The central claim is that a supervised regressor trained on RELATE predicts human judgments of text–audio relevance better than CLAPScore, the established reference that computes cosine similarity in a pretrained audio–text embedding space. On the held-out test set the proposed model reaches a Spearman rank correlation of 0.383 with human scores, compared with 0.351 for LAION-CLAP and 0.181 for MS-CLAP, and it also improves linear correlation, Kendall's tau, and mean squared error. The improvement persists in most of the eight top-level sound categories, with the stated exceptions of animal sounds and source-ambiguous sounds, where LAION-CLAP keeps a small advantage. The paper further claims that the released data, with 5,460 text–audio pairs and more than 17,000 relevance ratings, is a reusable benchmark for automatic text-to-audio evaluation.","pith_inferences":["An extension the paper does not pursue: the same anchor-screening and supervised-prediction recipe could be transferred to other generative domains, such as text-to-music or text-to-video, where a cheap proxy for human relevance is equally needed.","The per-category correlation tables suggest that a single global metric can mask uneven performance; reporting category-level correlations, as this paper does, is likely to become the norm for TTA evaluation.","Because the model includes a listener embedding, a natural next step would be to condition it on a target listener's attributes and predict that listener's own score rather than only the crowd average."],"forward_implications":["A deployed version of the trained model can rank unseen text–audio pairs by predicted human relevance, making it practical to screen large synthesized corpora before human listening.","Researchers can use RELATE as a fixed ground-truth benchmark, so relevance scores reported by different text-to-audio systems become comparable across studies.","The category-level results imply that the supervised approach captures human judgment beyond a global bias, since its advantage over CLAPScore appears within most sound categories.","The released inclusion and order scores, collected but not yet modeled, provide a ready target for future work on finer-grained text-to-audio evaluation."],"supporting_citations":[{"why":"Supplies the original audio-caption pairs that anchor both the original and synthesized text-audio material and the mismatched pairs used for screening.","marker":"[13]"},{"why":"Introduces CLAPScore, the unsupervised cosine-similarity baseline that the proposed model must beat.","marker":"[7]"},{"why":"Provides the LAION-CLAP variant used to compute one of the two CLAPScore baselines.","marker":"[30]"},{"why":"Provides the MS-CLAP variant used to compute the other CLAPScore baseline.","marker":"[31]"},{"why":"One of the four pretrained text-to-audio models whose outputs are rated in the dataset.","marker":"[14]"},{"why":"One of the four pretrained text-to-audio models whose outputs are rated in the dataset.","marker":"[15]"},{"why":"One of the four pretrained text-to-audio models whose outputs are rated in the dataset.","marker":"[16]"},{"why":"One of the four pretrained text-to-audio models whose outputs are rated in the dataset.","marker":"[17]"},{"why":"Pretrained BYOL-A audio encoder that provides the audio features in the prediction model.","marker":"[24]"},{"why":"Class-balanced loss used to down-weight overrepresented score classes during training.","marker":"[28]"}],"fun_headline_variants":["RELATE dataset pushes text-audio relevance scoring past CLAP","New benchmark RELATE beats CLAPScore on relevance prediction","Supervised relevance model outperforms CLAPScore on RELATE benchmark","RELATE: human-rated pairs improve automatic relevance scoring","Text-audio relevance model gains edge over CLAPScore"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The dataset is only as trustworthy as the assumption that the intentionally mismatched anchor pairs are genuinely low in relevance and that the listener-exclusion thresholds remove unreliable raters, an assumption the paper does not test with inter-annotator agreement.","fun_headline_variants_meta":{"raw":{"variants":["RELATE dataset pushes text-audio relevance scoring past CLAP","New benchmark RELATE beats CLAPScore on relevance prediction","Supervised relevance model outperforms CLAPScore on RELATE benchmark","RELATE: human-rated pairs improve automatic relevance scoring","Text-audio relevance model gains edge over CLAPScore"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000223,"raw_usage":{"total_tokens":1394,"prompt_tokens":821,"completion_tokens":573,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":437,"completion_tokens_details":{"reasoning_tokens":500}},"tokens_in":437,"tokens_out":573,"duration_ms":5732,"temperature":1.0,"reasoning_tokens":500,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T21:37:22.606240+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have a fresh group of listeners rate a random sample of RELATE pairs without anchor-based screening, then retrain the predictor on those unscreened labels; if the advantage over CLAPScore shrinks substantially, or if the new mean scores shift by more than a point on the 11-point scale, the benchmark result depends on the screening assumption.","supporting_citations":[{"cited_title":"Subjective-aligned dataset and metric for text-to-video quality assessment,","cited_arxiv_id":null,"evidence_quote":"Supplies the original audio-caption pairs that anchor both the original and synthesized text-audio material and the mismatched pairs used for screening."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces CLAPScore, the unsupervised cosine-similarity baseline that the proposed model must beat."},{"cited_title":"A new readability yardstick","cited_arxiv_id":null,"evidence_quote":"Provides the MS-CLAP variant used to compute the other CLAPScore baseline."},{"cited_title":"Environmental sound synthesis from vocal imitations and sound event labels,","cited_arxiv_id":null,"evidence_quote":"One of the four pretrained text-to-audio models whose outputs are rated in the dataset."},{"cited_title":"Make-an-audio: Text-to-audio gen- eration with prompt-enhanced diffusion models,","cited_arxiv_id":null,"evidence_quote":"One of the four pretrained text-to-audio models whose outputs are rated in the dataset."},{"cited_title":"Text-to- audio generation using instruction guided latent diffusion model,","cited_arxiv_id":null,"evidence_quote":"One of the four pretrained text-to-audio models whose outputs are rated in the dataset."},{"cited_title":"AudioLDM 2: Learn- ing holistic audio generation with self-supervised pretraining,","cited_arxiv_id":null,"evidence_quote":"Pretrained BYOL-A audio encoder that provides the audio features in the prediction model."}],"review_version":1}