{"id":"f16789be-b91e-45e0-97b0-0d0474e642af","arxiv_id":"2506.05984","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Gemini-2.5-Pro can act as an automatic judge of speaking style in speech, with human agreement comparable to human-human agreement.","lead":"This paper tests whether audio-aware large language models can score how well a spoken sentence matches requested emotional and prosodic styles. It finds that Gemini-2.5-Pro's scores agree with human ratings about as much as human raters agree with each other.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim rests on correlations with a 4-rater human baseline and no significance testing; on role-playing human-human r=0.253, so Gemini-human r=0.319 may be within noise.","rationale":"The paper is a useful and honest attempt to validate ALLMs as judges of speaking styles, and the comparison framework is reasonable. The single load-bearing weakness is that the central claim is an equivalence claim between two correlation coefficients, but no inferential statistics are reported. On role-playing, human-human agreement is only 0.253, so a Gemini-human correlation of 0.319, while numerically higher, is not compelling evidence of judge quality without knowing its sampling variability. The 80 scores are not independent because they come from 20 contexts and 4 SLMs, which further reduces the effective sample size. A bootstrap confidence interval for the difference between the two correlations would directly test whether 'comparable' is more than noise. The paper's own limitation section is candid about subjectivity but does not provide this quantification. I therefore do not think the conclusion should be rejected; it should remain conditional on additional statistical validation, which is exactly the reader's verdict. My read does not change that verdict, so I recommend UNCHANGED.","tokens_in":12923,"tokens_out":6165,"duration_ms":63730,"concrete_test":"Bootstrap the Table 2 correlations clustered by instance: resample the 20 role-playing contexts (and, separately, the 20 voice-style IF instances) with replacement, keeping all four SLM outputs and all judge scores per context, and compute the 95% bootstrap CI for the difference between mean Gemini-human Pearson r and mean human-human Pearson r. If the CI includes zero or negative values, the 'comparable' claim is unsupported. Also report the cluster-adjusted effective sample size or ICC(2,1) among the four human raters; if ICC is below 0.4 on role-playing, the gold standard is too weak to anchor the comparison.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline claim ('the agreement between Gemini and human judges is comparable to the agreement between human evaluators') is supported by point estimates of Pearson's r in Table 2, but the paper reports no confidence intervals, significance tests, or inter-rater reliability statistics. On the role-playing task, human-human agreement is r=0.253 (Section 4.3), which the paper itself calls 'somewhat subjective'; the reported Gemini-human r=0.319 is only 0.066 higher. The correlations are computed over 80 scores that are not independent (20 contexts x 4 SLMs), so the effective sample size is smaller than 80 and the observed difference is plausibly sampling noise. On voice style IF, Gemini-human (0.640) is numerically above human-human (0.596), but again there is no uncertainty quantification, so 'comparable' is not established. If the human baseline is noisy, an ALLM-human correlation that merely equals or slightly exceeds it does not show the ALLM is a reliable judge; it may only show that the ALLM tracks a noisy gold standard. The paper's own limitation section is candid, but it does not address this inferential gap.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces StyleSet, two evaluation tasks for speaking-style controllability of spoken language models (SLMs): voice style instruction following and role-playing. Four SLMs generate speech, and two audio-aware large language models (ALLMs), GPT-4o-audio and Gemini-2.5-Pro, judge the generated speech on 5-point Likert scales (plus a binary realism rating for role-playing). The judgments are compared with four human raters using average Pearson correlations. The central claim is that Gemini-2.5-Pro agrees with human judges at a level comparable to human-human agreement (voice style IF: 0.640 vs. 0.596; role-playing style: 0.319 vs. 0.253), while GPT-4o-audio agrees less well (0.355 and 0.305). The paper also reports that current SLMs, including GPT-4o-audio, still struggle with fine-grained style control and realistic dialogue.","tokens_in":65,"tokens_out":4976,"duration_ms":102834,"significance":"If the central claim survives proper statistical scrutiny, the paper provides a useful automatic evaluation tool for a speech attribute—speaking style—that is otherwise expensive and subjective to evaluate. The authors release two tasks that can be reused, compare ALLM judgments against a human-human baseline, and are candid about limitations. The use of a human-human agreement reference is conceptually appropriate, and the experiments cover multiple SLM families and both instruction-following and interactive settings. However, the support for the headline claim currently rests on point estimates of Pearson correlations without confidence intervals or significance tests, and on a role-playing task with very low human-human agreement. The significance of the contribution therefore depends on whether the inferential gaps can be closed.","major_comments":[{"comment":"The headline claim that Gemini-human agreement is comparable to human-human agreement compares two point estimates, r=0.640 and r=0.596, with no confidence intervals or significance test. Moreover, the correlations are computed over 80 scores that are not independent, because the same 20 contexts are reused across four SLMs. A naive sample size of 80 overstates the precision of the estimates. I ask the authors to report cluster-bootstrapped confidence intervals (resampling by context, and also by SLM) or an equivalent uncertainty-aware analysis, and to state whether the difference is statistically detectable.","section":"Section 4.2, Table 2"},{"comment":"On the role-playing style task, the human-human Pearson correlation is r=0.253, which the paper itself characterizes as 'somewhat subjective.' The Gemini-human correlation is r=0.319, only 0.066 higher. Given only four raters and non-independent items, this difference is well within plausible sampling noise. The conclusion that 'evaluating role-playing with ALLM judges is at least as good as using human evaluators' is not supported without uncertainty quantification, and the low human-human baseline weakens the interpretation that an ALLM-human correlation of roughly 0.3 indicates reliable judging. Please provide confidence intervals, and ideally a comparison against a null distribution obtained by correlating one human rater's scores with the average of the remaining three raters on the same items.","section":"Section 4.3, Table 2"},{"comment":"The correlations pool 80 scores from four SLMs with substantially different mean quality. Pearson's r over such pooled data can be inflated by between-model separation, so the reported values may overstate instance-level agreement. To support the claim that ALLMs track human judgments on individual items, please report within-SLM correlations (or within-context correlations), and consider a variance decomposition that separates between-model from within-model agreement.","section":"Section 4.2, Table 2"},{"comment":"GPT-4o-audio is used both as one of the four SLMs whose outputs are judged and as one of the two ALLM judges. The paper acknowledges potential self-enhancement bias, but because GPT-4o-audio is presented as a candidate judge, the human-4o correlation of 0.355 could be affected by this overlap. I ask the authors to analyze this confound concretely, for example by reporting per-SLM correlations or by checking whether the 4o judge's scores for its own outputs differ systematically from human scores in ways that the other judge does not exhibit.","section":"Section 4.2, Section 4.3, Table 1"}],"minor_comments":[{"comment":"There is a typo in the introduction: 'Invoice style instruction following' should read 'Voice style instruction following.'","section":"Section 1"},{"comment":"The headings contain an extra space in 'V oice style IF' and similar phrases; please fix the spacing.","section":"Throughout"},{"comment":"The prompt for SLM 2 reads 'pretend that you are [role_2] and I am [role_2]', which appears to be a typo; presumably the second role should be [role_1]. Also, 'dialoge_context' should be 'dialogue_context'.","section":"Appendix A.2.1"},{"comment":"The aggregation of the five sampled judge responses is described as 'ensemble the verdicts', but the exact aggregation rule (majority vote, average, or something else) is not specified. Please state it explicitly for reproducibility.","section":"Section 4.1 and Appendix C"},{"comment":"The caption should clarify that 'Human–4o' and 'Human–Gemini' are averaged correlations between the ALLM judge and each of the four human raters, while 'Human–Human' is the average pairwise correlation among human raters; this is implied in the text but should be stated in the table.","section":"Table 2 caption"},{"comment":"The temperature-sensitivity analysis for Gemini (ranging from 0.640 to 0.649) is useful, but it quantifies sensitivity to decoding hyperparameters, not sampling variability due to items or raters; the paper should not present this as a substitute for confidence intervals.","section":"Section 4.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope and addresses an important evaluation gap, but the central claim currently rests on weak statistical footing. The revisions I request—uncertainty quantification, handling of non-independence, within-SLM correlations, and analysis of the GPT-4o-audio judge overlap—are feasible within the manuscript's scope and would make the contribution much stronger. I do not see a fundamental flaw that would require rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: the paper gives the speech community a cheap automatic judge for speaking styles, and the main empirical finding—Gemini-2.5-Pro agrees with human raters about as well as humans agree with each other—is new and worth taking seriously. The StyleSet construction is thoughtful: the voice-style task breaks instructions into fine-grained, testable units (emphasis, pace changes, non-verbal insertions) rather than coarse emotional labels, and the paper correctly notes that current SLMs, including GPT-4o-audio, still fail many of them. The comparison against prior failures (Jiang et al. 2025) is useful framing.\n\nThe core evidence, however, is thinner than the abstract suggests. On voice style IF, human-human r=0.596 and Gemini-human r=0.640—fine, but there are no confidence intervals or significance tests, and the 80 scores are not independent (4 models × 20 instances). On role-playing, human-human r=0.253, and Gemini-human r=0.319 is only 0.066 higher. The paper itself calls style evaluation 'somewhat subjective' here, but it nonetheless concludes that ALLMs are at least as good as humans. That conclusion is not supported by the point estimate alone; with four human raters, the noise in the human-human baseline is large, and a correlation of 0.319 against a noisy gold standard does not establish reliability. The self-enhancement risk of GPT-4o judging its own outputs is acknowledged but not controlled; the fact that Gemini agrees with the human ranking of 4o as best mitigates it, but only partially.\n\nThe paper is candid about limitations—no full-duplex dialogue, no pairwise comparison, English only—and I don't think any of those are load-bearing. The dataset is promised for release, but no code or data is currently shipped, which makes the 80-item result set hard to inspect independently.\n\nWho is this for? Researchers working on speech evaluation or SLM development. It's a useful empirical result, not a methodological breakthrough. With proper uncertainty quantification, a released dataset, and a more careful treatment of the role-playing baseline, it could become a standard reference. I'd send it to peer review as a solid submission, and I'd ask for the statistical strengthening as a major revision. A desk rejection would be wrong.","headline":"A promising but under-powered demonstration that Gemini-2.5-Pro can replace human raters for speaking-style evaluation; the voice-style result is solid, the role-playing result is too noisy to carry the claim.","tokens_in":6,"tokens_out":1984,"would_cite":true,"duration_ms":47106,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An audio-aware large language model can rate speaking style as consistently as human raters rate one another.","keywords":["audio-aware large language models","spoken language models","speaking style evaluation","LLM-as-a-judge","voice style instruction following","role-playing dialogue","speech generation evaluation","StyleSet"],"falsifier":"Recruit a much larger panel of human raters, say 20 to 50 per instance, for the same StyleSet audio, take the panel mean as the gold standard, and recompute Gemini's correlation; the central claim fails if Gemini's correlation against this more reliable standard drops below the pairwise human-human correlation measured on the same audio.","tokens_in":12722,"feed_emoji":"🎙️","tokens_out":11127,"duration_ms":102593,"temperature":0.7,"pith_summary":"This paper argues that an off-the-shelf audio-aware large language model, Gemini-2.5-Pro, can judge the speaking style of generated speech about as well as human raters agree with one another. Across two new tasks in the released StyleSet benchmark—voice style instruction following and role-playing—Gemini's scores correlate with human scores at $r = 0.640$ and $r = 0.319$, while the average pairwise human-human correlations are $r = 0.596$ and $r = 0.253$. The comparison is made against four human raters per task, and the same audio is scored by humans and by the audio-aware model using matching instructions. If the claim holds, spoken language models can be evaluated automatically on expressive, non-textual qualities instead of relying exclusively on paid human evaluation.","feed_headline":"Rates speaking style as consistently as human raters do","feed_subtitle":"Gemini-2.5-Pro scores correlate 0.640 with humans vs 0.596 between human raters.","key_machinery":"StyleSet is the central object: two 20-instance tasks, voice style instruction following and role-playing, that elicit speaking styles including emotion, volume, pace, pitch, word emphasis, and non-verbal elements such as laughter and sighs. The argument is carried by a judge pipeline in which each speech sample is scored from a text prompt containing the target style or role context, with chain-of-thought reasoning (step-by-step reasoning before the numeric verdict), five sampled verdicts ensembled by self-consistency, and regular-expression extraction of the numeric score. The load-bearing measurement is the average pairwise human-human Pearson correlation, treated as the reliability ceiling that an automatic judge must approach; the authors compare each audio-aware judge's correlation with each human rater against that ceiling.","core_discovery":"The paper's central claim is that audio-aware large language models can assess non-textual speaking style, and that the agreement between Gemini-2.5-Pro and human evaluators is comparable to the agreement among human evaluators themselves. On voice style instruction following, the average pairwise human-human Pearson correlation is 0.596, while the average human-Gemini correlation is 0.640; on role-playing style, the corresponding values are 0.253 and 0.319. GPT-4o-audio as a judge reaches 0.355 on the first task and 0.305 on the second. The paper takes these numbers as evidence that a general-purpose audio-aware LLM can stand in for at least one human rater when scoring generated speech for style adherence, role-appropriate emotion, prosody, pace, volume, emphasis, and non-verbal elements.","pith_inferences":["A practical downstream consequence the paper does not test: the four-rater human average is a thin anchor, and re-running the comparison with a larger rater pool would show whether Gemini's edge over human-human agreement survives a more reliable gold standard.","The same judge pipeline could be pointed at other speech attributes the paper explicitly leaves out, such as intelligibility, speaker similarity, or full-duplex dialogue, and the release of StyleSet makes such extensions directly testable.","Because the paper demonstrates correlation but not absolute-score calibration, practical use of audio-aware judges as accept or reject gatekeepers would require mapping automatic scores onto the human scale means, not just relying on rank agreement.","Translating StyleSet to non-English languages would test whether Gemini's judging advantage persists; the authors note that Step-Audio and Qwen-2.5-Omni may be stronger in Chinese, so language coverage could change both spoken model rankings and judge reliability."],"forward_implications":["Gemini-2.5-Pro can serve as a cheap, reproducible first-pass judge for voice style instruction following, replacing one human rater whose agreement with other humans would typically be lower or equal.","StyleSet can be run without hiring human evaluators, allowing repeated automatic comparison of spoken language model versions as they are updated.","Audio-aware judges and humans agree on the coarse verdict that GPT-4o-audio is the best of the four tested spoken language models, but none of the models fully masters style control: the best human-rated average is 3.65 out of 5 on voice style instruction following.","Human-rated realism of the best spoken language model role-play is 0.51, far below the 0.95 realism score of human-recorded dialogues, so dialogue naturalness remains a major open problem even if style judging becomes automatic.","Because Gemini-human agreement exceeds the human-human baseline on both tasks, the consistency of automatic style judging is not merely a copy of one rater's taste."],"supporting_citations":[{"why":"Defines GPT-4o and its audio mode, used both as the strongest spoken language model under evaluation and as one of the two audio-aware judges.","marker":"OpenAI, 2024"},{"why":"Supplies Gemini-2.5-Pro, the audio-aware judge whose scores match or exceed human-human correlation.","marker":"Google, 2025"},{"why":"Provides IEMOCAP, the source of the 20 role-playing contexts used in StyleSet.","marker":"Busso et al., 2008"},{"why":"Supports the choice to let judges produce chain-of-thought reasoning, which raises LLM-human agreement.","marker":"Chiang and Lee, 2023b"},{"why":"Introduces chain-of-thought prompting, the reasoning step the judge pipeline relies on before score extraction.","marker":"Wei et al., 2022"},{"why":"Supplies the self-consistency ensemble method used to aggregate five judge responses per instance.","marker":"Wang et al., 2023"},{"why":"Prior attempt at audio-aware judging of paralinguistic instruction following found poor alignment with humans, the negative baseline this paper must beat.","marker":"Jiang et al., 2025"},{"why":"Shows audio-aware LLMs can be fine-tuned for mean opinion score prediction, the narrower prior scope that this paper's general style judging extends beyond.","marker":"Chen et al., 2025"},{"why":"Documents self-enhancement bias in LLM-as-a-judge, used to interpret why GPT-4o-audio judging itself scores as it does.","marker":"Zheng et al., 2023"},{"why":"Justifies using pairwise correlation rather than raw agreement to measure inter-evaluator reliability.","marker":"Amidei et al., 2019"}],"fun_headline_variants":["Gemini-2.5-Pro judges speech style with human-level agreement","Audio LLM judge matches human-human agreement for style scoring","ALLM rates speaking style consistently with human evaluators","AI judge for speech style: Gemini rivals human raters","Speaking style assessment by ALLM aligns with human judgment"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim rests on treating the average of just four human raters per task as reliable ground truth, even though those raters agree only weakly with each other on role-playing style (correlation 0.253); if the human yardstick is noisy, matching or beating it is only weak evidence that an audio-aware model is a good judge.","fun_headline_variants_meta":{"raw":{"variants":["Gemini-2.5-Pro judges speech style with human-level agreement","Audio LLM judge matches human-human agreement for style scoring","ALLM rates speaking style consistently with human evaluators","AI judge for speech style: Gemini rivals human raters","Speaking style assessment by ALLM aligns with human judgment"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000275,"raw_usage":{"total_tokens":1624,"prompt_tokens":905,"completion_tokens":719,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":521,"completion_tokens_details":{"reasoning_tokens":637}},"tokens_in":521,"tokens_out":719,"duration_ms":7081,"temperature":1.0,"reasoning_tokens":637,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:12:46.679355+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recruit a much larger panel of human raters, say 20 to 50 per instance, for the same StyleSet audio, take the panel mean as the gold standard, and recompute Gemini's correlation; the central claim fails if Gemini's correlation against this more reliable standard drops below the pairwise human-human correlation measured on the same audio.","supporting_citations":[],"review_version":1}