{"id":"47206271-491c-4faf-b3e5-07a97de56367","arxiv_id":"2608.06718","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Counterfactual audits show that current audio-language model judges often recover paralinguistic contrasts when explicitly compared, but fail to use the same cues in native one-context judgments, so aggregate accuracy overstates judge reliability.","lead":"This paper tests whether audio-language models actually use tone and emotion when judging assistant responses, by keeping the words identical and changing only how they are spoken. Most models, including strong Gemini systems, succeed when the two audio versions are shown side by side but fail when judging one audio alone, which undermines their use as automatic judges.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Single-turn audit labels rest on thin per-item human validation; the flagship pointwise collapse is conditional on item solvability, though the positional task provides independent support.","rationale":"The reader identified the same weak point: the single-turn validation is thin. I agree that this is the most load-bearing premise for the flagship single-turn collapse. However, the paper has independent support that keeps the central claim plausible: the positional task's human validation is 88-92%, and Gemini's pointwise-transition accuracies are 51.6-54.5% while pairwise-transition reaches 66.4-79.4%, so the protocol-dependence finding is not solely an artifact of the single-turn set. The aggregate-accuracy-hides-mechanism claim also has support from Figure 3, though item-level (P,O,J) state masses are not tabulated in the text. My verdict remains CONDITIONAL rather than ACCEPT or REJECT: a per-item human confirmation study and an item-level state table would settle the concern. Since the reader already assigned CONDITIONAL for essentially this reason, no verdict change is needed.","tokens_in":30013,"tokens_out":8518,"duration_ms":101311,"concrete_test":"Take a random sample of 100 items from SINGLE-TURN-EMOTIONS and obtain 10 independent native-English-speaker judgments per item under the POINTWISE + Hard Cue protocol. Report per-item majority accuracy and Fleiss' kappa. Recompute the Gemini pointwise vs pairwise protocol gaps from Table 3 on the subset of items whose majority human judgment is correct. If the gap persists (e.g., Gemini-3-Pro remains near 65% pointwise vs 91% pairwise), then item ambiguity is not the driver and the concern is resolved; if pointwise accuracy rises toward pairwise on the human-confirmed subset, the headline collapse is an artifact of unsolvable items.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing assumption is that the single-turn audit labels are solvable under the POINTWISE protocol, so that near-chance model accuracy indicates failure to use paralinguistic evidence. Section 4.3 and Appendix E support this with only five annotators per task, each seeing a random 50-item subset, and single-turn accuracy ranging from 62% to 100%; per-item agreement is not reported. The paper states that labels are fixed by construction rather than by human vote, but this does not settle the solvability question: a response written to fit emotion A can fail to be more appropriate than the response written for emotion B for a substantial share of items, and the 62% annotator floor shows some items are not clear. The pairwise success of the same Gemini models (74.3-91.0%) mitigates the concern, because solving PAIRWISE requires the cue to be present, but PAIRWISE may be solvable by explicit-contrast or low-level matching strategies that do not entail a stable POINTWISE label. Since the single-turn pointwise collapse is the paper's flagship example, the headline claim is conditional on per-item human solvability.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces counterfactual audits for evaluating audio-language models (ALMs) used as judges of spoken interactions. Each audit item fixes the transcript while varying the paralinguistic realization (affect, prosody, or the timing of an affective shift), forcing a valid judge to use the audio cue. Judges are evaluated in a native pointwise one-context protocol and a contrastive pairwise recoverability control, and each item is decomposed into three probes: perception (P), oracle response-mapping (O), and native judgment (J). Across Gemini, GPT, and open audio models, the authors report that Gemini models often achieve high pairwise accuracy (up to 91.0%) but near-chance pointwise accuracy (as low as 52.1%), while most other models remain near chance in both protocols. The diagnostic state analysis shows that similar aggregate accuracies can hide different failure modes, including 'Potemkin' failures where perception and response-mapping succeed but the native judgment fails. Robustness checks include generalization to real speech, generator-substitution, and response-quality balancing. The paper concludes that ALM judges should not be certified by accuracy alone.","tokens_in":30314,"tokens_out":6648,"duration_ms":71514,"significance":"If the results hold, this is a valuable methodological contribution to the evaluation of audio-language models: it provides a reusable instrument-auditing framework with a diagnostic state space that separates perception, response-mapping, and orchestration failures. The empirical finding that contrastive success can overstate native reliability is important for the growing use of ALMs as judges and reward models. The paper includes careful uncertainty reporting (Wilson CIs, paired bootstrap), multiple robustness checks, and a transparent limitations section. The promise of code release will aid reproducibility. The main risk is the load-bearing assumption that single-turn pointwise items are solvable by careful listeners, which is supported only by thin human validation.","major_comments":[{"comment":"The single-turn pointwise solvability of the audit items is not adequately established. Only five annotators each judged a random subset of 50 items, with accuracy ranging from 62% to 100%, and no per-item agreement is reported. Since the headline single-turn pointwise collapse in Table 3 interprets near-chance model accuracy as a failure to use paralinguistic evidence, item ambiguity would directly inflate the measured protocol gap. Please provide per-item human labels or majority-vote reliability (e.g., per-item agreement or kappa) on a larger sample, or re-analyze the headline results on a subset of items with high human agreement. This is load-bearing for the Section 5.1 claim that the relevant contrast is recoverable but not deployed in the native setting.","section":"Section 4.3 / Appendix E"},{"comment":"The diagnostic state masses π_poj are reported without uncertainty, and the marginal P and O probes are near chance for many models (e.g., GPT-4o-mini P=50.5% in Table 14). The joint state distribution is therefore potentially dominated by probe noise, so the 'Potemkin' mass and the claimed differences in failure modes between judges may not be statistically meaningful. Please report bootstrap confidence intervals for the state masses, or a sensitivity analysis under chance-corrected scoring of the P and O probes.","section":"Section 5.2 / Figure 3"},{"comment":"The sample sizes for single-turn pairwise evaluations are inconsistent with the stated 189 pairwise items (e.g., Gemini-2.5-Pro Pairwise Hard n=93; Gemini-3-Pro Pairwise No n=185). The paper should explain the exclusions (e.g., response-parsing failures, API errors) and confirm that the same item subset is used for all judges; otherwise the paired bootstrap protocol-gap comparisons in Table 6 may be computed over different item subsets for different judges.","section":"Table 7 / Section 4.1"},{"comment":"Using claude-haiku-4-5 as the LLM judge to rule out lexical-quality confounds introduces a circularity risk: the audit is designed to scrutinize LLM judges, yet another LLM judge is used to validate the absence of confounds. While this is a secondary robustness check, the paper should either use human ratings for the balance check or provide evidence that claude-haiku-4-5's quality judgments are themselves validated against human judgments.","section":"Section 5.3 / Table 4"}],"minor_comments":[{"comment":"The term 'Potemkin' is used throughout; please define it explicitly at first use (the reference to Mancoridis et al. 2025 is helpful).","section":"Section 2"},{"comment":"For the single-turn task, the Accuracy (%) column should clarify that the 62-100% range is per-annotator accuracy on a random 50-item subset, not per-item agreement; consider reporting the distribution of item-level agreement as well.","section":"Appendix E, Table 21"},{"comment":"There are minor typos in the appendix prompts and examples: 'hesistant' should be 'hesitant' in Listing 1, and 'Saurday' appears in the example transcript in Listing 8.","section":"Appendix B, Listings 1 and 8"},{"comment":"Please clarify whether the 378 pointwise instances are exactly the two audio realizations of each of the 189 items, and whether all judges evaluate the same set of pairwise items; the n values in Table 7 suggest otherwise.","section":"Section 4.1"},{"comment":"Please state more precisely how the GeminiGen vs GPTGen substitution was performed (e.g., whether only the annotation-generation LLM was swapped, or also the TTS rendering), and why the figure reports only the positional-emotion results.","section":"Figure 5"},{"comment":"The appendix prompt invokes Winoground and Winograd Schema; a citation for Winoground (e.g., Thrush et al., 2022) is missing from the reference list.","section":"Appendix F / References"}],"recommendation":"major_revision","confidential_remarks":"The core idea is strong and timely, and the positional-emotion results provide independent support for the protocol-collapse phenomenon even if the single-turn labels are imperfect. In the revision, I would prioritize strengthening the human-validation evidence for the single-turn pointwise items and adding uncertainty quantification for the diagnostic state masses. If the authors cannot provide per-item human agreement, they should temper the single-turn claims and frame the positional task as the primary evidence for contrastive overstatement."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The central claim here is solid: ALM judges, especially the Gemini family, perform well on pairwise contrastive matching but degrade sharply in the pointwise one-context setting that deployment actually requires. The numbers are consistent and the confidence intervals and paired bootstrap checks give me reasonable confidence in the protocol-collapse gap. What is actually new is the P/O/J diagnostic cascade, which lets the authors attribute failures to perception, response-mapping, or orchestration rather than just reporting an accuracy score. That is a genuine contribution to the audio-judge auditing literature.\n\nThe paper is also careful on validity. The counterfactual construction is sound, the real-vs-synthetic speech check correlates strongly with the synthetic results, the generator-substitution check addresses the obvious circularity concern, and the response-quality balance check rules out confounds. The related work is honest and does not oversell the novelty.\n\nThe soft spots are real but not load-bearing enough to sink the paper. The human validation for the single-turn task is thin: five annotators, 50 items each, and accuracy ranging from 62% to 100%. That means some items are genuinely hard or ambiguous, which weakens the interpretation that near-chance pointwise performance is purely a model failure. However, the positional task human validation is much stronger (88-92%), and the pairwise success of the same Gemini models mitigates the concern because solving pairwise requires the cue to be detectable. Also, the phrase \"collapse toward chance\" overstates what the tables show: the pointwise numbers are 60-65%, not 50%, so the model has some signal, it just is not deployable. The code and data are promised but not yet released, and the item-level state masses are not tabulated, which are both addressable.\n\nThe circularity from using Gemini to generate responses and judge them is controlled, not fatal. I would not reject this on those grounds.\n\nBottom line: this deserves a serious referee. The framework is useful, the empirical story is credible, and the limitations are stated rather than hidden. For revisions, I would want per-item human solvability data or a pre-registered filter, and released artifacts. I would bring it to a reading group and would likely cite the protocol-collapse result in work on audio evaluation.","headline":"A careful audit with a real central finding: audio-judge models can solve pairwise contrasts but collapse in native one-context judgment; the main caveat is thin human validation for the single-turn task.","tokens_in":30761,"tokens_out":1526,"would_cite":true,"duration_ms":19954,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that audio-language-model judges can pass contrastive matching yet fail native one-context judgments, so accuracy alone mis-certifies them.","keywords":["counterfactual audit","audio-language model judges","paralinguistic reasoning","response evaluation","pointwise vs pairwise judgment","diagnostic state decomposition","speech-to-speech evaluation","Potemkin failure"],"falsifier":"Run a larger human validation on all single-turn pointwise items under the hard cue: if a substantial fraction of the items that Gemini fails are also answered at chance by a panel of native English listeners, the pointwise collapse could reflect item ambiguity rather than model failure, whereas reliable human accuracy on the same items would confirm the collapse as a genuine judge limitation.","tokens_in":119,"feed_emoji":"🎧","tokens_out":4196,"duration_ms":57338,"temperature":0.7,"pith_summary":"The paper tries to establish that current audio-language model judges, when asked to pick the response appropriate to how a user sounded, frequently fail despite appearing capable. It introduces counterfactual audits: each item holds the transcript fixed and varies only affect, prosody, or the timing of an affective shift, so any valid judge must use the audio cue rather than lexical content. The headline result is that contrastive success overstates native judge reliability—Gemini models often match the two counterfactual worlds when both are shown, but collapse toward chance when only one audio context is given. The paper also shows that two judges with similar aggregate accuracy can fail in different places, which becomes visible through per-item diagnostic states. The upshot defended is that ALM judges should not be certified by end-to-end accuracy alone; deployment requires behavioral audits.","feed_headline":"Audio judges pass contrastive tests, fail one-context decisions","feed_subtitle":"Pairwise success overstates reliability; component probes pinpoint where audio judges fail.","key_machinery":"The central object is the counterfactual audit item, a tuple holding a fixed transcript or conversation history, two audio renderings of the same words, and two responses appropriate to each rendering. The load-bearing comparison is between pointwise native judgment, where the judge sees one audio context and two responses, and pairwise contrastive recoverability, where both contexts and both responses are visible and must be matched. Each item is then decomposed into probes for perception, oracle response-mapping, and native judgment, giving an eight-state diagnostic whose mass distribution separates perception bottlenecks, response-mapping bottlenecks, Potemkin orchestration failures, shortcut successes, and reliable integrated judgment.","core_discovery":"The central discovery is a protocol gap: a judge can recover a counterfactual contrast in a pairwise setting and still be unreliable in the native pointwise setting. Concretely, Gemini-3-Pro reaches 91.0% pairwise accuracy on single-turn emotion items with a hard cue but only 65.3% pointwise; Gemini-2.5-Pro drops from 86.0% to 60.1%. In the positional multi-turn task the same models hover near chance pointwise even when they solve the pairwise matching with a transition cue (Gemini-3-Flash: 79.4% pairwise vs 53.0% pointwise). The diagnostic decomposition into perception, oracle response-mapping, and native judgment further shows that similar aggregate accuracies hide different failure modes: Gemini models often pass perception and oracle response-mapping yet fail native judgment, which the paper calls a Potemkin failure, while GPT models show earlier perception bottlenecks. The authors state the conclusion directly: a model may distinguish the two counterfactuals when both are shown, but fail when the same audio cue must control a single-context decision.","pith_inferences":["If the pointwise collapse reflects a genuine competence–deployment gap, then a natural testable extension is to fine-tune ALMs on pointwise paralinguistic judgment using pairwise or oracle supervision; the diagnostic state distribution predicts which intervention—perception training, response-mapping training, or orchestration training—should help.","A consequence the paper does not draw is that the same audit could be applied to human listeners, since the five-annotator validation showed wide single-turn variation; reporting diagnostic states for humans would clarify whether annotator disagreement is perception-level or response-mapping-level.","Another implicit extension is to use the state masses as a regression-testing signal during deployment: monitoring the Potemkin failure state across model versions could detect when an audio upgrade fails to integrate into actual decisions."],"forward_implications":["An ALM judge that scores well on pairwise or contrastive evaluations should not be assumed reliable for single-context deployment; the pairwise–pointwise gap should be reported as a separate metric.","Aggregate accuracy is insufficient as a certification metric; judges should be audited at the component level to identify whether failures are perceptual, mapping, or orchestration.","For temporal-causal paralinguistic judgment, current ALM judges are especially brittle: near-chance pointwise performance on positional-emotion shows they cannot reliably localize and use an affective shift.","Even models with strong component skills can be unreliable end-to-end, since Gemini models often pass perception and oracle response-mapping but fail native audio judgment, so scaffolding such as explicit cues may be needed.","The findings extend to real speech: the synthetic-to-real comparison yields highly correlated diagnostic profiles, suggesting the failure modes are not simply artifacts of synthesized audio."],"supporting_citations":[{"why":"Supplies the EmoCF source data and the counterfactual response-evaluation idea that the single-turn audit builds on.","marker":"[Held et al., 2025]"},{"why":"Provides the concept of Potemkin understanding that the paper uses to interpret the perception-and-mapping-succeed-but-judgment-fails state.","marker":"[Mancoridis et al., 2025]"},{"why":"Supplies the pairwise test-time matching protocol that the paper adapts as its contrastive recoverability control.","marker":"[Zhu et al., 2026]"},{"why":"Motivates the use of audio-language models as automatic speech judges, the setting being audited.","marker":"[Manakul et al., 2025]"},{"why":"Provides the OD3 source conversations used to construct the positional multi-turn audit items.","marker":"[Chan et al., 2024]"},{"why":"Provides the emotion2vec scores used for audio-rendering quality control in dataset construction.","marker":"[Ma et al., 2024]"},{"why":"Provides DNSMOS, used in the response-quality balance check to rule out irrelevant audio-quality confounders.","marker":"[Reddy et al., 2021]"}],"fun_headline_variants":["Counterfactual audits expose audio judge blind spots","Pairwise wins hide audio judge failures in real use","Accuracy alone hides audio judge failure modes","Gemini passes perception yet fails native audio judgment","Potemkin failure: audio judges ace contrast, fail pointwise"],"cache_read_input_tokens":32896,"weakest_assumption_plain":"The audit labels are assumed to be solvable by careful listeners, so that near-chance model accuracy indicates model failure rather than item ambiguity; this rests on only five annotators per task, with single-turn accuracy ranging from 62% to 100%.","fun_headline_variants_meta":{"raw":{"variants":["Counterfactual audits expose audio judge blind spots","Pairwise wins hide audio judge failures in real use","Accuracy alone hides audio judge failure modes","Gemini passes perception yet fails native audio judgment","Potemkin failure: audio judges ace contrast, fail pointwise"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000276,"raw_usage":{"total_tokens":1640,"prompt_tokens":934,"completion_tokens":706,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":550,"completion_tokens_details":{"reasoning_tokens":632}},"tokens_in":550,"tokens_out":706,"duration_ms":6419,"temperature":1.0,"reasoning_tokens":632,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:38:23.551059+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a larger human validation on all single-turn pointwise items under the hard cue: if a substantial fraction of the items that Gemini fails are also answered at chance by a panel of native English listeners, the pointwise collapse could reflect item ambiguity rather than model failure, whereas reliable human accuracy on the same items would confirm the collapse as a genuine judge limitation.","supporting_citations":[],"review_version":2}