REVIEW 4 major objections 5 minor 34 references
High agreement with human ratings does not guarantee that large audio-language model judges are grounded in the audio; the paper shows several judges instead follow protocol-supplied labels or slot positions, so judge validity depends on th
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 05:03 UTC pith:STGVCSJ4
load-bearing objection A genuinely useful audit with two strong findings — wrong specialist labels override audio on emotion for most judges, and Qwen3-Omni-Thinking locks to a slot in concat A/B — but the 'protocol-level' generalization is softer than the abstract claims, and the missing error bars and proxy cells need fixing before publication. the 4 major comments →
Auditing Protocol-Level Shortcuts in Large Audio Language Model Judges for Speech Evaluation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On the paper's own terms, the central discovery is that protocol-level shortcuts are real, measurable, and common: in feature-blueprint judging, when a structured text description of acoustic features contains a wrong specialist label, five of six LALM judges copy it (emotion accuracy falls to 0.10 or below) even when the real audio is also supplied; in reference-conditioned judging, wrong reference labels are followed at rates that depend on where in the prompt the label appears; and in pairwise A/B comparison, judges can lock onto the same slot under order swaps — Qwen3-Omni-Thinking and Audio-Flamingo-3 do so on every trial. The pattern recurs on a second attribute for reference-following
What carries the argument
The central machinery is a matched counter-condition audit. For each deployment protocol the paper constructs a perturbation that should leave an audio-grounded verdict unchanged — a deliberately wrong specialist label in the blueprint, a wrong reference label placed in four different prompt positions, and AB/BA order swaps for pairwise comparison — and measures how often the verdict tracks the perturbation instead of the audio. Three metrics carry the argument: the reference-follow rate (and its chance-normalized version), the position-lock rate (same slot chosen in both orders), and the gold-aligned rate (chosen slot holds the better clip). These make a latent failure observable: a judge t
Load-bearing premise
The audit's conclusions rest on the premise that an audio judge ought to treat the audio as the ground truth and any supplied label or hint as at most a suggestion; if those labels are actually meant to be authoritative instructions, then following a wrong label would be instruction-following rather than a validity failure, and the central shortcut claim would not follow.
What would settle it
Run the same counter-conditions with an explicit prompt instruction that the supplied labels are fallible and should be ignored if they conflict with what the judge hears. If the wrong-label and wrong-reference following largely disappears, then the observed behavior is instruction-following under ambiguity rather than a stable shortcut; if it persists, the paper's shortcut reading is supported.
If this is right
- Aggregate agreement with human ratings should no longer be treated as sufficient evidence that an LALM judge is audio-grounded.
- Each model–protocol combination needs a matched shortcut probe — wrong-label blocks, placement sweeps, and order-swapped formats — before its scores are used.
- Reference-conditioned judging is prompt-placement-dependent; a wrong label in a rubric hint or system instruction is followed far more often than in other positions.
- Feature-blueprint judging is not reliable for attributes a judge cannot infer from audio; a wrong specialist label can drive accuracy to zero even when audio is present.
- Pairwise A/B results should be reported with position-lock and format (concatenated vs. native multi-audio input) information, since slot-locking can masquerade as a genuine preference.
Where Pith is reading between the lines
- If this pattern generalizes, a cheap pre-deployment screen for any new LALM judge would be to measure its audio-only accuracy per attribute and skip or distrust attributes near the base rate; the paper's capability-dependence result suggests shortcut proneness is partly predictable from that number.
- The finding that changing the input format (concat vs. native) can reduce slot-locking without improving gold-aligned accuracy implies that format fixes alone will not produce trustworthy preference judgments; engagement and preference need separate probes.
- The audit's criterion — audio as ground truth, side information as hint — is a normative choice. An explicit test would vary the prompt instruction (e.g., 'ignore labels that conflict with what you hear') and see whether label-following drops; if it does, the effect is closer to instruction-following under ambiguity than to a hard-wired shortcut.
- The matched-counter-condition template could transfer to other judge settings (image, video, text), where reference labels, specialist metadata, and position are similarly available as side channels.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes a protocol-level audit of LALM-as-a-judge for speech evaluation. It defines three deployment protocols—feature-blueprint judging, reference-conditioned judging, and pairwise A/B comparison—and pairs each with counter-conditions: wrong specialist labels with and without audio and with/without 'a classifier says' framing; wrong reference labels at four prompt placements; and AB/BA order swaps under concat and native audio input. Across six LALMs and attributes (emotion, language, naturalness, speaker similarity), the experiments show that several judges copy wrong specialist emotion labels even when audio is present, follow wrong references at placement-dependent rates, and lock onto the same A/B slot under order swaps, while language judgment resists the same manipulations for audio-capable judges. The paper concludes that aggregate agreement with human ratings is insufficient and that validity must be assessed for each model–protocol pair with a matched shortcut probe.
Significance. If the empirical patterns are robust, the paper makes a useful contribution to the LALM-judge literature by shifting validation from aggregate accuracy to protocol-level diagnostics. Its strengths are the matched counter-condition design, the cross-attribute replication (emotion/language for pointwise protocols; four tasks for pairwise), the inclusion of a no-frame condition for stated authority, the cue-conflict synthesis, and a seed-sweep stability check for the main blueprint table. The central caution—that a judge can score high by copying protocol-supplied labels rather than by listening—is important and actionable. However, the paper's interpretive claim rests on a normative premise about the role of side information, and several quantitative claims lack uncertainty quantification; these need work before the conclusions can be fully trusted.
major comments (4)
- [§III, first paragraph; §V-B/V-C] The audit's shortcut interpretation rests on the stated premise that a trustworthy judge should treat audio as ground truth and protocol side information as at most a hint. This premise is asserted, not defended, and the RQ1/RQ2 counter-conditions phrase the side information as an authoritative assertion ('a classifier has labeled this clip as X'; 'the reference label is X'). Under an alternative reading, a judge that trusts supplied labels is following the protocol's explicit input, not taking a shortcut. The wrong+audio condition is more decisive, but for judges with near-chance audio-only emotion accuracy (Qwen3-Omni-Thinking 0.13, Voxtral-Small 0.17, Table II), following the label could reflect a capability limit rather than a shortcut, as the paper itself acknowledges. This ambiguity directly affects the abstract claim that aggregate agreement 'does not guarantee' audio grounding. I
- [Tables II–V and §V-B] Rates are reported without confidence intervals or significance tests for the 30–200-trial measurements. The seed sweep covers only Table II; the lock/gold rates in Table IV and the associated four-mode taxonomy are based on small samples, and the absence of uncertainty makes it impossible to tell which differences are reliable. The paper should report exact binomial CIs or bootstrap intervals, and should test key contrasts (e.g., concat vs native lock for GPT-Audio; most- vs least-following placement in Table III). This is load-bearing because the qualitative conclusions are drawn from small-n rates that can be close to chance.
- [Table II footnote and §V-B] GPT-Audio's no-audio cells (skeleton, a.desc., true, wrong, no frame) are generated by gpt-4o, not by GPT-Audio. Consequently, the statement that 'inserting the true block raises every judge to the specialist's own accuracy' is misleading for GPT-Audio: the true-block cell does not come from the judge under test. The wrong+audio cell for GPT-Audio is valid, but the table's cross-condition comparison for this judge should either be removed or re-run with a text-capable checkpoint of the same model family.
- [Table IV NATIVE rows and §V-D] Native-format results are reported for only two judges (Gemini-3-Flash and GPT-Audio). The conclusion that concat lock is partly a format artifact is therefore supported by one judge's contrast. For the other four judges, there is no evidence about whether native input would change lock rates. The conclusion in §VI that pairwise judging is 'slot- or format-dependent' should be scaled back, or native multi-audio inputs should be obtained for the open-weight judges.
minor comments (5)
- [§V-C and Table III] The sentence 'The rubric slot triggers the largest reference-follow rate on both attributes' is not supported by Table III: for Audio-Flamingo-3 on emotion, the 'after' placement (0.41) exceeds rubric (0.33); for Gemini-3-Flash, system (0.85) exceeds rubric (0.58).
- [Table III] The 'match' column is never defined. Please specify whether it is accuracy under the true-reference condition and how it relates to Eq. (1).
- [Table II and §V-B] The 'no frame' column is not discussed in the results text. State what it shows for the authority-framing question; the current table suggests that dropping 'a classifier says' changes copying behavior substantially for several judges.
- [Abstract] The phrase 'incorrect specialist labels reduce five judges' emotion accuracy to 0.10 or below' is ambiguous: in the text-only 'wrong' column all six judges fall to 0.00–0.03, while the 0.10-or-below claim refers to the 'wrong+audio' condition. Please specify the condition in the abstract.
- [General] The paper lacks a limitations subsection acknowledging native-format coverage, sample sizes, and the GPT-Audio substitution. The prompt schematics in Fig. 1 are also hard to parse; including full prompts in an appendix would improve reproducibility.
Circularity Check
No significant circularity: the audit is an empirical, self-contained measurement study; its conclusions are not derived by definition or from self-referential fits.
full rationale
This paper is an empirical audit, not a derivation. The central claim—that several LALM judges rely on protocol-level shortcuts—is supported by measurements under deliberately altered protocols: wrong specialist labels in blueprints (Table II), wrong reference labels at varied prompt positions (Table III), and A/B order swaps (Table IV). These conditions are counter-factual probes: if a judge's verdict tracks a deliberately wrong label while audio is fixed, it is reasonable to conclude the verdict is not grounded in the audio. That inference is logical, but it is not circular because the paper does not fit any parameter to a target conclusion or define X in terms of Y. The paper explicitly states its normative premise in Section III: 'A trustworthy judge should treat the audio as the ground truth and the protocol's side information as, at most, a hint.' This is an evaluative criterion, not a result derived from the data. One could dispute this premise—for instance, arguing that models should treat supplied labels as authoritative instructions—but that is a correctness/assumption risk, not circularity. The paper also does not rely on load-bearing self-citations: prior work by one author (CLAIR/CLAIR-A) is only cited as related work, not used to justify the shortcut conclusion. No uniqueness theorem, no ansatz-via-citation, and no renaming of a known result as a new derivation appears. The capability-dependent interpretation (§V-B) is a post-hoc explanation of measured accuracy patterns, not an input to the measurements. Overall, the derivation chain is self-contained: the experimental conditions are designed to expose copying, not to produce a predetermined outcome.
Axiom & Free-Parameter Ledger
axioms (3)
- domain assumption A trustworthy judge should treat audio as ground truth and protocol side information as at most a hint; deliberately wrong side information should be ignored.
- domain assumption The two attributes (emotion, language) suffice to separate protocol-level from capability-dependent shortcuts, with audio-only accuracy as the capability measure.
- domain assumption The wrong-label permutation and blueprint templates do not accidentally cue the correct answer or co-vary with clip identity.
read the original abstract
Large audio-language models (LALMs) are increasingly used as automatic judges for speech evaluation. However, high agreement with human ratings does not guarantee that their verdicts are grounded in the audio. A judge may instead rely on specialist labels or reference data supplied by the evaluation protocol itself, taking a shortcut in place of listening to the audio. In this paper, we audit such protocol-level ``shortcuts'' in LALM judges across three common deployment protocols: feature-blueprint judging, where the audio is replaced by a structured text description of acoustic features, reference-conditioned judging, and pairwise A/B comparison. Across six judges and four attributes, we find that several LALMs rely on protocol-level shortcuts. For example, in feature-blueprint judging, incorrect specialist labels reduce five judges' emotion accuracy to 0.10 or below, and in concatenated A/B comparisons, Qwen3-Omni-Thinking often picks the same slot regardless of order swaps. These results indicate that aggregate agreement can overstate the validity of LALM judges unless the model and the evaluation protocol are assessed jointly, and that each model-protocol pair should be evaluated with a matched shortcut probe.
Figures
Reference graph
Works this paper leans on
-
[1]
Y . Chu, J. Xu, Q. Yang, H. Wei, X. Wei, Z. Guo, Y . Leng, Y . Lv, J. He, J. Lin, C. Zhou, and J. Zhou, “Qwen2-audio technical report,”arXiv preprint arXiv:2407.10759, 2024
Pith/arXiv arXiv 2024
-
[2]
Qwen2.5-omni technical report,
J. Xu, Z. Guo, J. He, H. Hu, T. He, S. Bai, K. Chen, J. Wang, F. Yang, K. Danget al., “Qwen2.5-omni technical report,”arXiv preprint arXiv:2503.20215, 2025
Pith/arXiv arXiv 2025
-
[3]
J. Xu, Z. Guo, H. Hu, Y . Chu, X. Wang, J. He, Y . Wang, X. Shi, T. He, X. Zhuet al., “Qwen3-omni technical report,”arXiv preprint arXiv:2509.17765, 2025
Pith/arXiv arXiv 2025
-
[4]
Audio flamingo 3: Advancing audio intelligence with fully open large audio language models,
A. Goel, S. Ghosh, J. Kim, S. Kumar, Z. Kong, S.-g. Lee, C.-H. H. Yang, R. Duraiswami, D. Manocha, R. Valle, and B. Catanzaro, “Audio flamingo 3: Advancing audio intelligence with fully open large audio language models,” inProceedings of Advances in Neural Information Processing Systems (NeurIPS), 2025, spotlight
2025
- [5]
-
[6]
Judging LLM-as-a-judge with MT-Bench and chatbot arena,
L. Zheng, W.-L. Chiang, Y . Sheng, S. Zhuang, Z. Wu, Y . Zhuang, Z. Lin, Z. Li, D. Li, E. P. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica, “Judging LLM-as-a-judge with MT-Bench and chatbot arena,” inProceedings of Advances in Neural Information Processing Systems (NeurIPS), vol. 36, 2023, pp. 46 595–46 623
2023
-
[7]
CLAIR: Evaluating image captions with large language models,
D. M. Chan, S. Petryk, J. E. Gonzalez, T. Darrell, and J. Canny, “CLAIR: Evaluating image captions with large language models,” inProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2023, pp. 13 638–13 646
2023
-
[8]
CLAIR-A: Lever- aging large language models to judge audio captions,
T.-H. Wu, J. E. Gonzalez, T. Darrell, and D. M. Chan, “CLAIR-A: Lever- aging large language models to judge audio captions,” inProceedings of the IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2025
2025
-
[9]
Audio large language models can be descriptive speech quality evaluators,
C. Chen, Y . Hu, S. Wang, H. Wang, Z. Chen, C. Zhang, C.-H. H. Yang, and E. S. Chng, “Audio large language models can be descriptive speech quality evaluators,” inProceedings of the International Conference on Learning Representations (ICLR), 2025
2025
-
[10]
EmergentTTS- Eval: Evaluating TTS models on complex prosodic, expressiveness, and linguistic challenges using model-as-a-judge,
R. R. Manku, Y . Tang, X. Shi, M. Li, and A. Smola, “EmergentTTS- Eval: Evaluating TTS models on complex prosodic, expressiveness, and linguistic challenges using model-as-a-judge,” inProceedings of Advances in Neural Information Processing Systems (NeurIPS), 2025
2025
-
[11]
K. Huang, Q. Tu, L. Fan, C. Yang, D. Zhang, S. Li, Z. Fei, Q. Cheng, and X. Qiu, “InstructTTSEval: Benchmarking complex natural-language instruction following in text-to-speech systems,”arXiv preprint arXiv:2506.16381, 2025
Pith/arXiv arXiv 2025
-
[12]
SpeechLLM-as-judges: Towards general and inter- pretable speech quality evaluation,
H. Wanget al., “SpeechLLM-as-judges: Towards general and inter- pretable speech quality evaluation,” inProceedings of the Annual Meeting of the Association for Computational Linguistics (ACL), 2026
2026
-
[13]
AudioJudge: Understanding what works in large audio model based speech evaluation,
P. Manakul, W. H. Gan, M. J. Ryan, A. S. Khan, W. Sirichotedumrong, K. Pipatanakul, W. Held, and D. Yang, “AudioJudge: Understanding what works in large audio model based speech evaluation,”arXiv preprint arXiv:2507.12705, 2025
Pith/arXiv arXiv 2025
-
[14]
Hearing between the lines: Unlocking the reasoning power of LLMs for speech evaluation,
A. Chandra, K. Miller, V . Ravichandran, C. Papayiannis, and V . Saligrama, “Hearing between the lines: Unlocking the reasoning power of LLMs for speech evaluation,” inProceedings of the Findings of the Association for Computational Linguistics: EACL 2026, 2026
2026
-
[15]
L. H.-Y . Foo, C.-K. Yang, C.-A. Li, K.-H. Lu, and H.-y. Lee, “All that glitters is not audio: Rethinking text priors and audio reliance in audio- language evaluation,”arXiv preprint arXiv:2604.24401, 2026
Pith/arXiv arXiv 2026
-
[16]
Do audio LLMs really LISTEN, or just transcribe? measuring lexical vs. acoustic emotion cues reliance,
J. Chen, Z. Guo, J. Chun, P. Wang, A. Perrault, and M. Elsner, “Do audio LLMs really LISTEN, or just transcribe? measuring lexical vs. acoustic emotion cues reliance,” inProceedings of the Conference of the European Chapter of the Association for Computational Linguistics (EACL), 2026
2026
-
[17]
When audio-LLMs don’t listen: A cross-linguistic study of modality arbitration,
J. Billa, “When audio-LLMs don’t listen: A cross-linguistic study of modality arbitration,”arXiv preprint arXiv:2602.11488, 2026
arXiv 2026
-
[18]
Do audio LLMs listen or read? analyzing and mitigating paralinguistic failures with V oxParadox,
J. Pang, A. Chaubey, and M. Soleymani, “Do audio LLMs listen or read? analyzing and mitigating paralinguistic failures with V oxParadox,” inProceedings of the International Conference on Machine Learning (ICML), 2026
2026
-
[19]
LALM-as-a-Judge: Benchmarking large audio-language models for safety evaluation in multi-turn spoken di- alogues,
A. Ivry and S. Watanabe, “LALM-as-a-Judge: Benchmarking large audio-language models for safety evaluation in multi-turn spoken di- alogues,” inProceedings of the International Conference on Machine Learning (ICML), 2026
2026
-
[20]
Large language models are not fair evaluators,
P. Wang, L. Li, L. Chen, Z. Cai, D. Zhu, B. Lin, Y . Cao, L. Kong, Q. Liu, T. Liu, and Z. Sui, “Large language models are not fair evaluators,” in Proceedings of the 62nd Annual Meeting of the Association for Com- putational Linguistics (Volume 1: Long Papers). Bangkok, Thailand: Association for Computational Linguistics, 2024, pp. 9440–9450
2024
-
[21]
Justice or prejudice? quantifying biases in LLM-as-a-judge,
J. Ye, Y . Wang, Y . Huang, D. Chen, Q. Zhang, N. Moniz, T. Gao, W. Geyer, C. Huang, P.-Y . Chen, N. V . Chawla, and X. Zhang, “Justice or prejudice? quantifying biases in LLM-as-a-judge,” inProceedings of the International Conference on Learning Representations (ICLR), 2025
2025
-
[22]
Towards understanding sycophancy in language models,
M. Sharma, M. Tong, T. Korbak, D. Duvenaud, A. Askell, S. R. Bowman, E. Durmus, Z. Hatfield-Dodds, S. R. Johnston, S. M. Kravec et al., “Towards understanding sycophancy in language models,” inPro- ceedings of the International Conference on Learning Representations (ICLR), 2024
2024
-
[23]
SycEval: Evaluating LLM sycophancy,
A. Fanouset al., “SycEval: Evaluating LLM sycophancy,” inProceed- ings of the AAAI/ACM Conference on AI, Ethics, and Society (AIES), 2025, pp. 893–900
2025
-
[24]
Shortcut learning in deep neural networks,
R. Geirhos, J.-H. Jacobsen, C. Michaelis, R. Zemel, W. Brendel, M. Bethge, and F. A. Wichmann, “Shortcut learning in deep neural networks,”Nature Machine Intelligence, vol. 2, no. 11, pp. 665–673, 2020
2020
-
[25]
The pitfalls of simplicity bias in neural networks,
H. Shah, K. Tamuly, A. Raghunathan, P. Jain, and P. Netrapalli, “The pitfalls of simplicity bias in neural networks,” inProceedings of Advances in Neural Information Processing Systems (NeurIPS), vol. 33, 2020, pp. 9573–9585
2020
-
[26]
Large language models can be lazy learners: Analyze shortcuts in in-context learning,
R. Tang, D. Kong, L. Huang, and H. Xue, “Large language models can be lazy learners: Analyze shortcuts in in-context learning,” inFindings of the Association for Computational Linguistics: ACL 2023. Toronto, Canada: Association for Computational Linguistics, 2023, pp. 4645– 4657
2023
-
[27]
The Geneva minimalistic acoustic parameter set (GeMAPS) for voice research and affective computing,
F. Eyben, K. R. Scherer, B. W. Schuller, J. Sundberg, E. Andr ´e, C. Busso, L. Y . Devillers, J. Epps, P. Laukka, S. S. Narayanan, and K. P. Truong, “The Geneva minimalistic acoustic parameter set (GeMAPS) for voice research and affective computing,”IEEE Transactions on Affective Computing, vol. 7, no. 2, pp. 190–202, 2016
2016
-
[28]
The Ryerson Audio-Visual Database of Emotional Speech and Song (RA VDESS): A dynamic, multimodal set of facial and vocal expressions in North American English,
S. R. Livingstone and F. A. Russo, “The Ryerson Audio-Visual Database of Emotional Speech and Song (RA VDESS): A dynamic, multimodal set of facial and vocal expressions in North American English,”PLoS ONE, vol. 13, no. 5, p. e0196391, 2018
2018
-
[29]
The V oiceMOS Challenge 2022,
E. Cooper, W.-C. Huang, Y . Tsao, H.-M. Wang, T. Toda, and J. Ya- magishi, “The V oiceMOS Challenge 2022,” inProceedings of INTER- SPEECH, 2022
2022
-
[30]
V oxCeleb: A large-scale speaker identification dataset,
A. Nagrani, J. S. Chung, and A. Zisserman, “V oxCeleb: A large-scale speaker identification dataset,” inProceedings of INTERSPEECH, 2017, pp. 2616–2620
2017
-
[31]
emotion2vec: Self-supervised pre-training for speech emotion rep- resentation,
Z. Ma, Z. Zheng, J. Ye, J. Li, Z. Gao, S. Zhang, and X. Chen, “emotion2vec: Self-supervised pre-training for speech emotion rep- resentation,” inProceedings of the Findings of the Association for Computational Linguistics: ACL 2024, 2024, pp. 15 747–15 760
2024
-
[32]
Robust speech recognition via large-scale weak super- vision,
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak super- vision,” inProceedings of the International Conference on Machine Learning (ICML), ser. Proceedings of Machine Learning Research, vol
-
[33]
ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification,
B. Desplanques, J. Thienpondt, and K. Demuynck, “ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification,” inProc. Interspeech, 2020, pp. 3830–3834
2020
-
[202]
28 492–28 518
PMLR, 2023, pp. 28 492–28 518
2023
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.