Pith. sign in

REVIEW 4 major objections 5 minor 34 references

High agreement with human ratings does not guarantee that large audio-language model judges are grounded in the audio; the paper shows several judges instead follow protocol-supplied labels or slot positions, so judge validity depends on th

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 05:03 UTC pith:STGVCSJ4

load-bearing objection A genuinely useful audit with two strong findings — wrong specialist labels override audio on emotion for most judges, and Qwen3-Omni-Thinking locks to a slot in concat A/B — but the 'protocol-level' generalization is softer than the abstract claims, and the missing error bars and proxy cells need fixing before publication. the 4 major comments →

arxiv 2607.13477 v1 pith:STGVCSJ4 submitted 2026-07-15 cs.SD cs.CLeess.AS

Auditing Protocol-Level Shortcuts in Large Audio Language Model Judges for Speech Evaluation

classification cs.SD cs.CLeess.AS
keywords LALM-as-a-judgespeech evaluationshortcut learningevaluation protocolreference followingposition biasfeature blueprintaudio grounding
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper tries to establish that large audio-language model judges used for speech evaluation can score high agreement with human ratings while actually relying on side information supplied by the evaluation protocol rather than on the audio itself. Across six judges and four attributes, the authors show that a wrong specialist label in feature-blueprint judging drives five judges' emotion accuracy to 0.10 or below even when the audio is present, that wrong reference labels are followed at placement-dependent rates, and that some judges lock onto the same A/B slot regardless of order swaps. The conclusion is that judge validity is a property of the model–protocol pair, not of the model alone, and that aggregate agreement should not be trusted until each pair is probed with matched counter-conditions. A careful reader would care because these judges are used as scalable proxies for human speech evaluation, and the audit offers concrete probes for making that use safer.

Core claim

On the paper's own terms, the central discovery is that protocol-level shortcuts are real, measurable, and common: in feature-blueprint judging, when a structured text description of acoustic features contains a wrong specialist label, five of six LALM judges copy it (emotion accuracy falls to 0.10 or below) even when the real audio is also supplied; in reference-conditioned judging, wrong reference labels are followed at rates that depend on where in the prompt the label appears; and in pairwise A/B comparison, judges can lock onto the same slot under order swaps — Qwen3-Omni-Thinking and Audio-Flamingo-3 do so on every trial. The pattern recurs on a second attribute for reference-following

What carries the argument

The central machinery is a matched counter-condition audit. For each deployment protocol the paper constructs a perturbation that should leave an audio-grounded verdict unchanged — a deliberately wrong specialist label in the blueprint, a wrong reference label placed in four different prompt positions, and AB/BA order swaps for pairwise comparison — and measures how often the verdict tracks the perturbation instead of the audio. Three metrics carry the argument: the reference-follow rate (and its chance-normalized version), the position-lock rate (same slot chosen in both orders), and the gold-aligned rate (chosen slot holds the better clip). These make a latent failure observable: a judge t

Load-bearing premise

The audit's conclusions rest on the premise that an audio judge ought to treat the audio as the ground truth and any supplied label or hint as at most a suggestion; if those labels are actually meant to be authoritative instructions, then following a wrong label would be instruction-following rather than a validity failure, and the central shortcut claim would not follow.

What would settle it

Run the same counter-conditions with an explicit prompt instruction that the supplied labels are fallible and should be ignored if they conflict with what the judge hears. If the wrong-label and wrong-reference following largely disappears, then the observed behavior is instruction-following under ambiguity rather than a stable shortcut; if it persists, the paper's shortcut reading is supported.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Aggregate agreement with human ratings should no longer be treated as sufficient evidence that an LALM judge is audio-grounded.
  • Each model–protocol combination needs a matched shortcut probe — wrong-label blocks, placement sweeps, and order-swapped formats — before its scores are used.
  • Reference-conditioned judging is prompt-placement-dependent; a wrong label in a rubric hint or system instruction is followed far more often than in other positions.
  • Feature-blueprint judging is not reliable for attributes a judge cannot infer from audio; a wrong specialist label can drive accuracy to zero even when audio is present.
  • Pairwise A/B results should be reported with position-lock and format (concatenated vs. native multi-audio input) information, since slot-locking can masquerade as a genuine preference.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If this pattern generalizes, a cheap pre-deployment screen for any new LALM judge would be to measure its audio-only accuracy per attribute and skip or distrust attributes near the base rate; the paper's capability-dependence result suggests shortcut proneness is partly predictable from that number.
  • The finding that changing the input format (concat vs. native) can reduce slot-locking without improving gold-aligned accuracy implies that format fixes alone will not produce trustworthy preference judgments; engagement and preference need separate probes.
  • The audit's criterion — audio as ground truth, side information as hint — is a normative choice. An explicit test would vary the prompt instruction (e.g., 'ignore labels that conflict with what you hear') and see whether label-following drops; if it does, the effect is closer to instruction-following under ambiguity than to a hard-wired shortcut.
  • The matched-counter-condition template could transfer to other judge settings (image, video, text), where reference labels, specialist metadata, and position are similarly available as side channels.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper proposes a protocol-level audit of LALM-as-a-judge for speech evaluation. It defines three deployment protocols—feature-blueprint judging, reference-conditioned judging, and pairwise A/B comparison—and pairs each with counter-conditions: wrong specialist labels with and without audio and with/without 'a classifier says' framing; wrong reference labels at four prompt placements; and AB/BA order swaps under concat and native audio input. Across six LALMs and attributes (emotion, language, naturalness, speaker similarity), the experiments show that several judges copy wrong specialist emotion labels even when audio is present, follow wrong references at placement-dependent rates, and lock onto the same A/B slot under order swaps, while language judgment resists the same manipulations for audio-capable judges. The paper concludes that aggregate agreement with human ratings is insufficient and that validity must be assessed for each model–protocol pair with a matched shortcut probe.

Significance. If the empirical patterns are robust, the paper makes a useful contribution to the LALM-judge literature by shifting validation from aggregate accuracy to protocol-level diagnostics. Its strengths are the matched counter-condition design, the cross-attribute replication (emotion/language for pointwise protocols; four tasks for pairwise), the inclusion of a no-frame condition for stated authority, the cue-conflict synthesis, and a seed-sweep stability check for the main blueprint table. The central caution—that a judge can score high by copying protocol-supplied labels rather than by listening—is important and actionable. However, the paper's interpretive claim rests on a normative premise about the role of side information, and several quantitative claims lack uncertainty quantification; these need work before the conclusions can be fully trusted.

major comments (4)
  1. [§III, first paragraph; §V-B/V-C] The audit's shortcut interpretation rests on the stated premise that a trustworthy judge should treat audio as ground truth and protocol side information as at most a hint. This premise is asserted, not defended, and the RQ1/RQ2 counter-conditions phrase the side information as an authoritative assertion ('a classifier has labeled this clip as X'; 'the reference label is X'). Under an alternative reading, a judge that trusts supplied labels is following the protocol's explicit input, not taking a shortcut. The wrong+audio condition is more decisive, but for judges with near-chance audio-only emotion accuracy (Qwen3-Omni-Thinking 0.13, Voxtral-Small 0.17, Table II), following the label could reflect a capability limit rather than a shortcut, as the paper itself acknowledges. This ambiguity directly affects the abstract claim that aggregate agreement 'does not guarantee' audio grounding. I
  2. [Tables II–V and §V-B] Rates are reported without confidence intervals or significance tests for the 30–200-trial measurements. The seed sweep covers only Table II; the lock/gold rates in Table IV and the associated four-mode taxonomy are based on small samples, and the absence of uncertainty makes it impossible to tell which differences are reliable. The paper should report exact binomial CIs or bootstrap intervals, and should test key contrasts (e.g., concat vs native lock for GPT-Audio; most- vs least-following placement in Table III). This is load-bearing because the qualitative conclusions are drawn from small-n rates that can be close to chance.
  3. [Table II footnote and §V-B] GPT-Audio's no-audio cells (skeleton, a.desc., true, wrong, no frame) are generated by gpt-4o, not by GPT-Audio. Consequently, the statement that 'inserting the true block raises every judge to the specialist's own accuracy' is misleading for GPT-Audio: the true-block cell does not come from the judge under test. The wrong+audio cell for GPT-Audio is valid, but the table's cross-condition comparison for this judge should either be removed or re-run with a text-capable checkpoint of the same model family.
  4. [Table IV NATIVE rows and §V-D] Native-format results are reported for only two judges (Gemini-3-Flash and GPT-Audio). The conclusion that concat lock is partly a format artifact is therefore supported by one judge's contrast. For the other four judges, there is no evidence about whether native input would change lock rates. The conclusion in §VI that pairwise judging is 'slot- or format-dependent' should be scaled back, or native multi-audio inputs should be obtained for the open-weight judges.
minor comments (5)
  1. [§V-C and Table III] The sentence 'The rubric slot triggers the largest reference-follow rate on both attributes' is not supported by Table III: for Audio-Flamingo-3 on emotion, the 'after' placement (0.41) exceeds rubric (0.33); for Gemini-3-Flash, system (0.85) exceeds rubric (0.58).
  2. [Table III] The 'match' column is never defined. Please specify whether it is accuracy under the true-reference condition and how it relates to Eq. (1).
  3. [Table II and §V-B] The 'no frame' column is not discussed in the results text. State what it shows for the authority-framing question; the current table suggests that dropping 'a classifier says' changes copying behavior substantially for several judges.
  4. [Abstract] The phrase 'incorrect specialist labels reduce five judges' emotion accuracy to 0.10 or below' is ambiguous: in the text-only 'wrong' column all six judges fall to 0.00–0.03, while the 0.10-or-below claim refers to the 'wrong+audio' condition. Please specify the condition in the abstract.
  5. [General] The paper lacks a limitations subsection acknowledging native-format coverage, sample sizes, and the GPT-Audio substitution. The prompt schematics in Fig. 1 are also hard to parse; including full prompts in an appendix would improve reproducibility.

Circularity Check

0 steps flagged

No significant circularity: the audit is an empirical, self-contained measurement study; its conclusions are not derived by definition or from self-referential fits.

full rationale

This paper is an empirical audit, not a derivation. The central claim—that several LALM judges rely on protocol-level shortcuts—is supported by measurements under deliberately altered protocols: wrong specialist labels in blueprints (Table II), wrong reference labels at varied prompt positions (Table III), and A/B order swaps (Table IV). These conditions are counter-factual probes: if a judge's verdict tracks a deliberately wrong label while audio is fixed, it is reasonable to conclude the verdict is not grounded in the audio. That inference is logical, but it is not circular because the paper does not fit any parameter to a target conclusion or define X in terms of Y. The paper explicitly states its normative premise in Section III: 'A trustworthy judge should treat the audio as the ground truth and the protocol's side information as, at most, a hint.' This is an evaluative criterion, not a result derived from the data. One could dispute this premise—for instance, arguing that models should treat supplied labels as authoritative instructions—but that is a correctness/assumption risk, not circularity. The paper also does not rely on load-bearing self-citations: prior work by one author (CLAIR/CLAIR-A) is only cited as related work, not used to justify the shortcut conclusion. No uniqueness theorem, no ansatz-via-citation, and no renaming of a known result as a new derivation appears. The capability-dependent interpretation (§V-B) is a post-hoc explanation of measured accuracy patterns, not an input to the measurements. Overall, the derivation chain is self-contained: the experimental conditions are designed to expose copying, not to produce a predetermined outcome.

Axiom & Free-Parameter Ledger

0 free parameters · 3 axioms · 0 invented entities

No fitted parameters or invented entities. The audit rests on normative premises about judge behavior and on a two-attribute inference to separate protocol-level from capability-dependent shortcuts; these are domain assumptions rather than free parameters.

axioms (3)
  • domain assumption A trustworthy judge should treat audio as ground truth and protocol side information as at most a hint; deliberately wrong side information should be ignored.
    Stated in §III-A; this normative criterion defines what counts as a shortcut in all three RQs. If a judge is expected to follow supplied expert/reference labels, the counter-conditions would measure instruction-following rather than validity.
  • domain assumption The two attributes (emotion, language) suffice to separate protocol-level from capability-dependent shortcuts, with audio-only accuracy as the capability measure.
    Used in §V-B/C to attribute wrong-label copying to missing acoustic capability. K=8 vs K=4, different corpora, and specialist accuracy 0.80 vs 1.00 are not controlled.
  • domain assumption The wrong-label permutation and blueprint templates do not accidentally cue the correct answer or co-vary with clip identity.
    The fixed permutation ensures no class maps to itself, but the paper provides no check that the generated wrong labels are balanced across true labels or that the specialist block's phrasing does not interact with prompt position effects.

pith-pipeline@v1.3.0-alltime-deepseek · 12231 in / 12277 out tokens · 127168 ms · 2026-08-02T05:03:00.484217+00:00 · methodology

0 comments
read the original abstract

Large audio-language models (LALMs) are increasingly used as automatic judges for speech evaluation. However, high agreement with human ratings does not guarantee that their verdicts are grounded in the audio. A judge may instead rely on specialist labels or reference data supplied by the evaluation protocol itself, taking a shortcut in place of listening to the audio. In this paper, we audit such protocol-level ``shortcuts'' in LALM judges across three common deployment protocols: feature-blueprint judging, where the audio is replaced by a structured text description of acoustic features, reference-conditioned judging, and pairwise A/B comparison. Across six judges and four attributes, we find that several LALMs rely on protocol-level shortcuts. For example, in feature-blueprint judging, incorrect specialist labels reduce five judges' emotion accuracy to 0.10 or below, and in concatenated A/B comparisons, Qwen3-Omni-Thinking often picks the same slot regardless of order swaps. These results indicate that aggregate agreement can overstate the validity of LALM judges unless the model and the evaluation protocol are assessed jointly, and that each model-protocol pair should be evaluated with a matched shortcut probe.

Figures

Figures reproduced from arXiv: 2607.13477 by David M. Chan, Hiroshi Saruwatari, Joonyong Park, Yuki Saito.

Figure 1
Figure 1. Figure 1: Overview of protocol-level shortcut audits for LALM judges: feature-blueprint copying, reference-label following, and pairwise slot or format bias. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

34 extracted references · 7 linked inside Pith

  1. [1]

    Qwen2-audio technical report,

    Y . Chu, J. Xu, Q. Yang, H. Wei, X. Wei, Z. Guo, Y . Leng, Y . Lv, J. He, J. Lin, C. Zhou, and J. Zhou, “Qwen2-audio technical report,”arXiv preprint arXiv:2407.10759, 2024

  2. [2]

    Qwen2.5-omni technical report,

    J. Xu, Z. Guo, J. He, H. Hu, T. He, S. Bai, K. Chen, J. Wang, F. Yang, K. Danget al., “Qwen2.5-omni technical report,”arXiv preprint arXiv:2503.20215, 2025

  3. [3]

    Qwen3-omni technical report,

    J. Xu, Z. Guo, H. Hu, Y . Chu, X. Wang, J. He, Y . Wang, X. Shi, T. He, X. Zhuet al., “Qwen3-omni technical report,”arXiv preprint arXiv:2509.17765, 2025

  4. [4]

    Audio flamingo 3: Advancing audio intelligence with fully open large audio language models,

    A. Goel, S. Ghosh, J. Kim, S. Kumar, Z. Kong, S.-g. Lee, C.-H. H. Yang, R. Duraiswami, D. Manocha, R. Valle, and B. Catanzaro, “Audio flamingo 3: Advancing audio intelligence with fully open large audio language models,” inProceedings of Advances in Neural Information Processing Systems (NeurIPS), 2025, spotlight

  5. [5]

    V oxtral,

    A. H. Liuet al., “V oxtral,”arXiv preprint arXiv:2507.13264, 2025

  6. [6]

    Judging LLM-as-a-judge with MT-Bench and chatbot arena,

    L. Zheng, W.-L. Chiang, Y . Sheng, S. Zhuang, Z. Wu, Y . Zhuang, Z. Lin, Z. Li, D. Li, E. P. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica, “Judging LLM-as-a-judge with MT-Bench and chatbot arena,” inProceedings of Advances in Neural Information Processing Systems (NeurIPS), vol. 36, 2023, pp. 46 595–46 623

  7. [7]

    CLAIR: Evaluating image captions with large language models,

    D. M. Chan, S. Petryk, J. E. Gonzalez, T. Darrell, and J. Canny, “CLAIR: Evaluating image captions with large language models,” inProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2023, pp. 13 638–13 646

  8. [8]

    CLAIR-A: Lever- aging large language models to judge audio captions,

    T.-H. Wu, J. E. Gonzalez, T. Darrell, and D. M. Chan, “CLAIR-A: Lever- aging large language models to judge audio captions,” inProceedings of the IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2025

  9. [9]

    Audio large language models can be descriptive speech quality evaluators,

    C. Chen, Y . Hu, S. Wang, H. Wang, Z. Chen, C. Zhang, C.-H. H. Yang, and E. S. Chng, “Audio large language models can be descriptive speech quality evaluators,” inProceedings of the International Conference on Learning Representations (ICLR), 2025

  10. [10]

    EmergentTTS- Eval: Evaluating TTS models on complex prosodic, expressiveness, and linguistic challenges using model-as-a-judge,

    R. R. Manku, Y . Tang, X. Shi, M. Li, and A. Smola, “EmergentTTS- Eval: Evaluating TTS models on complex prosodic, expressiveness, and linguistic challenges using model-as-a-judge,” inProceedings of Advances in Neural Information Processing Systems (NeurIPS), 2025

  11. [11]

    InstructTTSEval: Benchmarking complex natural-language instruction following in text-to-speech systems,

    K. Huang, Q. Tu, L. Fan, C. Yang, D. Zhang, S. Li, Z. Fei, Q. Cheng, and X. Qiu, “InstructTTSEval: Benchmarking complex natural-language instruction following in text-to-speech systems,”arXiv preprint arXiv:2506.16381, 2025

  12. [12]

    SpeechLLM-as-judges: Towards general and inter- pretable speech quality evaluation,

    H. Wanget al., “SpeechLLM-as-judges: Towards general and inter- pretable speech quality evaluation,” inProceedings of the Annual Meeting of the Association for Computational Linguistics (ACL), 2026

  13. [13]

    AudioJudge: Understanding what works in large audio model based speech evaluation,

    P. Manakul, W. H. Gan, M. J. Ryan, A. S. Khan, W. Sirichotedumrong, K. Pipatanakul, W. Held, and D. Yang, “AudioJudge: Understanding what works in large audio model based speech evaluation,”arXiv preprint arXiv:2507.12705, 2025

  14. [14]

    Hearing between the lines: Unlocking the reasoning power of LLMs for speech evaluation,

    A. Chandra, K. Miller, V . Ravichandran, C. Papayiannis, and V . Saligrama, “Hearing between the lines: Unlocking the reasoning power of LLMs for speech evaluation,” inProceedings of the Findings of the Association for Computational Linguistics: EACL 2026, 2026

  15. [15]

    All that glitters is not audio: Rethinking text priors and audio reliance in audio- language evaluation,

    L. H.-Y . Foo, C.-K. Yang, C.-A. Li, K.-H. Lu, and H.-y. Lee, “All that glitters is not audio: Rethinking text priors and audio reliance in audio- language evaluation,”arXiv preprint arXiv:2604.24401, 2026

  16. [16]

    Do audio LLMs really LISTEN, or just transcribe? measuring lexical vs. acoustic emotion cues reliance,

    J. Chen, Z. Guo, J. Chun, P. Wang, A. Perrault, and M. Elsner, “Do audio LLMs really LISTEN, or just transcribe? measuring lexical vs. acoustic emotion cues reliance,” inProceedings of the Conference of the European Chapter of the Association for Computational Linguistics (EACL), 2026

  17. [17]

    When audio-LLMs don’t listen: A cross-linguistic study of modality arbitration,

    J. Billa, “When audio-LLMs don’t listen: A cross-linguistic study of modality arbitration,”arXiv preprint arXiv:2602.11488, 2026

  18. [18]

    Do audio LLMs listen or read? analyzing and mitigating paralinguistic failures with V oxParadox,

    J. Pang, A. Chaubey, and M. Soleymani, “Do audio LLMs listen or read? analyzing and mitigating paralinguistic failures with V oxParadox,” inProceedings of the International Conference on Machine Learning (ICML), 2026

  19. [19]

    LALM-as-a-Judge: Benchmarking large audio-language models for safety evaluation in multi-turn spoken di- alogues,

    A. Ivry and S. Watanabe, “LALM-as-a-Judge: Benchmarking large audio-language models for safety evaluation in multi-turn spoken di- alogues,” inProceedings of the International Conference on Machine Learning (ICML), 2026

  20. [20]

    Large language models are not fair evaluators,

    P. Wang, L. Li, L. Chen, Z. Cai, D. Zhu, B. Lin, Y . Cao, L. Kong, Q. Liu, T. Liu, and Z. Sui, “Large language models are not fair evaluators,” in Proceedings of the 62nd Annual Meeting of the Association for Com- putational Linguistics (Volume 1: Long Papers). Bangkok, Thailand: Association for Computational Linguistics, 2024, pp. 9440–9450

  21. [21]

    Justice or prejudice? quantifying biases in LLM-as-a-judge,

    J. Ye, Y . Wang, Y . Huang, D. Chen, Q. Zhang, N. Moniz, T. Gao, W. Geyer, C. Huang, P.-Y . Chen, N. V . Chawla, and X. Zhang, “Justice or prejudice? quantifying biases in LLM-as-a-judge,” inProceedings of the International Conference on Learning Representations (ICLR), 2025

  22. [22]

    Towards understanding sycophancy in language models,

    M. Sharma, M. Tong, T. Korbak, D. Duvenaud, A. Askell, S. R. Bowman, E. Durmus, Z. Hatfield-Dodds, S. R. Johnston, S. M. Kravec et al., “Towards understanding sycophancy in language models,” inPro- ceedings of the International Conference on Learning Representations (ICLR), 2024

  23. [23]

    SycEval: Evaluating LLM sycophancy,

    A. Fanouset al., “SycEval: Evaluating LLM sycophancy,” inProceed- ings of the AAAI/ACM Conference on AI, Ethics, and Society (AIES), 2025, pp. 893–900

  24. [24]

    Shortcut learning in deep neural networks,

    R. Geirhos, J.-H. Jacobsen, C. Michaelis, R. Zemel, W. Brendel, M. Bethge, and F. A. Wichmann, “Shortcut learning in deep neural networks,”Nature Machine Intelligence, vol. 2, no. 11, pp. 665–673, 2020

  25. [25]

    The pitfalls of simplicity bias in neural networks,

    H. Shah, K. Tamuly, A. Raghunathan, P. Jain, and P. Netrapalli, “The pitfalls of simplicity bias in neural networks,” inProceedings of Advances in Neural Information Processing Systems (NeurIPS), vol. 33, 2020, pp. 9573–9585

  26. [26]

    Large language models can be lazy learners: Analyze shortcuts in in-context learning,

    R. Tang, D. Kong, L. Huang, and H. Xue, “Large language models can be lazy learners: Analyze shortcuts in in-context learning,” inFindings of the Association for Computational Linguistics: ACL 2023. Toronto, Canada: Association for Computational Linguistics, 2023, pp. 4645– 4657

  27. [27]

    The Geneva minimalistic acoustic parameter set (GeMAPS) for voice research and affective computing,

    F. Eyben, K. R. Scherer, B. W. Schuller, J. Sundberg, E. Andr ´e, C. Busso, L. Y . Devillers, J. Epps, P. Laukka, S. S. Narayanan, and K. P. Truong, “The Geneva minimalistic acoustic parameter set (GeMAPS) for voice research and affective computing,”IEEE Transactions on Affective Computing, vol. 7, no. 2, pp. 190–202, 2016

  28. [28]

    The Ryerson Audio-Visual Database of Emotional Speech and Song (RA VDESS): A dynamic, multimodal set of facial and vocal expressions in North American English,

    S. R. Livingstone and F. A. Russo, “The Ryerson Audio-Visual Database of Emotional Speech and Song (RA VDESS): A dynamic, multimodal set of facial and vocal expressions in North American English,”PLoS ONE, vol. 13, no. 5, p. e0196391, 2018

  29. [29]

    The V oiceMOS Challenge 2022,

    E. Cooper, W.-C. Huang, Y . Tsao, H.-M. Wang, T. Toda, and J. Ya- magishi, “The V oiceMOS Challenge 2022,” inProceedings of INTER- SPEECH, 2022

  30. [30]

    V oxCeleb: A large-scale speaker identification dataset,

    A. Nagrani, J. S. Chung, and A. Zisserman, “V oxCeleb: A large-scale speaker identification dataset,” inProceedings of INTERSPEECH, 2017, pp. 2616–2620

  31. [31]

    emotion2vec: Self-supervised pre-training for speech emotion rep- resentation,

    Z. Ma, Z. Zheng, J. Ye, J. Li, Z. Gao, S. Zhang, and X. Chen, “emotion2vec: Self-supervised pre-training for speech emotion rep- resentation,” inProceedings of the Findings of the Association for Computational Linguistics: ACL 2024, 2024, pp. 15 747–15 760

  32. [32]

    Robust speech recognition via large-scale weak super- vision,

    A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak super- vision,” inProceedings of the International Conference on Machine Learning (ICML), ser. Proceedings of Machine Learning Research, vol

  33. [33]

    ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification,

    B. Desplanques, J. Thienpondt, and K. Demuynck, “ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification,” inProc. Interspeech, 2020, pp. 3830–3834

  34. [202]

    28 492–28 518

    PMLR, 2023, pp. 28 492–28 518