{"id":"cb2b708a-92d1-4ddf-9797-72f07d2a806a","arxiv_id":"2412.05167","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"ADU-Bench is a new audio dialogue benchmark with 20,715 items across general talk, twelve skills, nine languages, and four ambiguity types, and it shows current audio-language models perform poorly on prosodic ambiguity.","lead":"This paper introduces ADU-Bench, over 20,000 spoken question-and-answer pairs that test how well audio AI models handle everyday conversations, skills, nine languages, and ambiguity. It finds current models, including GPT-4o, struggle most with spoken math, roleplay, non-English languages, and meanings carried by tone, pauses, and homophones.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The ambiguity dataset's synthetic prosody is unvalidated; the real-vs-synthetic ablation in Section 5.4 covers only general dialogues, not the intonation- and pause-based items where TTS prosody is load-bearing.","rationale":"The paper's central novelty is the ADU-Ambiguity dataset, and the headline finding that LALMs fail on prosodic ambiguity rests on those 1,390 synthetic items. The weakest link is the construct validity of the synthetic prosody: Section A uses SSML <prosody> and <break> tags, and the only real-vs-synthetic ablation in Section 5.4 covers ADU-General, not ADU-Ambiguity. The reader's weakest_assumption identifies exactly this gap, and I agree. The concern is concrete: if the TTS renderings do not reliably convey the intended intonation or pause cues, the measured deficits are artifacts, not genuine failures of audio understanding. The manuscript's manual validation is asserted but not quantified, so it does not resolve the issue. A separate but secondary concern is that the GPT-4 evaluator sees only the textual transcription and may reward hedging responses that list multiple interpretations; Section 5.3 notes GPT-4o often does this, which could inflate its ambiguity score. However, even with such inflation, the ambiguity scores are far below the general scores, so the qualitative conclusion is not overturned. The correct disposition remains the reader's CONDITIONAL: accept pending a real-vs-synthetic test on the ambiguity items and a perception check. No change to the verdict is needed.","tokens_in":26957,"tokens_out":5917,"duration_ms":61385,"concrete_test":"For the 395 intonation-based and 250 pause-based items, record human speakers producing each intended reading (or source real audio from an existing ambiguity corpus such as SD-Eval), keeping transcriptions and references identical. Run all 16 LALMs plus the GPT-4 evaluator on both real and synthetic versions; if mean ambiguity scores shift by more than 1 point or model ranking changes, synthetic prosody is a confound. Separately, run a forced-choice perception test in which human listeners select the intended interpretation from the synthetic audio alone; require at least 95% accuracy to establish that the intended cue is reliably present.","verdict_should_be":"UNCHANGED","load_bearing_attack":"ADU-Ambiguity is the paper's principal novelty, but its 1,390 items are synthetic: Section A generates intonation-based items with SSML <prosody> tags and pause-based items with <break> tags. The claim that LALMs fail on phonetic ambiguity presupposes that these synthetic renderings actually carry the intended prosodic and pause cues in a form LALMs can perceive. Section 5.4's real-vs-synthetic ablation validates only 1,000 ADU-General dialogues ('randomly sample 1,000 real-world audio dialogues and generate synthetic audio from their transcriptions'); it never tests the ambiguity items, where synthetic prosody is load-bearing. If the TTS prosody is exaggerated or unnatural, low ambiguity scores could reflect a TTS artifact rather than an inability to handle real speech. The manual validation in Section A is asserted but not quantified (no accuracy, sample size, or inter-annotator agreement), so it does not establish that the cues are robust for evaluating models.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes ADU-Bench, a benchmark for evaluating open-ended audio dialogue understanding of large audio-language models (LALMs). The benchmark comprises four datasets: ADU-General (12,000 items across helpful questions, daily questions, and daily statements), ADU-Skill (3,725 items across 12 domains), ADU-Multilingual (3,600 items in 9 languages), and ADU-Ambiguity (1,390 items covering intonation-, pause-, homophone-, and repetition-based ambiguity). The evaluation uses GPT-4 as a text-based judge that scores model responses against references, with a second scoring pass after swapping reference and response positions to mitigate position bias, plus human evaluation alignment studies. The paper benchmarks 16 LALMs and reports that existing systems struggle with mathematical symbols and formulas, human behavior such as roleplay, multilingual comprehension, and ambiguity arising from phonetic elements.","tokens_in":27124,"tokens_out":5166,"duration_ms":53256,"significance":"If the ambiguity results are valid, ADU-Bench fills a real gap in LALM evaluation, providing a broad and much-needed instrument with over 20,000 open-ended audio dialogues. The strengths are the scale and diversity of the data, the systematic treatment of position bias through double scoring, the inclusion of multiple LLM judges and human-alignment checks, and the public release of code and data. The ambiguity dataset, however, is the paper's principal novelty and carries the main risk: its synthetic prosody is not quantitatively validated, and the GPT-4 judge never hears the audio, which weakens the causal claim that low ambiguity scores reflect failures of audio understanding rather than TTS artifacts or evaluation blind spots.","major_comments":[{"comment":"The construction of the ADU-Ambiguity dataset relies on synthetic audio generated with SSML <prosody> and <break> tags, and Appendix A asserts that 'a manual validation process' was conducted, but no numbers are reported: there is no sample size, no accuracy, and no inter-annotator agreement. Because the entire ambiguity claim presupposes that the synthesized intonation and pause cues are correctly perceived by human listeners as the intended interpretation, the paper should report a perceptual study in which listeners hear only the audio and select the intended meaning, with agreement rates per ambiguity type. Without such validation, the low model scores on intonation- and pause-based items could be a TTS artifact rather than a genuine limitation of LALM audio understanding.","section":"Section 3.2 and Appendix A"},{"comment":"The real-versus-synthetic ablation is run on 1,000 randomly sampled ADU-General dialogues only ('randomly sample 1,000 real-world audio dialogues and generate synthetic audio from their transcriptions'). This provides no evidence for the claim that 'both real-world audio and synthetic audio can effectively serve as evaluation sources' for the ADU-Ambiguity dataset, where synthetic prosody is the load-bearing variable. An analogous validation for ambiguity items, or at least a comparison of real and synthetic renditions of the same ambiguous sentences, is needed to support the paper's conclusions about phonetic ambiguity.","section":"Section 5.4"},{"comment":"The GPT-4 evaluator receives only the textual transcription of the audio query, which is identical for both interpretations of an ambiguity item. Consequently, the judge cannot verify that the reference corresponds to the prosody actually present in the audio, and responses that hedge by listing multiple interpretations may be over-scored as helpful while responses that commit to the wrong interpretation may be under-scored. This weakens the claim that the scores specifically measure audio understanding for ambiguity. The paper should validate the ambiguity scores against human raters who listen to the audio and judge whether the model response matches the intended meaning, rather than only reporting overall pairwise consistency on a small sample; Table 3's ADU-Ambiguity row is based on only 20 queries and uses pairwise preference, which does not test this audio-reference consistency.","section":"Section 4"},{"comment":"The number of ADU-Skill dialogues is inconsistent: Section 3.2 states the dataset comprises 3,750 audio dialogues, while Table 1 reports 3,725, and the overall total of 20,715 is consistent with 3,725. The discrepancy should be corrected and the totals cross-checked.","section":"Section 3.2 and Table 1"},{"comment":"The paper claims to 'firstly propose the evaluation of ambiguity handling in audio dialogues' (Abstract) and 'firstly analyze the ambiguity within audio dialogues' (Section 3.2). This claim of firstness is contradicted by the cited SD-Eval (Ao et al., 2024), which explicitly evaluates spoken dialogue understanding beyond words, including phenomena related to prosody and intonation. The authors should either remove the firstness claim or clearly delineate the specific difference between ADU-Ambiguity and SD-Eval's coverage.","section":"Abstract and Section 3.2"}],"minor_comments":[{"comment":"The rendered figure in the PDF appears to contain duplicated subfigures, repeated 'strategy' axis labels, and 'Loading [MathJax]' artifacts; the final figure should be cleaned and each panel should be labeled and referenced unambiguously.","section":"Figure 2"},{"comment":"The sentence 'BLSP stands out with the highest average score of 3.85 among all LALMs' is ambiguous because Step-Audio-Chat (5.21) and the cascaded models score higher; the authors should specify that this is among the open-sourced end-to-end LALMs or otherwise qualify the comparison.","section":"Section 5.2"},{"comment":"The caption 'Association between human judgment and each dataset in ADU-Bench of GPT-4 evaluation' is unclear; the columns are model-pair comparisons, so the caption should state that each entry is the percentage of pairwise judgments in which the GPT-4 evaluator agrees with human preference for that model pair and dataset.","section":"Table 3"},{"comment":"The Limitations section mentions only the limited number of LALMs; given the load-bearing role of synthetic prosody in the ambiguity dataset, the limitations should also acknowledge the need for perceptual validation of the synthesized ambiguity cues.","section":"Limitations"},{"comment":"In the human evaluation details, the sentence 'we carefully consider the ethical aspects and potential risks associated with the research involving human subjects' begins with a lowercase 'we' after a previous sentence ending in a period; this capitalization and sentence flow should be corrected.","section":"Appendix D"}],"recommendation":"major_revision","confidential_remarks":"This is a solid and useful benchmark contribution with broad coverage and a clean evaluation pipeline, but the principal novelty (the ambiguity dataset) rests on synthetic prosody whose validity is asserted but not demonstrated, and the evaluation judge is blind to the audio signal. Both issues are fixable within the scope of the manuscript through a perceptual validation study and by reporting ambiguity-specific human-listener agreement. The 'firstly' claim for ambiguity evaluation also needs to be moderated in light of SD-Eval. For a journal venue, the fit is reasonable if the benchmark's scope and novelty are carefully positioned against existing spoken dialogue understanding benchmarks."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take. ADU-Bench is a solid, useful benchmark contribution. The new bit is the ambiguity dataset—four types (intonation, pause, homophone, repetition) that share the same transcription but differ in phonetic cues. That fills a real gap, and the empirical result that most LALMs score near zero on intonation-based ambiguity is genuinely new and worth knowing. The rest is standard but competently done: 20k dialogues, 16 models, double-scored GPT-4 evaluation with position-swap, plus human alignment above 85% and two additional judge LLMs as checks. The paper is honest about its construction and doesn't oversell.\n\nThe soft spot is exactly where the novelty lives. The ambiguity items are synthetic, generated with SSML prosody and break tags. The real-vs-synthetic ablation in Section 5.4 only samples 1,000 general dialogues; it never touches the ambiguity items. So the claim that LALMs can't perceive intonational or pause-based meaning depends on the TTS actually conveying those cues in a natural, model-accessible way. The manual validation in Appendix A is asserted but not quantified—no sample size, no accuracy, no agreement. If the synthetic prosody is exaggerated or unnatural, the low scores could be a TTS artifact. That's not fatal, but the headline result needs that fixed.\n\nTwo smaller things. The GPT-4 evaluator only sees transcriptions, so it can't independently verify prosodic perception; it judges the response against a reference that encodes the intended meaning, which is reasonable but indirect. And the per-domain sample sizes are very unbalanced (Writing 40, Roleplay 20, Finance 60), so those per-domain conclusions are thin. The same-family issue of GPT-4 generating references and scoring responses is a mild concern, mitigated by the human alignment checks.\n\nWho gets value: anyone benchmarking audio-language models, especially for speech prosody and multilingual understanding. It deserves a serious referee. I'd ask for a quantified human listening study on the ambiguity subset, and ideally an ablation that checks whether models can at least distinguish the two prosodic readings in a forced-choice setup. With that, it could become a standard reference.","headline":"A genuinely useful audio dialogue benchmark with a novel ambiguity axis, but the central prosody result rests on synthetic audio that is only validated on the general split, not the ambiguity items.","tokens_in":27645,"tokens_out":2706,"would_cite":true,"duration_ms":26734,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper presents ADU-Bench, a 20,715-dialogue benchmark for open-ended audio dialogue understanding, and uses it to show that current large audio-language models systematically fail at math-heavy content, roleplay and common sense…","keywords":["audio dialogue understanding","large audio-language models","ADU-Bench","spoken ambiguity","intonation and pause","multilingual dialogue","LLM-as-judge","synthetic speech benchmark"],"falsifier":"Have human listeners mark the intended meaning of ADU-Ambiguity audio without seeing the reference; if agreement is low, or if re-recording the same items with human voice actors reverses model rankings, the claim that LALMs specifically fail on phonetic ambiguity would not be supported.","tokens_in":26761,"feed_emoji":"🎙️","tokens_out":7305,"duration_ms":65856,"temperature":0.7,"pith_summary":"The paper's aim is to give the field of large audio-language models (LALMs) a measuring instrument for what it is that users actually do: hold open-ended spoken conversations. The authors construct ADU-Bench, four datasets totaling 20,715 audio dialogues, spanning 3 general scenarios, 12 skill domains, 9 languages, and 4 types of phonetic ambiguity. Running 16 LALMs through it, they find the same sentence said with a different intonation or a different pause placement is systematically misread, mathematics and code suffer when spoken aloud, non-English languages drop sharply, and even the strongest model reaches only moderate scores on ambiguity. The paper's contribution is the benchmark and the pattern of failures it exposes, which together give developers concrete targets for the next generation of audio dialogue models.","feed_headline":"New audio benchmark exposes AI's blind spots in tone and pause","feed_subtitle":"ADU-Bench: 20,715 spoken dialogues show even GPT-4o misreads intonation, pauses, and non-English speech.","key_machinery":"The central object is ADU-Bench, a collection of (audio query, textual reference) tuples organized into four datasets: ADU-General (12,000 dialogues), ADU-Skill (3,725), ADU-Multilingual (3,600), and ADU-Ambiguity (1,390). The mechanism that carries the argument is a two-stage evaluation pipeline: first, LALMs receive the audio and produce a response; then a GPT-4 evaluator scores that response against a GPT-4-generated reference on a 0–10 scale, using the transcriptions as the query. To counter position bias, the pipeline runs the scoring twice with reference and response swapped and averages the two scores, and it cross-validates the judge against LLaMA-3-70B, Qwen-2-72B, and human pairwise comparisons. The ambiguity dataset is built on SSML speech synthesis, with prosody tags for intonation and break tags for pauses, which is what lets the same written sentence become audibly different utterances.","core_discovery":"On the paper's own terms, ADU-Bench establishes that open-ended audio dialogue understanding is a measurable capability distinct from speech recognition and from text instruction-following, and that today's LALMs are far better at fluent conversation than at catching the meaning carried by sound itself. The headline result is that the same literal sentence can express different intentions depending on intonation, pause position, or homophones, and that models—GPT-4o included—miss these cues: GPT-4o averages 8.16 overall but only 6.87 on the ambiguity dataset, while open-source models linger below 4 on a 0–10 scale. The skill results show mathematics, physics, chemistry, and coding score lowest, which the authors attribute to mathematical symbols and formulas being hard to follow in audio; common sense and roleplay are also weak, pointing to missing understanding of human behavior. In multilingual dialogue, English is best for every model and non-Indo-European languages are the worst, indicating that audio encoders and training data are not yet multilingual in practice.","pith_inferences":["The authors do not isolate whether ambiguity failures come from the audio encoder or from the downstream language model; a testable extension is to feed the same LALM the audio and the correct transcript side by side and see whether it then picks the intended meaning.","An implication left implicit is that the ambiguity protocol could be turned into a generation benchmark for synthetic speech evaluation: if TTS systems are graded by whether listeners recover the intended intonation-based meaning, SSML-style control becomes directly measurable.","The real-versus-synthetic validation is run on 1,000 general dialogues; running the same ablation on the ambiguity subset would show whether low scores are phonetic failures or artifacts of synthetic prosody, which is the paper's main unverified load-bearing assumption."],"forward_implications":["If the benchmark is right, model rankings should be reported separately per dataset, since a model that tops general dialogue can still fail ambiguity (for example, GPT-4o drops from above 8 to 6.87 on ambiguity).","Developers get a concrete tuning target: improving perception of intonation contours, pause placement, and homophone disambiguation should move the ambiguity score before it moves general dialogue scores.","The cascade pattern, where Whisper-plus-LLM pipelines beat most end-to-end LALMs, implies that the audio encoder is a primary bottleneck for dialogue understanding, not the language model's reasoning.","The multilingual gap implies that spoken dialogue support for non-English languages likely needs dedicated audio data, not just LLM knowledge, because the same model's base abilities do not transfer to audio in those languages."],"supporting_citations":[{"why":"Supplies GPT-4, the model that generates expected references and acts as the primary evaluator of LALM responses.","marker":"Achiam et al. (2023)"},{"why":"Provides the LLM-as-judge evaluation method and the position-bias concern that motivates double scoring with swapped order.","marker":"Zheng et al. (2023)"},{"why":"Prior audio benchmark (AirBench) whose GPT-4 scoring protocol ADU-Bench follows for evaluating audio comprehension.","marker":"Yang et al. (2024)"},{"why":"Defines SSML, the markup used to synthesize audio with controlled intonation (prosody) and pause (break) tags.","marker":"Taylor and Isard (1997)"},{"why":"Provides the public SSML synthesis service that converts the textual dialogues into the benchmark's synthetic audio.","marker":"Microsoft (2024)"},{"why":"Source for the phonetics and phonology principles behind the four ambiguity types in ADU-Ambiguity.","marker":"McMahon (2002)"},{"why":"Second phonetics/phonology source used to define intonation-, pause-, homophone-, and repetition-based ambiguity examples.","marker":"Carr (2019)"},{"why":"Supplies GSM8K math word problems used to construct the mathematics portion of ADU-Skill.","marker":"Cobbe et al. (2021)"},{"why":"Supplies MATH problems used for the mathematics and formula-heavy items in ADU-Skill.","marker":"Hendrycks et al. (2021)"},{"why":"Common Voice supplies real-world recordings used for daily statements and for the real-vs-synthetic audio ablation.","marker":"Ardila et al. (2019)"}],"fun_headline_variants":["Audio AI misreads tone, pauses, and homophones","Benchmark: GPT-4o fails to catch intent from intonation","Same sentence, different meaning: audio AI misses cues","20k dialogues show audio models struggle with ambiguity","Audio language models weak on math, roleplay, and accents"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark's ambiguity scores assume the synthetic audio really carries the intended intonation and pause cues; the paper validates synthetic audio only on 1,000 general dialogues, not on the ambiguity items where those cues are the whole point.","fun_headline_variants_meta":{"raw":{"variants":["Audio AI misreads tone, pauses, and homophones","Benchmark: GPT-4o fails to catch intent from intonation","Same sentence, different meaning: audio AI misses cues","20k dialogues show audio models struggle with ambiguity","Audio language models weak on math, roleplay, and accents"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000232,"raw_usage":{"total_tokens":1524,"prompt_tokens":1015,"completion_tokens":509,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":631,"completion_tokens_details":{"reasoning_tokens":441}},"tokens_in":631,"tokens_out":509,"duration_ms":5233,"temperature":1.0,"reasoning_tokens":441,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T20:48:52.986224+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have human listeners mark the intended meaning of ADU-Ambiguity audio without seeing the reference; if agreement is low, or if re-recording the same items with human voice actors reverses model rankings, the claim that LALMs specifically fail on phonetic ambiguity would not be supported.","supporting_citations":[],"review_version":1}