{"id":"e218232d-9ec4-45c8-a2ce-5540c7a4055c","arxiv_id":"2506.07502","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"DEBATE is a first-of-its-kind Mandarin speech-text dataset for studying how prosody resolves textual ambiguity, and current large speech-language models lag well behind humans on it.","lead":"The paper introduces DEBATE, a new Mandarin speech dataset with over 10,000 recordings of 1,001 ambiguous written sentences, each spoken by 10 native speakers to show how pronunciation, pauses, and stress can reveal intended meaning. It also benchmarks three large speech-language models and finds they perform far below human listeners, especially on stress-based disambiguation.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No text-only control in the Section 4 benchmark; the claim that LSLMs fail to use acoustic cues is underdetermined because audio-conditioned scores are never compared with text-only scores.","rationale":"I considered the reader's weakest_assumption, that speakers may not have realized the prescribed prosodic cues. This is a legitimate data-quality risk, and I partially agree. However, even perfect recordings would not save the central benchmark claim without a text-only control. The missing control is more load-bearing because it directly determines whether the observed model scores say anything about acoustic cue processing. It is also easily fixable. In addition, the paper's reliance on CER as evidence of 'high consistency between audio and text' does not validate prosodic cue fidelity, another gap, but secondary. My proposed test would settle the attribution: if text-only equals audio accuracy, the paper must restate its conclusion as a general ambiguity-resolution gap; if audio is significantly better, the current interpretation survives (though human numbers are still needed). I therefore keep the reader's CONDITIONAL verdict unchanged. This is an internal-control issue, not a disagreement with field consensus, and it is presented without questioning the authors' conduct.","tokens_in":10583,"tokens_out":7559,"duration_ms":101546,"concrete_test":"Run the Section 4.2 evaluation on the same 150-item set in three conditions: (a) audio+prompt as reported; (b) text-only (the ambiguous sentence and two options, no audio; use the underlying LLM or a blank/silent audio track for models that require audio input); (c) text+explicit prosodic annotation (stress/pause marks). Report per-task accuracy for the three models. Also report human accuracy per task and Fleiss' kappa among the three evaluators. If audio-conditioned accuracy is not significantly above text-only accuracy for T_Stress (and T_Pause), the 'failure to leverage speech' conclusion is unsupported. Low human agreement would further undermine the gold labels.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that LSLMs 'still struggle to effectively leverage acoustic information' and are far below humans on speech-based disambiguation (Section 4.2, Table 4, Figure 4). All model scores come from an audio-plus-prompt condition; no matched text-only condition is reported. Without such a control, accuracies of 51–68% cannot be attributed to (in)ability to use speech cues. A model might reach the same accuracy from lexical-semantic priors (e.g., '我想起来了' is more often 'I remembered'), or it might actually be harmed by the audio, or it might be using audio cues but still only slightly above chance. The paper also reports no human accuracy values or inter-annotator agreement for the 150-item evaluation, only violin plots, so the claimed 'clear and huge' human-model gap is not quantitatively documented. This is load-bearing because the paper's core contribution is a dataset plus evidence that current models specifically fail at speech-based disambiguation; the evidence as presented cannot distinguish 'fails at speech cues' from 'fails at ambiguous text'.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces DEBATE, a Mandarin speech-text dataset for studying how speech cues (pronunciation, pause, stress, intonation) resolve ambiguity that persists in written text. The dataset contains 1,001 ambiguous sentences, each recorded by 10 native speakers, yielding 10,010 audio samples totaling 9.66 hours, organized into three subtasks: polyphonic-character ambiguity, structural/pause ambiguity, and stress/focus ambiguity. The authors describe a data pipeline involving collection from corpora, social media, and exam banks; LLM-based augmentation with manual verification; a two-person recording protocol; and ASR-based quality checks. They benchmark three large speech-language models (Qwen2-Audio, Qwen2.5-Omni, Gemini 2.0 Flash) in a zero-shot, forced-choice setting and report accuracy and macro-F1 per subtask. They additionally compare models with three human volunteers on a 50-sample-per-subtask subset and report, via violin plots, that humans outperform all models, especially on stress-based disambiguation.","tokens_in":10771,"tokens_out":3791,"duration_ms":52969,"significance":"If the dataset is validated as claimed, DEBATE would be a useful and timely resource: to my knowledge it is the first public Mandarin dataset specifically designed for speech-based disambiguation, and the paper documents the collection pipeline in commendable detail. The release of data and code, the use of multiple speakers with demographic diversity, the two-person recording protocol, and the ASR-based quality checks are all concrete strengths that make the resource potentially reusable for training and evaluation. However, the benchmark evidence for the central claim—that current LSLMs fail to leverage acoustic information—is incomplete: there is no text-only control, the human evaluation is reported only as violin plots with no numeric accuracy or agreement statistics, and the recordings' prosodic validity is not directly established beyond thin manual sampling and ASR lexical alignment. These gaps are fixable within the paper's scope, and the dataset contribution itself remains defensible.","major_comments":[{"comment":"The central empirical claim that LSLMs 'still struggle to effectively leverage acoustic information' is underdetermined by the reported experiments. Table 4 reports only audio-plus-prompt accuracies; no matched text-only condition is shown. Without a text-only baseline (for example, feeding the same forced-choice prompt with the transcript in place of audio to a text LLM, or masking the audio), the accuracies of 51–68% could reflect lexical-semantic priors rather than a failure to use speech cues. A model might choose the statistically more frequent reading of an ambiguous sentence regardless of audio, and the audio could even be slightly harmful. Please add matched text-only baselines and report the audio-minus-text accuracy delta; the conclusion about acoustic information should be revised on that basis.","section":"§4.2, Table 4"},{"comment":"The human–model comparison is not quantitatively documented. Only three volunteers, 50 samples per subtask, are described, with no accuracy numbers, no confidence intervals, and no inter-annotator agreement. The claim of a 'clear and huge' gap cannot be evaluated from a violin plot alone. Please report the numeric human accuracy per subtask and per speaker, the corresponding model accuracies on the same 50-sample set, and an agreement measure such as Fleiss' kappa or pairwise Cohen's kappa. This is load-bearing because the gap between humans and models is a headline result of the paper.","section":"§4.2, Figure 4"},{"comment":"The dataset's validity rests on the assumption that the recruited speakers reliably produced the prescribed prosodic cues (pronunciation, pause position, stress). The ASR CER results in Table 1 verify lexical alignment but not prosodic realization: stress, in particular, typically does not change the character sequence and thus is invisible to CER. The manual check of ten randomly selected samples per task is thin, and no criteria or reliability statistics are reported. Please add direct validation of the prosodic cues—for example, acoustic measurements of pause durations and pitch/intensity contours, or a perception test in which naive listeners select the intended meaning from the audio. Without this, the human–model gap could reflect recording artifacts rather than inherent model limitations.","section":"§3.2.2–3.2.3, Table 1"}],"minor_comments":[{"comment":"There are typographical errors: 'TProun' should be 'TPronun' (or a consistent abbreviation), and 'Gemeni' should be 'Gemini'.","section":"Table 4, Figure 4"},{"comment":"No inter-annotator agreement is reported for the manual selection and semantic annotation of the ambiguous sentences; even a brief description of annotation guidelines and a kappa value would strengthen confidence in the classification into the three subtask types.","section":"§3.2.1"},{"comment":"The CER comparison with AISHELL-1 and AISHELL-2 is not apples-to-apples because DEBATE sentences are deliberately ambiguous and include polyphonic characters; please report CER per speaker and per subtask, and state the ASR decoding configuration used.","section":"§3.2.3, Table 1"},{"comment":"The percentages shown in Figure 1 (e.g., '40%', '20%') are not clearly defined; the text reports 200, 401, and 400 samples per subtask, so the figure should use consistent counts or percentages with explicit labels.","section":"Figure 1, Table 2"},{"comment":"The term 'polyphonic character ambiguity' is nonstandard; consider using 'polyphone ambiguity' or 'heteronym ambiguity' and define the term on first use to avoid confusion with musical polyphony.","section":"§3.1"},{"comment":"The inference setup lacks reproducibility details: please report the decoding parameters, number of runs, and whether the reported variance in Table 4 is across speakers only or across repeated inference runs.","section":"§4.1"}],"recommendation":"major_revision","confidential_remarks":"The dataset contribution is real and the pipeline is described in unusual detail for a resource paper. The main risk is that the headline claim about models' failure to use acoustic information is not yet supported by the experiments, because the missing text-only control and the thin human baseline leave alternative explanations open. These are fixable without changing the dataset itself. I see no novelty disclosure concerns: the related work is appropriately cited, and the 'first dataset' claim appears credible based on the surveyed literature."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The dataset is the real contribution here, and it is a good one: 1,001 ambiguous Mandarin sentences, each spoken by 10 native speakers, over 10k audio clips and 9.66 hours, targeting polyphonic, pause, and stress disambiguation. That fills a genuine gap in the speech-language resource space. The collection pipeline is described carefully — two-person recording sessions, manual verification against annotated prosodic cues, ASR-based quality control with low character error rates. The three ambiguity types are well chosen and the examples in Figure 1 are clear. If you work on Chinese TTS, spoken language understanding, or prosody, this is a resource worth having and worth citing.\n\nNow the soft spots. The benchmark evidence is weaker than the paper's conclusions. The central claim is that large speech-language models 'struggle to effectively leverage acoustic information.' But there is no matched text-only condition anywhere in Section 4. The models hear audio and get a prompt, and their 51–68% accuracies are compared only to a small human baseline. Without a text-only control, you cannot attribute the gap to audio-cue processing. The models might be doing slightly better or worse than they would on the transcript alone; the paper simply does not say. This is load-bearing for the benchmark claim, though not for the dataset itself.\n\nThe human baseline is also thin: three volunteers, 50 samples per subtask, no inter-annotator agreement, no numeric accuracy values in the text (only violin plots), and no chance-level or majority-class reference. For a dataset paper this is forgivable as a sanity check, but it is presented as showing a 'clear and huge' human-model gap, and the evidence is not quantitatively documented to that standard.\n\nThe stress-test note is right: the missing text-only control genuinely underdetermines the 'fails at speech cues' interpretation. That said, I would not hold the dataset hostage to the benchmark. The resource itself is built with care, the examples are linguistically sound, and the limitations section at least acknowledges the narrow model coverage.\n\nVerdict: solid dataset paper, overreaching benchmark conclusions. A serious referee should engage with it, but the benchmark section needs revision — add a text-only condition, expand and properly report the human evaluation, and temper the claims about acoustic-cue failure. I would accept it for review with major revision expected.","headline":"A genuinely useful Mandarin speech-ambiguity dataset, with a benchmark that overclaims its control over text-only priors.","tokens_in":11304,"tokens_out":1435,"would_cite":true,"duration_ms":20095,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The DEBATE dataset pairs 1,001 ambiguous Mandarin utterances with 10,010 recordings from ten native speakers and shows large speech-language models resolve spoken intent far worse than human listeners.","keywords":["speech-based disambiguation","Mandarin Chinese","speech-text dataset","prosody","polyphonic characters","structural ambiguity","stress and intonation","large speech-language models"],"falsifier":"Measure, for every stress item and every speaker, the fundamental-frequency peak, duration, and intensity of the intended stressed syllable relative to the same syllable in the alternative reading, and measure the actual silent intervals at the annotated pause boundaries in the pause items. If a substantial fraction of recordings lack a reliable acoustic difference for the annotated cue, the assumption that DEBATE encodes the targeted prosodic cues fails, and the benchmark numbers would need re-scoring before being read as model ability.","tokens_in":10400,"feed_emoji":"🎙️","tokens_out":8663,"duration_ms":90463,"temperature":0.7,"pith_summary":"The paper sets out to fill a gap: many written Chinese sentences are ambiguous—a character with two readings, a missing word boundary, or a shifted stress can change the meaning—and existing corpora pair such sentences only with text, not with speech. DEBATE pairs 1,001 ambiguous written utterances with 10,010 spoken recordings by ten native speakers, annotating the pronunciation, pause, and stress cues that disambiguate each item. The authors then benchmark three large speech-language models and report that none resolves the intended meaning nearly as well as human listeners, with stress-based ambiguity being the hardest. If the dataset and results hold, they give the field a standard evaluation and training resource for disambiguation through speech, with implications for speech understanding and text-to-speech systems.","feed_headline":"Mandarin speech dataset exposes AI blind spot for ambiguous phrases","feed_subtitle":"Best large speech-language models score 52–68 percent on spoken ambiguity; humans do far better, especially on stress cues.","key_machinery":"The load-bearing object is the DEBATE corpus, organized into three disambiguation scenarios: polyphonic characters, where pronunciation determines meaning; structural ambiguity, where pause boundaries determine sentence segmentation; and focus ambiguity, where stress and intonation determine semantic emphasis. Each item pairs a written sentence that is ambiguous on the page with a spoken recording annotated for the disambiguated meaning and the prosodic cues that carry it. This pairing turns disambiguation through speech into a measurable task: a model hears one audio recording and must choose the intended reading, while the human baseline on the same items quantifies what the acoustic cues actually convey.","core_discovery":"The paper claims that Mandarin's written ambiguity can be systematically disentangled by spoken cues, and introduces DEBATE as the first dataset built to study this. It curates 1,001 ambiguous sentences—200 polyphonic-character, 401 structural, and 400 focus-related—each read by 10 native speakers, yielding 10,010 recordings totaling 9.66 hours with semantic and prosodic annotations. On this corpus, three large speech-language models reach only about 51–68 percent accuracy, markedly below human listeners on the same items, with the largest deficit in stress-based disambiguation. The authors read this as evidence that current models can detect relatively explicit cues like pauses but not fine prosody such as stress and pitch movement.","pith_inferences":["Editorial inference: the paper's three-way taxonomy of pronunciation, pause, and stress suggests a difficulty gradient that may generalize to other languages without explicit word boundaries, so analogous datasets could be built for Cantonese, Japanese, or other isolating languages.","Editorial inference: an acoustic verification study could test whether the recorded stress cues are perceptually consistent across the ten speakers; if they are not, the stress-task numbers would need to be reinterpreted as a measure of recording prompt-following rather than model ability.","Editorial inference: the human–model gap may shrink fastest on pause tasks because the cue is discrete and temporally located, whereas stress tasks may need prosody-aware training signals that are largely absent from current automatically transcribed corpora."],"forward_implications":["Speech-language models have measurable, task-specific weaknesses: pause-based structure is the easiest cue for them, while stress and intonation is the hardest, so future acoustic-modeling work should concentrate on fine prosody.","DEBATE can serve as a common benchmark for disambiguation through speech, letting future models be compared against the same human baseline on the same utterances.","The corpus offers targeted training data for text-to-speech systems, particularly for pronouncing polyphonic characters correctly and rendering prosodic stress naturally.","The implicit regional variation among the ten speakers makes the dataset usable for studying Mandarin pronunciation variants across dialect backgrounds.","If the benchmark is reused, stress-based tasks may require new training objectives or prosody-labeled pretraining data before model performance approaches human levels."],"supporting_citations":[{"why":"supplies the Chinese ambiguity benchmark whose categories and example types inform the design of the DEBATE text corpus.","marker":"[38]"},{"why":"supplies character-level word sense disambiguation material that motivates the polyphonic-character scenario.","marker":"[33]"},{"why":"provides a multilingual word sense disambiguation task cited as groundwork for the text corpus.","marker":"[23]"},{"why":"supplies one of the two ASR systems used to compute character error rates as a quality check on the recordings.","marker":"[26]"},{"why":"supplies the other ASR system used in the character error rate quality check.","marker":"[2]"},{"why":"gives a standard Mandarin speech recognition corpus whose character error rates contextualize the DEBATE quality numbers.","marker":"[6]"},{"why":"gives a second standard Mandarin corpus used to contextualize the ASR quality numbers.","marker":"[14]"},{"why":"provides one of the three large speech-language models benchmarked zero-shot on DEBATE.","marker":"[11]"},{"why":"provides the second benchmarked model, with the strongest pause-task accuracy among the three.","marker":"[35]"},{"why":"provides the third benchmarked model, showing the stress-task deficit in the model–human comparison.","marker":"[12]"}],"fun_headline_variants":["AI stumbles on Mandarin ambiguity that speech easily resolves","DEBATE dataset shows AI misses prosodic cues in Mandarin","New speech dataset shows AI can't untangle ambiguous Mandarin","Humans hear intent, AI doesn't: Mandarin disambiguation gap","First Mandarin speech dataset reveals AI's disambiguation gap"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire benchmark depends on the ten volunteer speakers actually producing the prescribed pronunciations, pause placements, and stress patterns when they read the instructions; if the recorded prosody does not match the annotations, the reported human–model gap could be an artifact of the recordings rather than a measure of disambiguation ability.","fun_headline_variants_meta":{"raw":{"variants":["AI stumbles on Mandarin ambiguity that speech easily resolves","DEBATE dataset shows AI misses prosodic cues in Mandarin","New speech dataset shows AI can't untangle ambiguous Mandarin","Humans hear intent, AI doesn't: Mandarin disambiguation gap","First Mandarin speech dataset reveals AI's disambiguation gap"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000949,"raw_usage":{"total_tokens":4020,"prompt_tokens":886,"completion_tokens":3134,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":502,"completion_tokens_details":{"reasoning_tokens":3052}},"tokens_in":502,"tokens_out":3134,"duration_ms":31227,"temperature":1.0,"reasoning_tokens":3052,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:31:47.981500+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure, for every stress item and every speaker, the fundamental-frequency peak, duration, and intensity of the intended stressed syllable relative to the same syllable in the alternative reading, and measure the actual silent intervals at the annotated pause boundaries in the pause items. If a substantial fraction of recordings lack a reliable acoustic difference for the annotated cue, the assumption that DEBATE encodes the targeted prosodic cues fails, and the benchmark numbers would need re-scoring before being read as model ability.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the Chinese ambiguity benchmark whose categories and example types inform the design of the DEBATE text corpus."},{"cited_title":"Routledge","cited_arxiv_id":null,"evidence_quote":"supplies character-level word sense disambiguation material that motivates the polyphonic-character scenario."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"provides a multilingual word sense disambiguation task cited as groundwork for the text corpus."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies one of the two ASR systems used to compute character error rates as a quality check on the recordings."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the other ASR system used in the character error rate quality check."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"gives a standard Mandarin speech recognition corpus whose character error rates contextualize the DEBATE quality numbers."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"provides the third benchmarked model, showing the stress-task deficit in the model–human comparison."}],"review_version":1}