{"id":"029d09db-62a5-497d-8d07-79379dc4723d","arxiv_id":"2508.15407","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Large audio-language models show a systematic bias toward text over audio when the two conflict, according to the new MCR-BENCH benchmark.","lead":"This paper introduces MCR-BENCH, a new benchmark that tests audio-language AI models when audio and text disagree. It finds that these models tend to trust the written text over the audio, which causes errors in audio-focused tasks.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"MCR-BENCH's text-bias metric presupposes audio is the correct modality; the central claim may be an artifact if text is more informative in the benchmark.","rationale":"The reader correctly identified the audio-as-ground-truth assumption as the weakest point. I agree that this is the core issue, but I would further emphasize that the benchmark may not establish audio as the more reliable modality. Without access to the full text, we cannot verify whether the benchmark controls for modality informativeness or item difficulty. The proposed concrete test directly checks whether the bias metric is confounded. Since the reader already marked the paper as UNVERDICTED due to insufficient information, my concern does not change that verdict; it strengthens the need for a specific validation before the central claim can be accepted.","tokens_in":585,"tokens_out":2249,"duration_ms":25423,"concrete_test":"Take a stratified sample of MCR-BENCH items (e.g., 50 per task). For each item, measure (a) audio-only accuracy by ablating the text, (b) text-only accuracy by ablating the audio, and (c) combined accuracy, using a representative LALM and human/oracle annotators. If text-only accuracy >= audio-only accuracy on a substantial fraction of items, rerun the bias metric with accuracy-matched distractors. The 'text bias' claim holds only if models prefer text even when audio is equally or more reliable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that LALMs are biased toward text when audio and text conflict rests entirely on the construct validity of MCR-BENCH. To measure 'bias,' the benchmark must establish that the audio channel alone provides sufficient information for the correct answer, and that the conflicting text is a distractor. The abstract gives no evidence that this ground-truth criterion is met. If in some tasks text is more informative (e.g., contains lexical cues absent from the audio) or the audio is ambiguous, then a model preferring text is rationally optimizing accuracy, not exhibiting bias. Concretely, if text-only baseline accuracy exceeds audio-only accuracy, the reported 'significant bias' could conflate task difficulty with modality preference. Without item-level statistics and proper controls, the observed performance degradation in audio-centric tasks may be a benchmark-construction artifact rather than a genuine model deficiency.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces MCR-BENCH, a benchmark for evaluating how Large Audio-Language Models (LALMs) handle conflicting audio and text inputs. Based on this benchmark, the authors report that LALMs exhibit a significant bias toward textual information, often ignoring audio evidence, which degrades performance on audio-centric tasks. They also investigate factors influencing this bias, explore supervised finetuning as a mitigation, and analyze model confidence patterns showing overconfidence. The project is publicly available.","tokens_in":797,"tokens_out":1797,"duration_ms":21242,"significance":"If the central claim survives scrutiny, this work addresses an under-explored and practically important issue: modality reliability in multi-modal LALMs. The authors identify a systematic failure mode and propose a mitigating approach, which could inform future training and fusion designs. The claimed release of the benchmark is a positive step for reproducibility. However, the significance hinges entirely on the validity of MCR-BENCH as a measure of 'text bias,' which the abstract does not yet establish. The paper's strength is its scope: a first-of-its-kind benchmark covering diverse audio understanding tasks. The main risk is that the observed text preference may be an artifact of benchmark design rather than a genuine model deficiency.","major_comments":[{"comment":"The definition of 'text bias' presupposes that the audio modality is the correct ground truth in audio-centric tasks. To support the central claim, the paper must show that, for each benchmark item, the audio channel alone is sufficient for the correct answer and that the conflicting text is a distractor. If some items contain lexical cues in the text that are absent or ambiguous in the audio, a model preferring text is rationally optimizing accuracy, not exhibiting bias. The abstract provides no item-level statistics (e.g., audio-only vs. text-only baseline accuracies) to rule out this confound. This is a load-bearing point: without such controls, the reported 'significant bias' could be a benchmark-construction artifact.","section":"Abstract"},{"comment":"The abstract reports a 'significant bias' and 'substantial performance degradation' but offers no quantitative results, effect sizes, confidence intervals, or significance tests. The strength of the central claim is unclear; it could range from a small average shift to a large systematic effect. The manuscript should report aggregate bias metrics, per-model and per-task breakdowns, and statistical comparisons against control conditions (e.g., consistent audio-text pairs) to demonstrate that the degradation is due specifically to conflict.","section":"Abstract"},{"comment":"The abstract makes a general claim about 'LALMs' plural, yet provides no information about the selection and coverage of models. How many models were evaluated, what architectures, what size range, and what training corpora? Without showing that the set spans the current LALM landscape, the claim that 'LALMs display a significant bias' cannot be generalized. The benchmark should also report performance on a non-conflicting baseline to show that the models are competent on the audio tasks in the first place.","section":"Abstract"}],"minor_comments":[{"comment":"The term 'audio-centric tasks' is undefined. Please specify the task taxonomy and provide a concrete definition of what makes a task audio-centric versus text-centric.","section":"Abstract"},{"comment":"The abstract claims 'first comprehensive benchmark' without contrasting with or citing prior benchmarks for audio-text conflict or modality bias. A brief positioning statement is needed.","section":"Abstract"},{"comment":"The sentence about 'model confidence patterns that reveal persistent overconfidence even with contradictory inputs' is intriguing but is not tied to any metric or analysis in the abstract. Please clarify what confidence measure is used and how overconfidence is operationalized.","section":"Abstract"}],"recommendation":"uncertain","confidential_remarks":"This review is based solely on the abstract because the full text was not provided to the referee. The stress-test concern about construct validity is real and directly lands on the abstract's phrasing. However, the full manuscript may well include the necessary item-level controls and model details. The 'uncertain' recommendation reflects the lack of evidence to verify the central claim, not a judgment that the work is likely incorrect. I recommend obtaining the full text before a final editorial decision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is worth a serious look, but the abstract alone can't carry the claim. The headline result—LALMs trust text over audio when the two conflict—is exactly the kind of thing the field needs to know, and the authors have built a dedicated benchmark (MCR-BENCH) to test it. That alone is a real contribution: a reusable evaluation set, plus work on mitigation and confidence calibration, and code release. I have no reason to doubt the effort.\n\nThe soft spot is the construct validity of the 'text bias' measure. The stress-test note is right: to call it bias, the benchmark has to ensure that audio alone carries the correct answer and that the conflicting text is a distractor. If the text contains extra cues, or the audio is genuinely ambiguous, then a model that follows the text is not biased—it's just optimizing accuracy. The abstract gives no item-level statistics, no audio-only vs text-only baselines, no human judgments. Without those, the 'substantial performance degradation' could be a benchmark artifact. That's a load-bearing concern, not a nitpick. The circularity worry is minor by comparison: the benchmark looks independent of model outputs, and nothing in the abstract suggests the design pre-determines the finding.\n\nI also can't verify the 'first comprehensive benchmark' claim from the abstract, but that's a literature-review issue, not a fatal one.\n\nBottom line: the idea is important, the benchmark is the kind of thing the community can reuse, and the paper deserves peer review—but a referee should insist on the control analyses that establish audio as the ground-truth modality. I would not cite it yet, and I wouldn't bring it to reading group until the full text is out. If the control analyses hold up, this is a solid 6-ish contribution; if they don't, it's a cautionary tale about measuring modality bias.\n\nMy recommendation: send it to review, with a specific request to check the audio-correctness assumption and the item-level statistics.","headline":"Plausible and important claim about text bias in LALMs, but the abstract can't verify it; the key is whether the benchmark's ground truth truly makes audio the correct answer.","tokens_in":1120,"tokens_out":1683,"would_cite":false,"duration_ms":18271,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Audio-language models trust text over sound in conflicts","keywords":["large audio-language models","modality bias","audio-text conflict","benchmark","multimodal robustness","model confidence","supervised finetuning"],"falsifier":"An experiment that would settle the claim: run MCR-BENCH with the textual input replaced by a clearly irrelevant label; if a model's accuracy stays as high as with meaningful conflicting text, then the model is not actually using the text content, so the apparent bias is not a text preference. Conversely, if a model trained without conflict examples already follows audio as often as text, the reported systematic bias would not be confirmed.","tokens_in":577,"feed_emoji":"🎧","tokens_out":3045,"duration_ms":33189,"temperature":0.7,"pith_summary":"This paper asks what happens to large audio-language models when the audio and text they receive tell different stories. To answer it, the authors build MCR-BENCH, a benchmark of inconsistent audio-text pairs across audio-understanding tasks, and show that the models systematically side with the text, often discarding the audio even when the audio carries the task's true answer. This text bias degrades accuracy in audio-centric tasks and persists even when the text is clearly wrong. The finding matters because real-world audio assistants regularly face noisy transcripts and conflicting cues, and current models appear poorly balanced between the two modalities.","feed_headline":"Audio-language models trust text over sound in conflicts","feed_subtitle":"New benchmark MCR-BENCH shows they routinely ignore audio evidence in audio-centric tasks.","key_machinery":"The central object is MCR-BENCH, a benchmark built from inconsistent audio-text pairs. For each task the audio is the ground-truth signal, and the text is altered so that it disagrees; the benchmark then measures whether a model follows the audio or the text. This controlled disagreement makes the text bias visible and quantifiable.","core_discovery":"The paper's central discovery is a systematic modality imbalance: when audio and text conflict, large audio-language models exhibit a marked preference for textual input, frequently overriding evidence from the audio. This is established through MCR-BENCH, which pairs each audio clip with a mismatched textual transcript or description, and measures performance on audio-centric tasks. Across models and tasks, the authors find that the text channel dominates, leading to large accuracy drops compared with consistent audio-text input. The paper also reports that the bias is affected by factors such as model size and instruction phrasing, and that supervised finetuning on conflict examples can pa","pith_inferences":["The pattern likely extends beyond audio-language models: vision-language models are known to lean on language priors, and MCR-BENCH's design suggests a general 'text-first' inductive bias in multimodal transformers that share a language backbone.","A natural extension is adversarial editing of the text: instead of random mismatches, use plausible but misleading transcripts to measure how far the text bias can be exploited in real-world scenarios.","Because the paper finds finetuning on conflicts helps, a testable prediction is that larger-scale versions of such finetuning will reduce text bias but possibly at the cost of performance on consistent inputs, pointing to a trade-off.","The benchmark's metric could be refined to separate 'text bias' from 'task difficulty' by comparing accuracy with random text labels, giving a clean baseline for how much information the text channel actually contributes."],"forward_implications":["If LALMs systematically prefer text, real-world deployment on audio-first tasks such as speaker identification, emotion recognition, or acoustic scene analysis will be unreliable whenever an automated transcript or caption injects even slightly wrong text.","Training procedures should include explicit conflict examples so models learn to arbitrate between modalities rather than default to text.","Model confidence scores cannot be trusted when modalities conflict, since the models remain overconfident even on wrong, text-led answers.","Evaluation suites for multimodal models should routinely include inconsistent-modality tests, not just consistent ones, to expose modality bias.","Supervised finetuning on conflicting pairs offers a partial mitigation path, suggesting that targeted data can shift the balance toward audio."],"supporting_citations":[],"fun_headline_variants":["Audio models ignore audio when text disagrees","LALMs show text bias in audio-text conflicts","When audio and text clash, AI leans on text","New benchmark reveals text wins over sound in LALMs","MCR-BENCH: models trust text over sound in clashes"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The benchmark counts any preference for text over audio as bias, assuming the audio is always the correct source in audio-centric tasks; if some of those tasks are actually better answered from the text, the measured bias overstates the model's fault.","fun_headline_variants_meta":{"raw":{"variants":["Audio models ignore audio when text disagrees","LALMs show text bias in audio-text conflicts","When audio and text clash, AI leans on text","New benchmark reveals text wins over sound in LALMs","MCR-BENCH: models trust text over sound in clashes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000222,"raw_usage":{"total_tokens":1269,"prompt_tokens":703,"completion_tokens":566,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":447,"completion_tokens_details":{"reasoning_tokens":488}},"tokens_in":447,"tokens_out":566,"duration_ms":6505,"temperature":1.0,"reasoning_tokens":488,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T17:53:27.639873+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An experiment that would settle the claim: run MCR-BENCH with the textual input replaced by a clearly irrelevant label; if a model's accuracy stays as high as with meaningful conflicting text, then the model is not actually using the text content, so the apparent bias is not a text preference. Conversely, if a model trained without conflict examples already follows audio as often as text, the reported systematic bias would not be confirmed.","supporting_citations":[],"review_version":1}