{"id":"32f0dc51-4cc5-47a1-b2c1-912149b0a9d2","arxiv_id":"2603.12046","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Shapley attribution on six AVSR models shows persistent audio bias under noise, with SNR as the main driver of modality weighting and balance shifting during decoding.","lead":"This paper introduces Dr. SHAP-AV, a Shapley-value framework that measures how much audio versus video drives speech recognition models under noise. It matters because it claims current AVSR systems stay audio-biased even when sound is badly degraded, which would change how people design and diagnose multimodal speech systems.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Abstract-only review cannot validate whether SHAP modality attributions faithfully track causal audio/visual use under the paper's coalitions and baselines; that premise underpins every reported finding.","rationale":"The Reader correctly flags the faithfulness of Shapley modality attributions as the weakest assumption and correctly withholds a verdict because only the abstract is available. No stronger internal inconsistency is visible from the abstract alone; the concern is precisely the one the Reader named. An abstract-only stress test cannot raise a new load-bearing objection that would move the verdict, so UNCHANGED / UNVERDICTED remains appropriate. The concrete test above is the minimal check that would settle whether the concern lands once the full paper appears.","tokens_in":1997,"tokens_out":524,"duration_ms":4460,"concrete_test":"Once methods/code appear: on one model and one SNR condition, recompute Global SHAP with at least two alternative baselines (e.g., silence vs. Gaussian noise for audio; mean-face vs. blank for video) and compare against a direct leave-one-modality-out accuracy drop. If SHAP audio/visual ratios reverse or diverge >20% from the ablation ranking, the attribution is baseline-sensitive and the headline bias claim is unreliable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (persistent audio bias under noise; SNR as dominant driver of modality weighting) is entirely measured by Global/Generative/Temporal Alignment SHAP. For those numbers to mean what the abstract asserts, the cooperative-game credit assignment over audio and visual inputs (or modality-level coalitions) must track actual causal modality use rather than model- or baseline-specific artifacts. The abstract supplies no independent validation of that premise: no ablation against known modality ablations or occlusion, no sensitivity analysis of baseline choice (e.g., silence vs. noise vs. mean-feature for audio; blank vs. mean-face for video), no comparison to simpler leave-one-modality-out or gradient-based attributions, and no discussion of how feature dependence or the non-additive fusion layers typical of AVSR affect SHAP faithfulness. Without methods, figures, or code, it is impossible to check whether the reported high audio contributions under severe degradation are real model behavior or an artifact of the chosen value function and baselines. That is the single load-bearing soft spot; everything else (six models, two benchmarks, three SHAP variants) is secondary until faithfulness is established.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The manuscript proposes Dr. SHAP-AV, a Shapley-value framework for quantifying relative audio and visual contributions in AVSR. From the abstract, the authors apply three analyses—Global SHAP (overall modality balance), Generative SHAP (contribution dynamics during decoding), and Temporal Alignment SHAP (input–output correspondence)—to six models on two benchmarks across SNR levels. They report that models shift toward visual reliance under noise yet retain high audio contributions even under severe degradation, that modality balance evolves during generation, that temporal alignment holds under noise, and that SNR is the dominant driver of modality weighting, concluding a persistent audio bias that motivates ad-hoc modality weighting and Shapley-based diagnostics as standard AVSR practice.","tokens_in":2237,"tokens_out":911,"duration_ms":17610,"significance":"If the attributions are faithful and the multi-model results hold, the work would supply a reusable diagnostic suite for AVSR and document a systematic audio bias with practical implications for fusion design and training. The three complementary SHAP views and the multi-model, multi-benchmark, multi-SNR design are strengths on paper. Significance is contingent on establishing that the cooperative-game credit assignment tracks causal modality use rather than baseline- or coalition-specific artifacts; without that, the reported bias and SNR dominance remain uninterpretable as model behavior.","major_comments":[{"comment":"The central empirical claims (high audio contribution under severe degradation; persistent audio bias; SNR as dominant driver) are measured exclusively by Global/Generative/Temporal Alignment SHAP. The abstract supplies no independent validation that these attributions track causal modality use—e.g., comparison to leave-one-modality-out or occlusion, sensitivity to audio/visual baselines (silence vs. noise vs. mean feature; blank vs. mean face), or checks against gradient-based alternatives. For non-additive AVSR fusion, this faithfulness premise is load-bearing for every finding and must be established with concrete experiments and ablations in the full manuscript.","section":"Abstract"},{"comment":"Generalization to a ‘persistent audio bias’ in AVSR rests on six models and two benchmarks. The abstract does not identify the models, fusion architectures, or benchmarks, nor does it report error bars, statistical tests, or selection criteria. Without those details and supporting statistics, it is not possible to judge whether the sample supports the general claim rather than model- or dataset-specific behavior.","section":"Abstract"},{"comment":"Coalition structure, value function, and baseline choices for modality-level (or feature-level) SHAP are not specified. These choices critically affect SHAP values under feature dependence and non-additive fusion typical of AVSR; they must be stated, justified, and subjected to sensitivity analysis before the reported high audio contributions under severe degradation can be interpreted as model behavior rather than attribution artifacts.","section":"Abstract"}],"minor_comments":[{"comment":"The name ‘Dr. SHAP-AV’ is not expanded or motivated; a brief expansion would aid readability.","section":"Abstract"},{"comment":"‘Ad-hoc modality-weighting mechanisms’ are invoked as a motivation but not defined; a short clarification of what is meant would help.","section":"Abstract"},{"comment":"The three SHAP variants are named but their precise input–output scopes (e.g., what is the value function for Generative vs. Temporal Alignment SHAP) remain opaque from the abstract alone; clearer one-sentence definitions would improve accessibility.","section":"Abstract"}],"recommendation":"uncertain","confidential_remarks":"This is an abstract-only review; the full manuscript was not available. A proper accept/reject decision requires the methods, figures, model identities, baseline definitions, and any faithfulness checks. Scope (eess.AS / AVSR diagnostics) appears appropriate. The load-bearing risk is SHAP faithfulness under the paper’s coalitions and baselines—not novelty of applying Shapley values per se. If the full paper already contains leave-one-modality-out / baseline-sensitivity validation, the major comments can be largely discharged; if not, major revision would be the natural next recommendation."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"Punchline: this is a practical diagnostic for AVSR—Shapley modality attribution across models and SNRs—plus an empirical claim that systems keep leaning on audio even under severe degradation. Whether that claim is real depends on whether their SHAP setup tracks causal modality use. We only have the abstract, so that premise is still open.\n\nWhat is actually new is not Shapley itself. It is the systematic application: Global, Generative, and Temporal Alignment SHAP on six models, two benchmarks, and a noise sweep. If the numbers hold, that multi-model picture and the three analysis modes are a solid subfield contribution. Framing a persistent audio bias as something that motivates explicit reweighting is also useful; AVSR people care about exactly this robustness question. Circularity looks low on the face of it—they are measuring existing models with an external attribution tool, not fitting a theory to recover its own inputs.\n\nThe soft spot is the one the stress-test names, and it is load-bearing rather than minor. Every reported finding rides on cooperative-game credit over audio/visual coalitions. From the abstract there is no ablation against leave-one-modality-out or occlusion, no baseline sensitivity (silence vs noise, blank face vs mean face), and no comparison to gradient-style attributions. Non-additive fusion and feature dependence can make SHAP numbers look like “audio bias” when they are partly artifacts of the value function. That does not mean the paper is wrong; it means we cannot yet tell. Secondary gaps—error bars, model identities, statistical tests—also sit behind the full text.\n\nThis is for AVSR and multimodal ASR readers who want robustness diagnostics, not for people hunting a new recognition architecture. I would send it to peer review: the question matters in the subfield and the experimental scope is large enough to deserve referee time, with a clear ask that they validate attribution faithfulness. I would not put it in reading group until methods and figures exist, and I would not cite it yet without the full paper.","headline":"Useful AVSR diagnostic package with a real multi-model claim, but every finding rides on SHAP faithfulness we cannot check from the abstract alone.","tokens_in":2859,"tokens_out":505,"would_cite":false,"duration_ms":11171,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Shapley attribution shows AV speech models keep leaning on audio even when noise ruins it, while still shifting toward vision as SNR falls.","keywords":["audio-visual speech recognition","Shapley values","modality contribution","signal-to-noise ratio","AVSR diagnostics","modality bias","explainable AI"],"falsifier":"Recompute the same Global/Generative/Temporal SHAP attributions with alternate baselines or finer feature coalitions; if the reported audio bias and SNR-driven shift disappear or reverse, the central claim about modality contributions is undermined.","tokens_in":2854,"feed_emoji":"🗣️","tokens_out":809,"duration_ms":7178,"temperature":0.7,"pith_summary":"Audio-visual speech recognition systems are sold as robust because they can fall back on lip reading when sound is noisy, yet it has been unclear how much they actually use each stream. This paper introduces Dr. SHAP-AV, a Shapley-value framework that credits audio versus visual inputs for the model’s predictions, and applies it to six models on two benchmarks across many noise levels. The analyses show a consistent pattern: as signal-to-noise ratio drops, models do increase their reliance on vision, yet audio still retains a large share of the credit even under severe degradation. Contribution balance also changes step by step during decoding, while the temporal match between input frames and output tokens stays surprisingly stable. The practical upshot is that SNR, not architecture quirks alone, is the main driver of modality weighting, and that a stubborn audio bias remains—so the authors argue for explicit modality-weighting fixes and for treating Shapley diagnostics as routine in AVSR evaluation.","feed_headline":"AV speech models still credit audio even when noise ruins it","feed_subtitle":"Shapley analysis of six systems shows vision gains under noise, yet a stubborn audio bias remains","key_machinery":"Dr. SHAP-AV: three Shapley-value analyses (Global SHAP for overall modality balance, Generative SHAP for contribution dynamics during decoding, and Temporal Alignment SHAP for input–output correspondence) that treat audio and visual streams as players in a cooperative game and attribute prediction credit to each modality.","core_discovery":"Across six AVSR models and two benchmarks, Shapley-based attribution shows that models shift toward visual reliance as noise rises, yet continue to assign high credit to audio even under severe degradation; SNR dominates modality weighting, modality balance evolves during decoding, and temporal input–output alignment holds under noise, revealing a persistent audio bias.","pith_inferences":["If the audio bias is architectural rather than purely data-driven, simply adding more noisy training data may not close the gap without explicit re-weighting.","The same three SHAP views could be applied to other multimodal sequence models (e.g., audio-visual translation or emotion recognition) to test whether residual primary-modality bias is a general pattern.","A practical next measurement would track whether models that already implement learned modality gates show lower residual audio SHAP under the same severe-SNR conditions."],"forward_implications":["AVSR systems need explicit ad-hoc modality-weighting mechanisms rather than relying on fusion alone to correct residual audio bias under noise.","Shapley-based modality attribution can be adopted as a standard diagnostic alongside word-error-rate when comparing AVSR models.","Because SNR dominates weighting, noise-robust training curricula should target the SNR regimes where audio credit remains unjustifiably high.","Temporal Alignment SHAP remaining stable under noise implies that fusion layers preserve frame-to-token correspondence even when overall modality balance shifts."],"fun_headline_variants":["AVSR models keep weighting audio even under severe noise","Shapley shows AV systems shift to vision yet retain audio bias","Noise raises visual credit but audio still dominates AVSR","Six models expose persistent audio bias as SNR drops","AV speech systems credit audio highly despite ruined input"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That Shapley values computed over the paper’s chosen modality coalitions and baselines faithfully measure how much the model truly uses audio versus vision, rather than reflecting baseline or coalition artifacts.","fun_headline_variants_meta":{"raw":{"variants":["AVSR models keep weighting audio even under severe noise","Shapley shows AV systems shift to vision yet retain audio bias","Noise raises visual credit but audio still dominates AVSR","Six models expose persistent audio bias as SNR drops","AV speech systems credit audio highly despite ruined input"]},"model":"grok-4.5","effort":"low","cost_usd":0.002802,"raw_usage":{"total_tokens":963,"prompt_tokens":696,"num_sources_used":0,"completion_tokens":60,"cost_in_usd_ticks":28020000,"prompt_tokens_details":{"text_tokens":696,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":207,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":696,"tokens_out":60,"duration_ms":3040,"temperature":1.0,"reasoning_tokens":207,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T22:28:29.727184+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Recompute the same Global/Generative/Temporal SHAP attributions with alternate baselines or finer feature coalitions; if the reported audio bias and SNR-driven shift disappear or reverse, the central claim about modality contributions is undermined.","supporting_citations":[],"review_version":1}