REVIEW 3 major objections 3 minor
Dr. SHAP-AV: Decoding Relative Modality Contributions via Shapley Attribution in Audio-Visual Speech Recognition
T0 review · 3 major / 3 minor · reviewed 2026-07-14 · grok-4.5
Pith's one-line read Shapley attribution shows AV speech models keep leaning on audio even when noise ruins it, while still shifting toward vision as SNR falls.
desk verdict Useful AVSR diagnostic package with a real multi-model claim, but every finding rides on SHAP faithfulness we cannot check from the abstract alone. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Dr. SHAP-AV: three Shapley-value analyses (Global SHAP for overall modality balance, Generative SHAP for contribution dynamics during decoding, and Temporal Alignment SHAP for input–output correspondence) that treat audio and visual streams as players in a cooperative game and attribute prediction credit to each modality.
What would settle it
Recompute the same Global/Generative/Temporal SHAP attributions with alternate baselines or finer feature coalitions; if the reported audio bias and SNR-driven shift disappear or reverse, the central claim about modality contributions is undermined.
Extended reading notes
Core claim
Across six AVSR models and two benchmarks, Shapley-based attribution shows that models shift toward visual reliance as noise rises, yet continue to assign high credit to audio even under severe degradation; SNR dominates modality weighting, modality balance evolves during decoding, and temporal input–output alignment holds under noise, revealing a persistent audio bias.
Load-bearing premise
That Shapley values computed over the paper’s chosen modality coalitions and baselines faithfully measure how much the model truly uses audio versus vision, rather than reflecting baseline or coalition artifacts.
Editorial extensions
If this is right
- AVSR systems need explicit ad-hoc modality-weighting mechanisms rather than relying on fusion alone to correct residual audio bias under noise.
- Shapley-based modality attribution can be adopted as a standard diagnostic alongside word-error-rate when comparing AVSR models.
- Because SNR dominates weighting, noise-robust training curricula should target the SNR regimes where audio credit remains unjustifiably high.
- Temporal Alignment SHAP remaining stable under noise implies that fusion layers preserve frame-to-token correspondence even when overall modality balance shifts.
Reading between the lines
- If the audio bias is architectural rather than purely data-driven, simply adding more noisy training data may not close the gap without explicit re-weighting.
- The same three SHAP views could be applied to other multimodal sequence models (e.g., audio-visual translation or emotion recognition) to test whether residual primary-modality bias is a general pattern.
- A practical next measurement would track whether models that already implement learned modality gates show lower residual audio SHAP under the same severe-SNR conditions.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript proposes Dr. SHAP-AV, a Shapley-value framework for quantifying relative audio and visual contributions in AVSR. From the abstract, the authors apply three analyses—Global SHAP (overall modality balance), Generative SHAP (contribution dynamics during decoding), and Temporal Alignment SHAP (input–output correspondence)—to six models on two benchmarks across SNR levels. They report that models shift toward visual reliance under noise yet retain high audio contributions even under severe degradation, that modality balance evolves during generation, that temporal alignment holds under noise, and that SNR is the dominant driver of modality weighting, concluding a persistent audio bias that motivates ad-hoc modality weighting and Shapley-based diagnostics as standard AVSR practice.
Significance. If the attributions are faithful and the multi-model results hold, the work would supply a reusable diagnostic suite for AVSR and document a systematic audio bias with practical implications for fusion design and training. The three complementary SHAP views and the multi-model, multi-benchmark, multi-SNR design are strengths on paper. Significance is contingent on establishing that the cooperative-game credit assignment tracks causal modality use rather than baseline- or coalition-specific artifacts; without that, the reported bias and SNR dominance remain uninterpretable as model behavior.
major comments (3)
- [Abstract] The central empirical claims (high audio contribution under severe degradation; persistent audio bias; SNR as dominant driver) are measured exclusively by Global/Generative/Temporal Alignment SHAP. The abstract supplies no independent validation that these attributions track causal modality use—e.g., comparison to leave-one-modality-out or occlusion, sensitivity to audio/visual baselines (silence vs. noise vs. mean feature; blank vs. mean face), or checks against gradient-based alternatives. For non-additive AVSR fusion, this faithfulness premise is load-bearing for every finding and must be established with concrete experiments and ablations in the full manuscript.
- [Abstract] Generalization to a ‘persistent audio bias’ in AVSR rests on six models and two benchmarks. The abstract does not identify the models, fusion architectures, or benchmarks, nor does it report error bars, statistical tests, or selection criteria. Without those details and supporting statistics, it is not possible to judge whether the sample supports the general claim rather than model- or dataset-specific behavior.
- [Abstract] Coalition structure, value function, and baseline choices for modality-level (or feature-level) SHAP are not specified. These choices critically affect SHAP values under feature dependence and non-additive fusion typical of AVSR; they must be stated, justified, and subjected to sensitivity analysis before the reported high audio contributions under severe degradation can be interpreted as model behavior rather than attribution artifacts.
minor comments (3)
- [Abstract] The name ‘Dr. SHAP-AV’ is not expanded or motivated; a brief expansion would aid readability.
- [Abstract] ‘Ad-hoc modality-weighting mechanisms’ are invoked as a motivation but not defined; a short clarification of what is meant would help.
- [Abstract] The three SHAP variants are named but their precise input–output scopes (e.g., what is the value function for Generative vs. Temporal Alignment SHAP) remain opaque from the abstract alone; clearer one-sentence definitions would improve accessibility.
Circularity Check
No significant circularity: empirical SHAP diagnostics of existing AVSR models, not a self-referential derivation.
full rationale
Abstract-only review of an analysis paper that applies Shapley-value attribution (Global, Generative, Temporal Alignment SHAP) to six existing AVSR models on two benchmarks under varying SNR. The claimed findings—shift toward visual reliance under noise, persistent high audio contribution, SNR as dominant driver, temporal alignment holding—are reported measurements of model behavior under a chosen attribution method, not predictions recovered from fitted parameters or definitions that equal their inputs by construction. There is no self-definitional loop (X defined via Y then used to derive Y), no fitted constant renamed as a prediction, no load-bearing uniqueness theorem or ansatz imported solely via overlapping-author citation, and no renaming of a known empirical law as a new first-principles result. Method dependence of SHAP (baseline/coalition choices, faithfulness under non-additive fusion) is a validity concern for interpreting the numbers, not circularity of a derivation chain. With no equations or self-citation chain available in the abstract that reduce the central claims to their own inputs, the honest finding is score 0 and empty steps.
Assumptions & free parameters
assumptions (3)
- domain assumption Shapley values over modality (or feature) coalitions fairly attribute prediction credit to audio vs. visual inputs in AVSR models.
- domain assumption SNR-controlled noise on the evaluated benchmarks is a valid proxy for real acoustic degradation relevant to modality reweighting.
- ad hoc to paper Six models and two benchmarks suffice to support claims of a ‘persistent audio bias’ in AVSR generally.
invented entities (1)
-
Dr. SHAP-AV (Global / Generative / Temporal Alignment SHAP suite)
Cite this review
Pith. "Pith review of Dr. SHAP-AV: Decoding Relative Modality Contributions via Shapley Attribution in Audio-Visual Speech Recognition." pith.science (2026). https://pith.science/paper/ALDQSNIM
@misc{pith2026260312046,
author = {Pith},
title = {Pith review of: Dr. SHAP-AV: Decoding Relative Modality Contributions via Shapley Attribution in Audio-Visual Speech Recognition},
year = {2026},
howpublished = {\url{https://pith.science/paper/ALDQSNIM}},
note = {Machine review of arXiv:2603.12046}
}
read the original abstract
Audio-Visual Speech Recognition (AVSR) leverages both acoustic and visual information for robust recognition under noise. However, how models balance these modalities remains unclear. We present Dr. SHAP-AV, a framework using Shapley values to analyze modality contributions in AVSR. Through experiments on six models across two benchmarks and varying SNR levels, we introduce three analyses: Global SHAP for overall modality balance, Generative SHAP for contribution dynamics during decoding, and Temporal Alignment SHAP for input-output correspondence. Our findings reveal that models shift toward visual reliance under noise yet maintain high audio contributions even under severe degradation. Modality balance evolves during generation, temporal alignment holds under noise, and SNR is the dominant factor driving modality weighting. These findings expose a persistent audio bias, motivating ad-hoc modality-weighting mechanisms and Shapley-based attribution as a standard AVSR diagnostic.
Reviewed July 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.