Pith. sign in

REVIEW 3 major objections 3 minor

Dr. SHAP-AV: Decoding Relative Modality Contributions via Shapley Attribution in Audio-Visual Speech Recognition

T0 review · 3 major / 3 minor · reviewed 2026-07-14 · grok-4.5

Pith's one-line read Shapley attribution shows AV speech models keep leaning on audio even when noise ruins it, while still shifting toward vision as SNR falls.

desk verdict Useful AVSR diagnostic package with a real multi-model claim, but every finding rides on SHAP faithfulness we cannot check from the abstract alone. read the letter →

arxiv 2603.12046 v2 pith:ALDQSNIM submitted 2026-03-12 eess.AS cs.CVcs.SD

classification eess.AScs.CVcs.SD
keywords audio-visualspeechrecognitionShapleyvaluesmodalitycontributionsignal-to-noiseratioAVSRdiagnosticsbiasexplainableAI
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Audio-visual speech recognition systems are sold as robust because they can fall back on lip reading when sound is noisy, yet it has been unclear how much they actually use each stream. This paper introduces Dr. SHAP-AV, a Shapley-value framework that credits audio versus visual inputs for the model’s predictions, and applies it to six models on two benchmarks across many noise levels. The analyses show a consistent pattern: as signal-to-noise ratio drops, models do increase their reliance on vision, yet audio still retains a large share of the credit even under severe degradation. Contribution balance also changes step by step during decoding, while the temporal match between input frames and output tokens stays surprisingly stable. The practical upshot is that SNR, not architecture quirks alone, is the main driver of modality weighting, and that a stubborn audio bias remains—so the authors argue for explicit modality-weighting fixes and for treating Shapley diagnostics as routine in AVSR evaluation.

What carries the argument

Dr. SHAP-AV: three Shapley-value analyses (Global SHAP for overall modality balance, Generative SHAP for contribution dynamics during decoding, and Temporal Alignment SHAP for input–output correspondence) that treat audio and visual streams as players in a cooperative game and attribute prediction credit to each modality.

What would settle it

Recompute the same Global/Generative/Temporal SHAP attributions with alternate baselines or finer feature coalitions; if the reported audio bias and SNR-driven shift disappear or reverse, the central claim about modality contributions is undermined.

Watch

Extended reading notes

Core claim

Across six AVSR models and two benchmarks, Shapley-based attribution shows that models shift toward visual reliance as noise rises, yet continue to assign high credit to audio even under severe degradation; SNR dominates modality weighting, modality balance evolves during decoding, and temporal input–output alignment holds under noise, revealing a persistent audio bias.

Load-bearing premise

That Shapley values computed over the paper’s chosen modality coalitions and baselines faithfully measure how much the model truly uses audio versus vision, rather than reflecting baseline or coalition artifacts.

Editorial extensions

If this is right

  • AVSR systems need explicit ad-hoc modality-weighting mechanisms rather than relying on fusion alone to correct residual audio bias under noise.
  • Shapley-based modality attribution can be adopted as a standard diagnostic alongside word-error-rate when comparing AVSR models.
  • Because SNR dominates weighting, noise-robust training curricula should target the SNR regimes where audio credit remains unjustifiably high.
  • Temporal Alignment SHAP remaining stable under noise implies that fusion layers preserve frame-to-token correspondence even when overall modality balance shifts.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the audio bias is architectural rather than purely data-driven, simply adding more noisy training data may not close the gap without explicit re-weighting.
  • The same three SHAP views could be applied to other multimodal sequence models (e.g., audio-visual translation or emotion recognition) to test whether residual primary-modality bias is a general pattern.
  • A practical next measurement would track whether models that already implement learned modality gates show lower residual audio SHAP under the same severe-SNR conditions.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The manuscript proposes Dr. SHAP-AV, a Shapley-value framework for quantifying relative audio and visual contributions in AVSR. From the abstract, the authors apply three analyses—Global SHAP (overall modality balance), Generative SHAP (contribution dynamics during decoding), and Temporal Alignment SHAP (input–output correspondence)—to six models on two benchmarks across SNR levels. They report that models shift toward visual reliance under noise yet retain high audio contributions even under severe degradation, that modality balance evolves during generation, that temporal alignment holds under noise, and that SNR is the dominant driver of modality weighting, concluding a persistent audio bias that motivates ad-hoc modality weighting and Shapley-based diagnostics as standard AVSR practice.

Significance. If the attributions are faithful and the multi-model results hold, the work would supply a reusable diagnostic suite for AVSR and document a systematic audio bias with practical implications for fusion design and training. The three complementary SHAP views and the multi-model, multi-benchmark, multi-SNR design are strengths on paper. Significance is contingent on establishing that the cooperative-game credit assignment tracks causal modality use rather than baseline- or coalition-specific artifacts; without that, the reported bias and SNR dominance remain uninterpretable as model behavior.

major comments (3)
  1. [Abstract] The central empirical claims (high audio contribution under severe degradation; persistent audio bias; SNR as dominant driver) are measured exclusively by Global/Generative/Temporal Alignment SHAP. The abstract supplies no independent validation that these attributions track causal modality use—e.g., comparison to leave-one-modality-out or occlusion, sensitivity to audio/visual baselines (silence vs. noise vs. mean feature; blank vs. mean face), or checks against gradient-based alternatives. For non-additive AVSR fusion, this faithfulness premise is load-bearing for every finding and must be established with concrete experiments and ablations in the full manuscript.
  2. [Abstract] Generalization to a ‘persistent audio bias’ in AVSR rests on six models and two benchmarks. The abstract does not identify the models, fusion architectures, or benchmarks, nor does it report error bars, statistical tests, or selection criteria. Without those details and supporting statistics, it is not possible to judge whether the sample supports the general claim rather than model- or dataset-specific behavior.
  3. [Abstract] Coalition structure, value function, and baseline choices for modality-level (or feature-level) SHAP are not specified. These choices critically affect SHAP values under feature dependence and non-additive fusion typical of AVSR; they must be stated, justified, and subjected to sensitivity analysis before the reported high audio contributions under severe degradation can be interpreted as model behavior rather than attribution artifacts.
minor comments (3)
  1. [Abstract] The name ‘Dr. SHAP-AV’ is not expanded or motivated; a brief expansion would aid readability.
  2. [Abstract] ‘Ad-hoc modality-weighting mechanisms’ are invoked as a motivation but not defined; a short clarification of what is meant would help.
  3. [Abstract] The three SHAP variants are named but their precise input–output scopes (e.g., what is the value function for Generative vs. Temporal Alignment SHAP) remain opaque from the abstract alone; clearer one-sentence definitions would improve accessibility.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: empirical SHAP diagnostics of existing AVSR models, not a self-referential derivation.

full rationale

Abstract-only review of an analysis paper that applies Shapley-value attribution (Global, Generative, Temporal Alignment SHAP) to six existing AVSR models on two benchmarks under varying SNR. The claimed findings—shift toward visual reliance under noise, persistent high audio contribution, SNR as dominant driver, temporal alignment holding—are reported measurements of model behavior under a chosen attribution method, not predictions recovered from fitted parameters or definitions that equal their inputs by construction. There is no self-definitional loop (X defined via Y then used to derive Y), no fitted constant renamed as a prediction, no load-bearing uniqueness theorem or ansatz imported solely via overlapping-author citation, and no renaming of a known empirical law as a new first-principles result. Method dependence of SHAP (baseline/coalition choices, faithfulness under non-additive fusion) is a validity concern for interpreting the numbers, not circularity of a derivation chain. With no equations or self-citation chain available in the abstract that reduce the central claims to their own inputs, the honest finding is score 0 and empty steps.

Assumptions & free parameters 0 free parameters · 3 assumptions · 1 invented entities

Abstract-only ledger: the central empirical claims rest on standard SHAP/cooperative-game assumptions applied to deep AVSR models, plus domain assumptions about SNR as a noise proxy and that the chosen models/benchmarks represent the class. No free numerical constants are stated in the abstract; the main invented construct is the Dr. SHAP-AV analysis suite itself, which is a method rather than a physical entity.

assumptions (3)
  • domain assumption Shapley values over modality (or feature) coalitions fairly attribute prediction credit to audio vs. visual inputs in AVSR models.
    Load-bearing for Global/Generative/Temporal Alignment SHAP; standard in XAI but contested for deep correlated multimodal features.
  • domain assumption SNR-controlled noise on the evaluated benchmarks is a valid proxy for real acoustic degradation relevant to modality reweighting.
    Abstract treats SNR as the dominant driver; generalizing beyond synthetic SNR depends on this.
  • ad hoc to paper Six models and two benchmarks suffice to support claims of a ‘persistent audio bias’ in AVSR generally.
    Scope of induction from the experimental sample to the class of AVSR systems is assumed, not derived.
invented entities (1)
  • Dr. SHAP-AV (Global / Generative / Temporal Alignment SHAP suite)
    purpose: Package Shapley attribution into three AVSR-specific analyses of modality balance, decoding dynamics, and input-output alignment.
    Methodological construct introduced by the paper; independent evidence would be external replications or predictive utility beyond the reported experiments, which the abstract does not supply.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Dr. SHAP-AV: Decoding Relative Modality Contributions via Shapley Attribution in Audio-Visual Speech Recognition." pith.science (2026). https://pith.science/paper/ALDQSNIM

@misc{pith2026260312046,
  author       = {Pith},
  title        = {Pith review of: Dr. SHAP-AV: Decoding Relative Modality Contributions via Shapley Attribution in Audio-Visual Speech Recognition},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ALDQSNIM}},
  note         = {Machine review of arXiv:2603.12046}
}
read the original abstract

Audio-Visual Speech Recognition (AVSR) leverages both acoustic and visual information for robust recognition under noise. However, how models balance these modalities remains unclear. We present Dr. SHAP-AV, a framework using Shapley values to analyze modality contributions in AVSR. Through experiments on six models across two benchmarks and varying SNR levels, we introduce three analyses: Global SHAP for overall modality balance, Generative SHAP for contribution dynamics during decoding, and Temporal Alignment SHAP for input-output correspondence. Our findings reveal that models shift toward visual reliance under noise yet maintain high audio contributions even under severe degradation. Modality balance evolves during generation, temporal alignment holds under noise, and SNR is the dominant factor driving modality weighting. These findings expose a persistent audio bias, motivating ad-hoc modality-weighting mechanisms and Shapley-based attribution as a standard AVSR diagnostic.

Discussion (0). Continue with ORCID to comment.

Pith tools

Reviewed July 14, 2026 · model on record in the stance chip above.