Standard feature attribution methods give inconsistent explanations across retrained speech classifiers, except for word-level perturbation on word-based tasks.
On the reliability of feature attribution methods for speech classification
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
As the capabilities of large-scale pre-trained models evolve, understanding the determinants of their outputs becomes more important. Feature attribution aims to reveal which parts of the input elements contribute the most to model outputs. In speech processing, the unique characteristics of the input signal make the application of feature attribution methods challenging. We study how factors such as input type and aggregation and perturbation timespan impact the reliability of standard feature attribution methods, and how these factors interact with characteristics of each classification task. We find that standard approaches to feature attribution are generally unreliable when applied to the speech domain, with the exception of word-aligned perturbation methods when applied to word-based classification tasks.
citation-role summary
citation-polarity summary
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
support 1representative citing papers
citing papers explorer
-
On the reliability of feature attribution methods for speech classification
Standard feature attribution methods give inconsistent explanations across retrained speech classifiers, except for word-level perturbation on word-based tasks.