REVIEW 5 cited by
audioLIME: Listenable Explanations Using Source Separation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
audioLIME: Listenable Explanations Using Source Separation
read the original abstract
Deep neural networks (DNNs) are successfully applied in a wide variety of music information retrieval (MIR) tasks but their predictions are usually not interpretable. We propose audioLIME, a method based on Local Interpretable Model-agnostic Explanations (LIME) extended by a musical definition of locality. The perturbations used in LIME are created by switching on/off components extracted by source separation which makes our explanations listenable. We validate audioLIME on two different music tagging systems and show that it produces sensible explanations in situations where a competing method cannot.
Forward citations
Cited by 5 Pith papers
-
If It's Good Enough for You, It's Good Enough for Me: Transferability of Audio Sufficiencies across Models
Transferability analysis finds that minimal sufficient signals transfer across audio models at rates varying by task, around 26% for music genre classification, with some deepfake models showing distinct behaviors not...
-
Activation-Deactivation: A General Framework for Robust Post-hoc Explainable AI
ConvAD replaces input occlusion in post-hoc explanation with neuron deactivation in a CNN forward pass, yielding more robust causal explanations with no retraining.
-
Investigating Modality Contribution in Audio LLMs for Music
Adapts MM-SHAP to quantify modality contributions in two Audio LLMs on MuChoMusic, showing text dominance alongside limited audio localization of key events.
-
SIGMA: Saliency-Guided Sparse Mask Attacks for Speech Emotion Recognition
SIGMA applies post-hoc XAI saliency maps to define reusable sparse masks for magnitude-bounded perturbations on self-supervised speech features, evaluated on IEMOCAP and TESS for competitive attack success with explan...
-
Evaluating the Temporal Detection Capability of Integrated Gradients Applied on Sound Classifier
Integrated gradients on a 10-class domestic sound classifier yields 0.39 mean IoU, 0.52 frame F1 and 82.6% Pointing Game accuracy for temporal event detection, approaching weakly and strongly supervised framewise CNN ...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.