REVIEW 3 cited by
AudioMNIST: Exploring Explainable Artificial Intelligence for Audio Analysis on a Simple Benchmark
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
AudioMNIST: Exploring Explainable Artificial Intelligence for Audio Analysis on a Simple Benchmark
read the original abstract
Explainable Artificial Intelligence (XAI) is targeted at understanding how models perform feature selection and derive their classification decisions. This paper explores post-hoc explanations for deep neural networks in the audio domain. Notably, we present a novel Open Source audio dataset consisting of 30,000 audio samples of English spoken digits which we use for classification tasks on spoken digits and speakers' biological sex. We use the popular XAI technique Layer-wise Relevance Propagation (LRP) to identify relevant features for two neural network architectures that process either waveform or spectrogram representations of the data. Based on the relevance scores obtained from LRP, hypotheses about the neural networks' feature selection are derived and subsequently tested through systematic manipulations of the input data. Further, we take a step beyond visual explanations and introduce audible heatmaps. We demonstrate the superior interpretability of audible explanations over visual ones in a human user study.
Forward citations
Cited by 3 Pith papers
-
TabPFN beyond Tabular Data: Calibration and Accuracy on Multimodal Embeddings
TabPFN as a zero-gradient head on frozen multimodal embeddings ranks best on NLL and ECE across 22 820 episodes while matching accuracy in mid-shot, mid-dimension regimes and also fixes miscalibration after fine-tuning.
-
SpeechGuard: Online Defense against Backdoor Attacks on Speech Recognition Models
SpeechGuard combines an SNR-adapted STRIP detector with an autoencoder that learns time-frequency masks to suppress backdoor triggers in speech, but purifier training needs oracle poisoned/clean pairs.
-
TabPFN beyond Tabular Data: Calibration and Accuracy on Multimodal Embeddings
TabPFN as a training-free head on PCA-reduced frozen multimodal embeddings broadly improves calibration (NLL, ECE) over classical heads, with an accuracy edge only for k≥50 shots and d≤32 features.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.