REVIEW 5 cited by
AudioMNIST: Exploring Explainable Artificial Intelligence for Audio Analysis on a Simple Benchmark
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Explainable Artificial Intelligence (XAI) is targeted at understanding how models perform feature selection and derive their classification decisions. This paper explores post-hoc explanations for deep neural networks in the audio domain. Notably, we present a novel Open Source audio dataset consisting of 30,000 audio samples of English spoken digits which we use for classification tasks on spoken digits and speakers' biological sex. We use the popular XAI technique Layer-wise Relevance Propagation (LRP) to identify relevant features for two neural network architectures that process either waveform or spectrogram representations of the data. Based on the relevance scores obtained from LRP, hypotheses about the neural networks' feature selection are derived and subsequently tested through systematic manipulations of the input data. Further, we take a step beyond visual explanations and introduce audible heatmaps. We demonstrate the superior interpretability of audible explanations over visual ones in a human user study.
Forward citations
Cited by 5 Pith papers
-
PASS: Private Attributes Protection with Stochastic Data Substitution
PASS learns a stochastic substitution mapping that drives private-attribute inference to chance level across images, audio, and sensor data while keeping useful attributes mostly intact.
-
TabPFN beyond Tabular Data: Calibration and Accuracy on Multimodal Embeddings
TabPFN as a zero-gradient head on frozen multimodal embeddings ranks best on NLL and ECE across 22 820 episodes while matching accuracy in mid-shot, mid-dimension regimes and also fixes miscalibration after fine-tuning.
-
Kill Two Birds with One Stone! Trajectory enabled Unified Online Detection of Adversarial Examples and Backdoor Attacks
UniGuard detects both adversarial examples and backdoor-triggered inputs at inference time by treating each input's layer-by-layer path as a time series and flagging anomalies.
-
SpeechGuard: Online Defense against Backdoor Attacks on Speech Recognition Models
SpeechGuard combines an SNR-adapted STRIP detector with an autoencoder that learns time-frequency masks to suppress backdoor triggers in speech, but purifier training needs oracle poisoned/clean pairs.
-
Mettle: Meta-Token Learning for Memory-Efficient Audio-Visual Adaptation
Mettle distills frozen transformer layer features into compact meta-tokens via parallel cross-attention and linear projection, cutting training memory dramatically while retaining competitive accuracy on three audio-v...
Discussion (0). Sign in to comment.