Pith. sign in

REVIEW 3 cited by

AudioMNIST: Exploring Explainable Artificial Intelligence for Audio Analysis on a Simple Benchmark

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1807.03418 v3 pith:V2KNBZ72 submitted 2018-07-09 cs.SD cs.AIcs.LGeess.AS

AudioMNIST: Exploring Explainable Artificial Intelligence for Audio Analysis on a Simple Benchmark

classification cs.SD cs.AIcs.LGeess.AS
keywords audioexplanationsneuralartificialaudibleclassificationdatadigits
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Explainable Artificial Intelligence (XAI) is targeted at understanding how models perform feature selection and derive their classification decisions. This paper explores post-hoc explanations for deep neural networks in the audio domain. Notably, we present a novel Open Source audio dataset consisting of 30,000 audio samples of English spoken digits which we use for classification tasks on spoken digits and speakers' biological sex. We use the popular XAI technique Layer-wise Relevance Propagation (LRP) to identify relevant features for two neural network architectures that process either waveform or spectrogram representations of the data. Based on the relevance scores obtained from LRP, hypotheses about the neural networks' feature selection are derived and subsequently tested through systematic manipulations of the input data. Further, we take a step beyond visual explanations and introduce audible heatmaps. We demonstrate the superior interpretability of audible explanations over visual ones in a human user study.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. TabPFN beyond Tabular Data: Calibration and Accuracy on Multimodal Embeddings

    cs.LG 2026-07 accept novelty 6.0

    TabPFN as a zero-gradient head on frozen multimodal embeddings ranks best on NLL and ECE across 22 820 episodes while matching accuracy in mid-shot, mid-dimension regimes and also fixes miscalibration after fine-tuning.

  2. SpeechGuard: Online Defense against Backdoor Attacks on Speech Recognition Models

    cs.SD 2026-07 reject novelty 5.0

    SpeechGuard combines an SNR-adapted STRIP detector with an autoencoder that learns time-frequency masks to suppress backdoor triggers in speech, but purifier training needs oracle poisoned/clean pairs.

  3. TabPFN beyond Tabular Data: Calibration and Accuracy on Multimodal Embeddings

    cs.LG 2026-07 conditional novelty 5.0

    TabPFN as a training-free head on PCA-reduced frozen multimodal embeddings broadly improves calibration (NLL, ECE) over classical heads, with an accuracy edge only for k≥50 shots and d≤32 features.