Pith. sign in

REVIEW 5 cited by

AudioMNIST: Exploring Explainable Artificial Intelligence for Audio Analysis on a Simple Benchmark

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1807.03418 v3 pith:V2KNBZ72 submitted 2018-07-09 cs.SD cs.AIcs.LGeess.AS

classification cs.SDcs.AIcs.LGeess.AS
keywords audioexplanationsneuralartificialaudibleclassificationdatadigits
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Explainable Artificial Intelligence (XAI) is targeted at understanding how models perform feature selection and derive their classification decisions. This paper explores post-hoc explanations for deep neural networks in the audio domain. Notably, we present a novel Open Source audio dataset consisting of 30,000 audio samples of English spoken digits which we use for classification tasks on spoken digits and speakers' biological sex. We use the popular XAI technique Layer-wise Relevance Propagation (LRP) to identify relevant features for two neural network architectures that process either waveform or spectrogram representations of the data. Based on the relevance scores obtained from LRP, hypotheses about the neural networks' feature selection are derived and subsequently tested through systematic manipulations of the input data. Further, we take a step beyond visual explanations and introduce audible heatmaps. We demonstrate the superior interpretability of audible explanations over visual ones in a human user study.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PASS: Private Attributes Protection with Stochastic Data Substitution

    cs.LG 2025-06 conditional novelty 7.0 of 10

    PASS learns a stochastic substitution mapping that drives private-attribute inference to chance level across images, audio, and sensor data while keeping useful attributes mostly intact.

  2. TabPFN beyond Tabular Data: Calibration and Accuracy on Multimodal Embeddings

    cs.LG 2026-07 conditional novelty 6.0 of 10

    TabPFN as a zero-gradient head on frozen multimodal embeddings ranks best on NLL and ECE across 22 820 episodes while matching accuracy in mid-shot, mid-dimension regimes and also fixes miscalibration after fine-tuning.

  3. Kill Two Birds with One Stone! Trajectory enabled Unified Online Detection of Adversarial Examples and Backdoor Attacks

    cs.CR 2025-06 conditional novelty 6.0 of 10

    UniGuard detects both adversarial examples and backdoor-triggered inputs at inference time by treating each input's layer-by-layer path as a time series and flagging anomalies.

  4. SpeechGuard: Online Defense against Backdoor Attacks on Speech Recognition Models

    cs.SD 2026-07 reject novelty 5.0 of 10

    SpeechGuard combines an SNR-adapted STRIP detector with an autoencoder that learns time-frequency masks to suppress backdoor triggers in speech, but purifier training needs oracle poisoned/clean pairs.

  5. Mettle: Meta-Token Learning for Memory-Efficient Audio-Visual Adaptation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    Mettle distills frozen transformer layer features into compact meta-tokens via parallel cross-attention and linear projection, cutting training memory dramatically while retaining competitive accuracy on three audio-v...

Pith tools