Pith. sign in

REVIEW 1 cited by

Gradient-Adjusted Neuron Activation Profiles for Comprehensive Introspection of Convolutional Speech Recognition Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2002.08125 v1 pith:RTBW46TK submitted 2020-02-19 cs.LG cs.SDeess.ASstat.ML

classification cs.LGcs.SDeess.ASstat.ML
keywords gradnapsspeechannsdatarecognitionactivationdeepdifferent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deep Learning based Automatic Speech Recognition (ASR) models are very successful, but hard to interpret. To gain better understanding of how Artificial Neural Networks (ANNs) accomplish their tasks, introspection methods have been proposed. Adapting such techniques from computer vision to speech recognition is not straight-forward, because speech data is more complex and less interpretable than image data. In this work, we introduce Gradient-adjusted Neuron Activation Profiles (GradNAPs) as means to interpret features and representations in Deep Neural Networks. GradNAPs are characteristic responses of ANNs to particular groups of inputs, which incorporate the relevance of neurons for prediction. We show how to utilize GradNAPs to gain insight about how data is processed in ANNs. This includes different ways of visualizing features and clustering of GradNAPs to compare embeddings of different groups of inputs in any layer of a given network. We demonstrate our proposed techniques using a fully-convolutional ASR model.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Echoes of Phonetics: Unveiling Relevant Acoustic Cues for ASR via Feature Attribution

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A modern Conformer ASR model's predictions are tied to vowel formants (F1 and F2), sibilant fricative spectral peaks, and plosive release bursts, per SPES feature attributions on TIMIT.

Pith tools