Pith. sign in

REVIEW 1 cited by

Towards falsifiable interpretability research

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.12016 v1 pith:Y4HOJNDX submitted 2020-10-22 cs.CY cs.AIcs.CVcs.LGstat.ML

classification cs.CYcs.AIcs.CVcs.LGstat.ML
keywords interpretabilityresearchfalsifiablemethodsaddressarguednnsframework
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Methods for understanding the decisions of and mechanisms underlying deep neural networks (DNNs) typically rely on building intuition by emphasizing sensory or semantic features of individual examples. For instance, methods aim to visualize the components of an input which are "important" to a network's decision, or to measure the semantic properties of single neurons. Here, we argue that interpretability research suffers from an over-reliance on intuition-based approaches that risk-and in some cases have caused-illusory progress and misleading conclusions. We identify a set of limitations that we argue impede meaningful progress in interpretability research, and examine two popular classes of interpretability methods-saliency and single-neuron-based approaches-that serve as case studies for how overreliance on intuition and lack of falsifiability can undermine interpretability research. To address these concerns, we propose a strategy to address these impediments in the form of a framework for strongly falsifiable interpretability research. We encourage researchers to use their intuitions as a starting point to develop and test clear, falsifiable hypotheses, and hope that our framework yields robust, evidence-based interpretability methods that generate meaningful advances in our understanding of DNNs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Enhancing Uncertainty Estimation and Interpretability via Bayesian Non-negative Decision Layer

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A Bayesian non-negative decision layer with gamma priors and Weibull variational inference improves uncertainty estimation and interpretability for image classifiers.

Pith tools