Pith. sign in

REVIEW 13 cited by

Sanity Checks for Saliency Maps

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1810.03292 v3 pith:KZ63HUC5 submitted 2018-10-08 cs.CV cs.LGstat.ML

Sanity Checks for Saliency Maps

classification cs.CV cs.LGstat.ML
keywords modeldatamethodssaliencyfindingslearnedproposedvisual
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Saliency methods have emerged as a popular tool to highlight features in an input deemed relevant for the prediction of a learned model. Several saliency methods have been proposed, often guided by visual appeal on image data. In this work, we propose an actionable methodology to evaluate what kinds of explanations a given method can and cannot provide. We find that reliance, solely, on visual assessment can be misleading. Through extensive experiments we show that some existing saliency methods are independent both of the model and of the data generating process. Consequently, methods that fail the proposed tests are inadequate for tasks that are sensitive to either data or model, such as, finding outliers in the data, explaining the relationship between inputs and outputs that the model learned, and debugging the model. We interpret our findings through an analogy with edge detection in images, a technique that requires neither training data nor model. Theory in the case of a linear model and a single-layer convolutional neural network supports our experimental findings.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. The Statistical Cost of Adaptation in Multi-Source Transfer Learning

    math.ST 2026-05 unverdicted novelty 8.0

    Multi-source transfer learning incurs an intrinsic adaptation cost that can exceed one, with phase transitions separating regimes where bias-agnostic estimators match oracle performance from those where they cannot.

  2. From Mechanistic to Compositional Interpretability

    cs.LG 2026-05 unverdicted novelty 7.0

    Compositional interpretability defines explanations as commuting syntactic-semantic mapping pairs grounded in compositionality and minimum description length, with compressive refinement and a parsimony theorem guaran...

  3. From Mechanistic to Compositional Interpretability

    cs.LG 2026-05 unverdicted novelty 7.0

    The paper introduces compositional interpretability as a category-theoretic framework that casts mechanistic explanations as commuting syntactic-semantic mappings optimized under faithfulness and complexity constraint...

  4. ECG-InterpBench: Benchmarking the Interpretability of ECG Foundation Models with Matched-Scale Sparse Autoencoders

    cs.LG 2026-07 conditional novelty 6.0

    With matched-scale sparse autoencoders, HuBERT-ECG best preserves its ECG representation while ECG-JEPA best exposes clinical measurements through single features — a leader split that repeats on MIMIC-IV-ECG.

  5. Training Large Language Models for Self-Explanation Faithfulness

    cs.LG 2026-07 conditional novelty 6.0

    RL fine-tuning with a counterfactual mention/influence reward raises LLM self-explanation faithfulness (Phi-CCT) from near zero to ~0.66 in-distribution for two 8B models, with partial transfer to held-out tasks.

  6. XtrAIn: Training-Guided Occlusion for Feature Attribution

    cs.LG 2026-06 unverdicted novelty 6.0

    XtrAIn shifts occlusion from input space to parameter space along the training trajectory to produce cleaner feature attributions than standard methods.

  7. Scaling Vision Models Does Not Consistently Improve Localisation-Based Explanation Quality

    cs.CV 2026-05 accept novelty 6.0

    Scaling vision models by depth and parameter count does not consistently improve localisation-based explanation quality across architectures, datasets, and post-hoc methods; smaller models often perform comparably or better.

  8. Cognitive Alignment At No Cost: Inducing Human Attention Biases For Interpretable Vision Transformers

    cs.CV 2026-04 unverdicted novelty 6.0

    Fine-tuning ViT-B/16 attention on human saliency maps induces three human-like attention biases and improves alignment on five metrics with no loss in classification accuracy on ImageNet, ImageNet-C, and ObjectNet.

  9. Contrastive learning of extragalactic stellar streams. Sculpting a latent space of representations with DES DR2 photometry

    astro-ph.GA 2026-01 conditional novelty 6.0

    Applying NNCLR contrastive learning to DES DR2 galaxy cutouts yields embeddings that cluster major-merger galaxies but not stellar streams; a tiered sigmoid scaling redirects the network's saliency toward low-surface-...

  10. Variational Proximal Policy Optimization

    stat.ML 2026-06 unverdicted novelty 5.0

    VP2O maps PPO to SVGD in a MoE architecture using functional kernels and expert orthogonalization, claiming +179 ELO on Codeforces and 32% token reduction on AIME for a 33B/4B model.

  11. The impact of source and survey modelling on the connection between [O III] emitters and Ly $\alpha$ forest transmission at z ~ 6

    astro-ph.CO 2026-06 unverdicted novelty 5.0

    Empirical halo-to-[O III] emitter modeling with realistic JWST survey mocks produces cross-correlations consistent with z~6 data within large scatter, but with a ~10 cMpc offset in the 1D peak.

  12. Non-identifiability of Explanations from Model Behavior in Deep Networks of Image Authenticity Judgments

    cs.CV 2026-04 unverdicted novelty 5.0

    Models predicting human authenticity judgments produce inconsistent attribution maps across architectures, showing that explanations are non-identifiable.

  13. Learning to model pediatric asthma exacerbation from multiple risk factors: a case study in coastal Virginia

    cs.LG 2026-06 unverdicted novelty 4.0

    A case study develops a sparse dictionary learning approach to model pediatric asthma exacerbations from multiple risk factors and reports consensus on relative risks across statistical and machine learning models.