Pith. sign in

Right for the Right Reasons: Training Differentiable Models by Constraining their Explanations

8 Pith papers cite this work, alongside 7 external citations. Polarity classification is still indexing.

8 Pith papers citing it
7 external citations · Pith
abstract

Neural networks are among the most accurate supervised learning methods in use today, but their opacity makes them difficult to trust in critical applications, especially when conditions in training differ from those in test. Recent work on explanations for black-box models has produced tools (e.g. LIME) to show the implicit rules behind predictions, which can help us identify when models are right for the wrong reasons. However, these methods do not scale to explaining entire datasets and cannot correct the problems they reveal. We introduce a method for efficiently explaining and regularizing differentiable models by examining and selectively penalizing their input gradients, which provide a normal to the decision boundary. We apply these penalties both based on expert annotation and in an unsupervised fashion that encourages diverse models with qualitatively different decision boundaries for the same classification problem. On multiple datasets, we show our approach generates faithful explanations and models that generalize much better when conditions differ between training and test.

citation-role summary

background 1

citation-polarity summary

roles

background 1

polarities

background 1

representative citing papers

Landseer: Exploring the Machine Learning Defense Landscape

cs.CR · 2026-05-26 · unverdicted · novelty 6.0

Landseer offers a containerized modular system to integrate and evaluate combinations of machine learning defenses, with an initial analysis of 35 defenses highlighting replicability challenges.

Shortcut Mitigation via Spurious-Positive Samples

cs.LG · 2026-05-13 · unverdicted · novelty 5.0

A method uses spurious-positive samples to identify and regularize neurons that rely on spurious features, improving model robustness without extra annotations or balanced data.

citing papers explorer

Showing 8 of 8 citing papers.