An explainability-aware L0 penalty yields coherent counterfactual edits, and the same geometry defines a Tolerance-Region Confusion Matrix that quantifies class-to-class fragility under interpretable perturbations.
Getting a CLUE: A Method for Explaining Uncertainty Estimates, March 2021
3 Pith papers cite this work. Polarity classification is still indexing.
abstract
Both uncertainty estimation and interpretability are important factors for trustworthy machine learning systems. However, there is little work at the intersection of these two areas. We address this gap by proposing a novel method for interpreting uncertainty estimates from differentiable probabilistic models, like Bayesian Neural Networks (BNNs). Our method, Counterfactual Latent Uncertainty Explanations (CLUE), indicates how to change an input, while keeping it on the data manifold, such that a BNN becomes more confident about the input's prediction. We validate CLUE through 1) a novel framework for evaluating counterfactual explanations of uncertainty, 2) a series of ablation experiments, and 3) a user study. Our experiments show that CLUE outperforms baselines and enables practitioners to better understand which input patterns are responsible for predictive uncertainty.
fields
cs.LG 3representative citing papers
Framework embeds aleatoric and epistemic uncertainties into BNN parameter variances and applies moment propagation for sampling-free variational inference in lightweight networks.
TRUST searches for minimal input changes that achieve a user-defined confidence target in PTM models, claiming perfect robustness and low cost on benchmarks versus standard boundary-crossing methods.
citing papers explorer
-
Optimized Instance Alteration for Explaining and Assessing Robustness of Classifiers
An explainability-aware L0 penalty yields coherent counterfactual edits, and the same geometry defines a Tolerance-Region Confusion Matrix that quantifies class-to-class fragility under interpretable perturbations.
-
A Framework for Variational Inference of Lightweight Bayesian Neural Networks with Heteroscedastic Uncertainties
Framework embeds aleatoric and epistemic uncertainties into BNN parameter variances and applies moment propagation for sampling-free variational inference in lightweight networks.
-
Target-confidence Recourse Using tSeTlin machines: TRUST
TRUST searches for minimal input changes that achieve a user-defined confidence target in PTM models, claiming perfect robustness and low cost on benchmarks versus standard boundary-crossing methods.