REVIEW 9 cited by
The Disagreement Problem in Explainable Machine Learning: A Practitioner's Perspective
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
As various post hoc explanation methods are increasingly being leveraged to explain complex models in high-stakes settings, it becomes critical to develop a deeper understanding of whether and when the explanations output by these methods disagree with each other, and how such disagreements are resolved in practice. However, there is little to no research that provides answers to these critical questions. In this work, we formalize and study the disagreement problem in explainable machine learning. More specifically, we define the notion of disagreement between explanations, analyze how often such disagreements occur in practice, and how practitioners resolve these disagreements. We first conduct interviews with data scientists to understand what constitutes disagreement between explanations generated by different methods for the same model prediction, and introduce a novel quantitative framework to formalize this understanding. We then leverage this framework to carry out a rigorous empirical analysis with four real-world datasets, six state-of-the-art post hoc explanation methods, and six different predictive models, to measure the extent of disagreement between the explanations generated by various popular explanation methods. In addition, we carry out an online user study with data scientists to understand how they resolve the aforementioned disagreements. Our results indicate that (1) state-of-the-art explanation methods often disagree in terms of the explanations they output, and (2) machine learning practitioners often employ ad hoc heuristics when resolving such disagreements. These findings suggest that practitioners may be relying on misleading explanations when making consequential decisions. They also underscore the importance of developing principled frameworks for effectively evaluating and comparing explanations output by various explanation techniques.
Forward citations
Cited by 9 Pith papers
-
Your Model Is Unfair, Are You Even Aware? Inverse Relationship Between Comprehension and Trust in Explainability Visualizations of Biased ML Models
More comprehensible explainability visualizations increase perceived model bias and decrease trust, with bias perception mediating the negative comprehension-trust relationship.
-
On Spectral Properties of Gradient-based Explanation Methods
Gradient-based explanations behave like frequency-band selectors: the gradient acts as a high-pass filter, perturbation as a low-pass filter, and their combination creates explanations that shift with the perturbation scale.
-
Multi-criteria Rank-based Aggregation for Explainable AI
A multi-criteria rank-based aggregation method that combines LIME, SHAP, and ANCHOR explanations, weighted by new rank-based complexity, faithfulness, and stability metrics, is proposed and tested on five datasets.
-
Explainable AI for the EU Right to Explanation: A Systematic Review of the Law-XAI Translation Gap
A systematic review of 57 post-2024 papers shows only 19 integrate EU law and XAI, most misidentify the GDPR basis, and the authors propose an addressee/purpose framework and a four-phase operationalization blueprint.
-
Circuit Claims Depend on What Is Extracted and How It Is Compared
On a synthetic Lean tactic-prediction task, exact circuit edge lists barely overlap across dense and sparse checkpoints while attention-head sets and size rankings do, so a circuit claim is well defined only once grap...
-
Interpretable Text Classification Applied to the Detection of LLM-generated Creative Writing
Simple word-count classifiers distinguish 1920s–30s detective fiction from GPT-4.1 rewrites with ~98% accuracy, and the main cue is the AI's broader synonym use plus modernized phrasing.
-
HattriQ: Designing Integrated Gradients for Feature Attribution in Quantum Machine Learning
HattriQ computes input-feature attributions for amplitude-encoded quantum classifiers by estimating amplitude gradients with Hadamard-test circuits and integrating them from a baseline image.
-
Position: Explainability Research Must Prioritize Foundations over Ad-hoc Methods
Explainable AI research should prioritize definitions, properties, evaluations, and actionability over new ad-hoc methods, on evidence from 617 papers and 34 practitioners.
-
CASE: Contrastive Activation for Saliency Estimation
CASE removes gradient components shared with confused classes to produce more class-distinct saliency maps, validated on a top-k overlap diagnostic where many existing methods show class-insensitive behavior.
Discussion (0). Continue with ORCID to comment.