ExaGPT uses span-level similarity retrieval from human and LLM datastores to detect machine-generated text while supplying the matching spans as human-interpretable evidence, achieving up to 37-point accuracy gains over prior interpretable detectors at 1% FPR.
hub
"Why Should I Trust You?": Explaining the Predictions of Any Classifier
27 Pith papers cite this work, alongside 347 external citations. Polarity classification is still indexing.
abstract
Despite widespread adoption, machine learning models remain mostly black boxes. Understanding the reasons behind predictions is, however, quite important in assessing trust, which is fundamental if one plans to take action based on a prediction, or when choosing whether to deploy a new model. Such understanding also provides insights into the model, which can be used to transform an untrustworthy model or prediction into a trustworthy one. In this work, we propose LIME, a novel explanation technique that explains the predictions of any classifier in an interpretable and faithful manner, by learning an interpretable model locally around the prediction. We also propose a method to explain models by presenting representative individual predictions and their explanations in a non-redundant way, framing the task as a submodular optimization problem. We demonstrate the flexibility of these methods by explaining different models for text (e.g. random forests) and image classification (e.g. neural networks). We show the utility of explanations via novel experiments, both simulated and with human subjects, on various scenarios that require trust: deciding if one should trust a prediction, choosing between models, improving an untrustworthy classifier, and identifying why a classifier should not be trusted.
hub tools
citation-role summary
citation-polarity summary
fields
cs.LG 8 cs.AI 5 cs.CV 2 cs.SE 2 eess.IV 2 stat.ML 2 cond-mat.mtrl-sci 1 cond-mat.soft 1 cs.CE 1 cs.CL 1roles
background 4polarities
background 4representative citing papers
TIRA attacks with PMiS and PRSMP push fairness metrics to ideal values and reduce SHAP attribution for protected features to zero in black-box settings.
Defines MCT as the weakest confidence an abductive explanation can guarantee and proposes an optimization-based algorithm to generate minimal explanations meeting a target confidence threshold for boosted tree classifiers.
CodeQ aggregates token rationales into code categories to enable global interpretability of LLMs, claiming over 50% entropy reduction and revealing model preference for syntactic cues plus human misalignment in a 37-person study.
Scaling vision models by depth and parameter count does not consistently improve localisation-based explanation quality across architectures, datasets, and post-hoc methods; smaller models often perform comparably or better.
The paper presents a taxonomy of seven production-specific failure modes for agentic AI, demonstrates that existing metrics fail to detect four of them entirely, and proposes the PAEF five-dimension framework for continuous production evaluation with an open-source implementation.
A new scale-aware diagnostic framework shows that unconstrained diffusion generative models exhibit structural freezing and instability instead of smooth physical responses under multiscale perturbations.
LSI and ζ best distinguish temperature-dependent local structure in supercooled TIP4P/2005 water under a unified neural temperature-classification benchmark, with H-bond network descriptors next.
Cross-modal averaging maps ECG model attributions to CineECG 3D space, raising Dice overlap with expert annotations from 0.47 to 0.56 on 20 cases while filtering attribution noise.
The authors define interpretability for machine learning, specify when it is required, and propose a taxonomy for its rigorous evaluation while identifying open research questions.
VP2O maps PPO to SVGD in a MoE architecture using functional kernels and expert orthogonalization, claiming +179 ELO on Codeforces and 32% token reduction on AIME for a 33B/4B model.
ECPO is a listwise policy optimization method that couples ranking utility with span-level evidence certificate validity and a deterministic verifier reward on MAVEN-ERE and RAMS datasets.
A method that translates causal relationships into a Bipolar Argumentation Framework and applies semi-stable semantics to generate explanatory feature sets for machine learning predictions.
LLM reasoning traces and post-hoc explanations increase false trust in incorrect predictions, whereas contrastive dual explanations enhance users' ability to distinguish correct from incorrect AI outputs.
A data-derived baseline using feature effects on binary outcomes provides a model-agnostic way to check if machine learning explanations align with the underlying data structure.
Developed 19 metamorphic relations to test correlation detection and LSTM forecasting in an outage prediction application, uncovering 8 unknown issues in the live system and detecting 65.9% of injected bugs via mutation testing.
An optimized KernelSHAP method for 3D medical image segmentation restricts computation to ROI and receptive fields, uses patch logit caching for 15-30% savings, and compares organ units versus supervoxels for clinically interpretable attributions.
Proposes combining causal ML with interpretable models to achieve competitive prediction performance and transparency on causal structures for decision support.
IViT applies quadratic programming to a pre-trained Vision Transformer with a multi-objective loss, achieving 93.80% accuracy on six skin disease datasets (0.21% below baseline) while reducing feature redundancy by 29.5% and producing clinically consistent activations.
Spatial Learning Entropy Maps derived from MLP weight adaptations during spatial pixel prediction tasks highlight image points with high learning impact.
XpertXAI is an expert-driven multi-pathology CBM that outperforms post-hoc XAI methods and unsupervised CBMs in both accuracy and alignment with radiologist annotations on public chest X-ray data.
A Giant-Step Baby-Step Classifier for real-time anomaly detection in ICS via linearization of sensor-actuator interactions, achieving millisecond response and 97.72% accuracy on a water treatment testbed with built-in explainability.
Industry AI practitioners view model quality through nine attributes with context-dependent priorities, where data imbalance is a key challenge addressed by strategies like active learning, as confirmed by interviews and a follow-up survey.
Benchmark of local explainability methods on tabular data finds explanation quality driven primarily by dataset complexity rather than model predictive performance.
citing papers explorer
-
ExaGPT: Example-Based Machine-Generated Text Detection for Human Interpretability
ExaGPT uses span-level similarity retrieval from human and LLM datastores to detect machine-generated text while supplying the matching spans as human-interpretable evidence, achieving up to 37-point accuracy gains over prior interpretable detectors at 1% FPR.
-
The Unseen Hand: Manipulating Model Fairness and SHAP with Targeted Identity Re-Association Attacks
TIRA attacks with PMiS and PRSMP push fairness metrics to ideal values and reduce SHAP attribution for protected features to zero in black-box settings.
-
Beyond Explaining Predictions: Logic-Based Explanations for Confidence in Machine Learning Models
Defines MCT as the weakest confidence an abductive explanation can guarantee and proposes an optimization-based algorithm to generate minimal explanations meeting a target confidence threshold for boosted tree classifiers.
-
Enabling Global, Human-Centered Explanations for LLMs:From Tokens to Interpretable Code and Test Generation
CodeQ aggregates token rationales into code categories to enable global interpretability of LLMs, claiming over 50% entropy reduction and revealing model preference for syntactic cues plus human misalignment in a 37-person study.
-
Scaling Vision Models Does Not Consistently Improve Localisation-Based Explanation Quality
Scaling vision models by depth and parameter count does not consistently improve localisation-based explanation quality across architectures, datasets, and post-hoc methods; smaller models often perform comparably or better.
-
Evaluating Agentic AI in the Wild: Failure Modes, Drift Patterns, and a Production Evaluation Framework
The paper presents a taxonomy of seven production-specific failure modes for agentic AI, demonstrates that existing metrics fail to detect four of them entirely, and proposes the PAEF five-dimension framework for continuous production evaluation with an open-source implementation.
-
Scale-Aware Adversarial Analysis: A Diagnostic for Generative AI in Multiscale Complex Systems
A new scale-aware diagnostic framework shows that unconstrained diffusion generative models exhibit structural freezing and instability instead of smooth physical responses under multiscale perturbations.
-
Machine learning evaluation of structural descriptors for supercooled water
LSI and ζ best distinguish temperature-dependent local structure in supercooled TIP4P/2005 water under a unified neural temperature-classification benchmark, with H-bond network descriptors next.
-
Validating the Clinical Utility of CineECG 3D Reconstructions through Cross-Modal Feature Attribution
Cross-modal averaging maps ECG model attributions to CineECG 3D space, raising Dice overlap with expert annotations from 0.47 to 0.56 on 20 cases while filtering attribution noise.
-
Towards A Rigorous Science of Interpretable Machine Learning
The authors define interpretability for machine learning, specify when it is required, and propose a taxonomy for its rigorous evaluation while identifying open research questions.
-
Variational Proximal Policy Optimization
VP2O maps PPO to SVGD in a MoE architecture using functional kernels and expert orthogonalization, claiming +179 ELO on Codeforces and 32% token reduction on AIME for a 33B/4B model.
-
ECPO: Evidence-Coupled Policy Optimization for Evidence-Certified Candidate Ranking
ECPO is a listwise policy optimization method that couples ranking utility with span-level evidence certificate validity and a deterministic verifier reward on MAVEN-ERE and RAMS datasets.
-
A Causal Argumentation Method for Explainability of Machine Learning Models
A method that translates causal relationships into a Bipolar Argumentation Framework and applies semi-stable semantics to generate explanatory feature sets for machine learning predictions.
-
Evaluating the False Trust Engendered by LLM Explanations
LLM reasoning traces and post-hoc explanations increase false trust in incorrect predictions, whereas contrastive dual explanations enhance users' ability to distinguish correct from incorrect AI outputs.
-
Does the Model Say What the Data Says? A Simple Heuristic for Model Data Alignment
A data-derived baseline using feature effects on binary outcomes provides a model-agnostic way to check if machine learning explanations align with the underlying data structure.
-
Metamorphic Testing of a Deep Learning based Forecaster
Developed 19 metamorphic relations to test correlation detection and LSTM forecasting in an outage prediction application, uncovering 8 unknown issues in the live system and detecting 65.9% of injected bugs via mutation testing.
-
Efficient KernelSHAP Explanations for Patch-based 3D Medical Image Segmentation
An optimized KernelSHAP method for 3D medical image segmentation restricts computation to ROI and receptive fields, uses patch logit caching for 15-30% savings, and compares organ units versus supervoxels for clinically interpretable attributions.
-
A Step Towards Inherently Interpretable Causal Machine Learning Models For Decision Support
Proposes combining causal ML with interpretable models to achieve competitive prediction performance and transparency on causal structures for decision support.
-
IViT: A Novel Interpretable Visual Transformer for Skin Disease Detection
IViT applies quadratic programming to a pre-trained Vision Transformer with a multi-objective loss, achieving 93.80% accuracy on six skin disease datasets (0.21% below baseline) while reducing feature redundancy by 29.5% and producing clinically consistent activations.
-
Learning Entropy and Spatial Adaptation Dynamics of Multilayer Perceptrons for Structural Point Extraction
Spatial Learning Entropy Maps derived from MLP weight adaptations during spatial pixel prediction tasks highlight image points with high learning impact.
-
Explainability Through Human-Centric Design for XAI in Lung Cancer Detection
XpertXAI is an expert-driven multi-pathology CBM that outperforms post-hoc XAI methods and unsupervised CBMs in both accuracy and alignment with radiologist annotations on public chest X-ray data.
-
A Giant-Step Baby-Step Classifier For Scalable and Real-Time Anomaly Detection In Industrial Control Systems and Water Treatment Systems
A Giant-Step Baby-Step Classifier for real-time anomaly detection in ICS via linearization of sensor-actuator interactions, achieving millisecond response and 97.72% accuracy on a water treatment testbed with built-in explainability.
-
Industry Practitioners Perspectives on AI Model Quality: Perceptions, Challenges, and Solutions
Industry AI practitioners view model quality through nine attributes with context-dependent priorities, where data imbalance is a key challenge addressed by strategies like active learning, as confirmed by interviews and a follow-up survey.
-
Evaluating Local Explainability Metrics for Machine Learning Models on Tabular Data
Benchmark of local explainability methods on tabular data finds explanation quality driven primarily by dataset complexity rather than model predictive performance.
-
Platonic Projection Structures: Operator-Induced Observability in Representation Learning
The paper introduces 'Platonic Projection Structures,' a reformulation of standard PSD operator theory applied to representation learning, with experiments that verify definitions rather than test predictions.
-
Self-Explainability in Self-Adaptive and Self-Organising Systems: Status and Research Directions
A systematic literature review defines self-explainability, proposes a taxonomy and levels framework, and reports that most approaches are conceptual with no standard evaluation method.
-
Six Open Questions in Machine-Learned Interatomic Potential Foundation Models
This perspective article develops a definition of foundational MLIPs and poses six open questions that the authors believe will define future research in machine-learned interatomic potentials.