Pith. sign in

REVIEW 17 cited by

Opportunities and Challenges in Explainable Artificial Intelligence (XAI): A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.11371 v2 pith:3M22GA5O submitted 2020-06-16 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords algorithmsdeepartificialintelligencechallengescriticalexplainableexplanation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Nowadays, deep neural networks are widely used in mission critical systems such as healthcare, self-driving vehicles, and military which have direct impact on human lives. However, the black-box nature of deep neural networks challenges its use in mission critical applications, raising ethical and judicial concerns inducing lack of trust. Explainable Artificial Intelligence (XAI) is a field of Artificial Intelligence (AI) that promotes a set of tools, techniques, and algorithms that can generate high-quality interpretable, intuitive, human-understandable explanations of AI decisions. In addition to providing a holistic view of the current XAI landscape in deep learning, this paper provides mathematical summaries of seminal work. We start by proposing a taxonomy and categorizing the XAI techniques based on their scope of explanations, methodology behind the algorithms, and explanation level or usage which helps build trustworthy, interpretable, and self-explanatory deep learning models. We then describe the main principles used in XAI research and present the historical timeline for landmark studies in XAI from 2007 to 2020. After explaining each category of algorithms and approaches in detail, we then evaluate the explanation maps generated by eight XAI algorithms on image data, discuss the limitations of this approach, and provide potential future directions to improve XAI evaluation.

Discussion (0). Sign in to comment.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 498 citations worldwide. Full citation record

  1. From Attribution to Action: A Human-Centered Application of Activation Steering

    cs.AI 2026-04 unverdicted novelty 6.5 of 10

    Activation steering paired with attribution enables intervention-based debugging in vision models, as all 8 interviewed experts shifted to hypothesis testing, most trusted observed responses, and highlighted risks lik...

  2. DCFO: Density-Based Counterfactuals for Outliers -- Additional Material

    cs.LG 2025-12 conditional novelty 6.0 of 10

    DCFO partitions the feature space by nearest-neighbour structure to make LOF scores differentiable, then uses gradient-based search to find the closest change that turns an outlier into an inlier.

  3. Attribution Graphs and Causal Probing for Mechanistic Discovery and Bias Repair in Multimodal Generative Learning

    cs.LG 2025-10 unverdicted novelty 6.0 of 10

    The paper claims a multimodal generative explainability framework with Attribution Graphs, Causal Probing, and a Cognitive Alignment Score, but the submitted text appears to describe a different, smaller multimodal cl...

  4. Robust Explanations Through Uncertainty Decomposition: A Path to Trustworthier AI

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Aleatoric uncertainty selects between counterfactual and feature-importance explanations, epistemic uncertainty rejects unreliable explanations, and correlation experiments support the rule.

  5. DeltaSHAP: Explaining Prediction Evolutions in Online Patient Monitoring with Shapley Values

    cs.LG 2025-07 conditional novelty 6.0 of 10

    DeltaSHAP attributes a monitoring model's prediction change between time steps to individual features using sampled Shapley values over the latest observed measurements.

  6. Iterative Self-Improvement of Vision Language Models for Image Scoring and Self-Explanation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    An iterative self-training method using DPO on self-generated score-conditioned explanations improves both image scoring accuracy and score-explanation consistency in VLMs.

  7. (EC)2: Event-Centric Explainability for Cybersecurity Through Multi-Agent LLM Investigations

    cs.CR 2026-07 reject novelty 5.0 of 10

    An event-centric, multi-agent LLM framework explains network alerts through hypothesis-driven, retrieval-augmented investigation and claims to improve explanation quality and boundary-case classification.

  8. Concept-based Visual Counterfactual Explanations with Diffusion Models

    cs.AI 2026-05 conditional novelty 5.0 of 10

    C-VCE embeds a concept-bottleneck classifier inside a diffusion generator so counterfactual edits are steered by interpretable attributes and a gradient mask, beating L-DVCE on proximity and realism but not on flip ra...

  9. ELASTIC: Event-Tracking Data Synchronization in Soccer Without Annotated Event Locations

    cs.DB 2025-08 unverdicted novelty 5.0 of 10

    The abstract claims ELASTIC outperforms prior soccer event-tracking synchronizers on 2,134 annotated events, but the submitted full text is a different manuscript, leaving the central claim unverifiable.

  10. "So, Tell Me About Your Policy...": Distillation of interpretable policies from Deep Reinforcement Learning agents

    cs.LG 2025-07 conditional novelty 5.0 of 10

    EXPLAIN trains an interpretable linear policy from an expert's offline trajectories by combining advantage-weighted policy gradients with a behavioral cloning regularizer.

  11. Revisiting LRP: Positional Attribution as the Missing Ingredient for Transformer Explainability

    cs.LG 2025-06 conditional novelty 5.0 of 10

    PA-LRP extends Layer-wise Relevance Propagation to attribute relevance to positional encodings in Transformers, improving faithfulness of explanations.

  12. From Plausible to Actionable: A Position on LLM Self-Explanations

    cs.CL 2026-07 conditional novelty 4.0 of 10

    Self-explanations from LLMs should be evaluated by their actionability for stakeholders rather than by plausibility or faithfulness alone.

  13. A Multi-level Analysis of Factors Associated with Student Performance: A Machine Learning Approach to the SAEB Microdata

    cs.LG 2025-10 conditional novelty 4.0 of 10

    On 6.48 million Brazilian students, a Random Forest model predicted above/below-average proficiency with 90.2% accuracy, and the school's average socioeconomic level was the most influential predictor.

  14. SHAP-Guided Regularization in Machine Learning Models

    cs.LG 2025-07 reject novelty 4.0 of 10

    A SHAP entropy and stability regularization for LightGBM is proposed, with small aggregate accuracy gains but no algorithm details or error bars.

  15. A Taxonomy for Design and Evaluation of Prompt-Based Natural Language Explanations

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A three-part taxonomy for prompt-based natural language explanations, covering context, generation and presentation, and evaluation with 15 desirable properties.

  16. Human-Centered Explainability in Interactive Information Systems: A Survey

    cs.HC 2025-07 conditional novelty 4.0 of 10

    A systematic review of 100 empirical user studies synthesizes explainability research into five conceptual dimensions, a design classification, and six measurement categories.

  17. Enhancing Interpretability of Quantum-Assisted Blockchain Clustering via AI Agent-Based Qualitative Analysis

    quant-ph 2025-06 reject novelty 4.0 of 10

    A two-stage framework uses clustering metrics and an LLM-based agent to interpret quantum-assisted blockchain clustering, reporting K=3 as optimal on MCO2 transaction data.

Pith tools