Activation steering of SAE-attributed components lets practitioners move from correlational inspection to causal hypothesis testing on CLIP failures, with trust shifting to observed model responses (N=8 experts).
hub
arXiv preprint arXiv:2006.11371 , year=
12 Pith papers cite this work, alongside 498 external citations. Polarity classification is still indexing.
hub tools
citation-role summary
citation-polarity summary
roles
background 3polarities
background 3representative citing papers
TIRA attacks with PMiS and PRSMP push fairness metrics to ideal values and reduce SHAP attribution for protected features to zero in black-box settings.
A latent mediation framework with sparse autoencoders enables non-additive token-level influence attribution in LLMs by learning orthogonal features and back-propagating attributions.
X-SYS is a reference architecture for interactive explanation systems organized around STAR quality attributes and five service components, demonstrated via SemanticLens for vision-language models.
Delta-XAI wraps existing XAI methods for online time series and introduces SWING to explain prediction changes while accounting for temporal dependencies.
The paper claims a multimodal generative explainability framework with Attribution Graphs, Causal Probing, and a Cognitive Alignment Score, but the submitted text appears to describe a different, smaller multimodal classification system evaluated on MNIST variants.
OPTIMUS generates minimal and sufficient concept-based visual explanations for deep classifiers using prime implicant theory to enforce logical sufficiency and minimality.
UNR-Explainer applies MCTS to find subgraphs that change k-NN relations in unsupervised node embeddings, claiming superior performance on GraphSAGE and DGI across datasets.
This survey synthesizes XAI methods with surrogate modeling workflows for simulations and outlines a research agenda to embed explainability into simulation-driven design and decision-making.
FAMeX creates a graph of feature associations to explain AI classification decisions and outperforms SHAP and permutation feature importance on eight benchmark datasets.
Infusing learning theories into the XAI lifecycle offers a learner-centered path to improve human agency and mitigate explanation-related risks in AI systems.
Industry AI practitioners view model quality through nine attributes with context-dependent priorities, where data imbalance is a key challenge addressed by strategies like active learning, as confirmed by interviews and a follow-up survey.
citing papers explorer
-
From Attribution to Action: A Human-Centered Application of Activation Steering
Activation steering of SAE-attributed components lets practitioners move from correlational inspection to causal hypothesis testing on CLIP failures, with trust shifting to observed model responses (N=8 experts).
-
The Unseen Hand: Manipulating Model Fairness and SHAP with Targeted Identity Re-Association Attacks
TIRA attacks with PMiS and PRSMP push fairness metrics to ideal values and reduce SHAP attribution for protected features to zero in black-box settings.
-
Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces
A latent mediation framework with sparse autoencoders enables non-additive token-level influence attribution in LLMs by learning orthogonal features and back-propagating attributions.
-
X-SYS: A Reference Architecture for Interactive Explanation Systems
X-SYS is a reference architecture for interactive explanation systems organized around STAR quality attributes and five service components, demonstrated via SemanticLens for vision-language models.
-
Delta-XAI: A Unified Framework for Explaining Prediction Changes in Online Time Series Monitoring
Delta-XAI wraps existing XAI methods for online time series and introduces SWING to explain prediction changes while accounting for temporal dependencies.
-
Attribution Graphs and Causal Probing for Mechanistic Discovery and Bias Repair in Multimodal Generative Learning
The paper claims a multimodal generative explainability framework with Attribution Graphs, Causal Probing, and a Cognitive Alignment Score, but the submitted text appears to describe a different, smaller multimodal classification system evaluated on MNIST variants.
-
OPTIMUS-Prime: Minimal and Sufficient Concept Explanations for Deep Vision Models
OPTIMUS generates minimal and sufficient concept-based visual explanations for deep classifiers using prime implicant theory to enforce logical sufficiency and minimality.
-
UNR-Explainer: Counterfactual Explanations for Unsupervised Node Representation Learning Models
UNR-Explainer applies MCTS to find subgraphs that change k-NN relations in unsupervised node embeddings, claiming superior performance on GraphSAGE and DGI across datasets.
-
Interpretable and Explainable Surrogate Modeling for Simulations: A State-of-the-Art Survey and Perspectives on Explainable AI for Decision-Making
This survey synthesizes XAI methods with surrogate modeling workflows for simulations and outlines a research agenda to embed explainability into simulation-driven design and decision-making.
-
A New Technique for AI Explainability using Feature Association Map
FAMeX creates a graph of feature associations to explain AI classification decisions and outperforms SHAP and permutation feature importance on eight benchmark datasets.
-
Using Learning Theories to Evolve Human-Centered XAI: Future Perspectives and Challenges
Infusing learning theories into the XAI lifecycle offers a learner-centered path to improve human agency and mitigate explanation-related risks in AI systems.
-
Industry Practitioners Perspectives on AI Model Quality: Perceptions, Challenges, and Solutions
Industry AI practitioners view model quality through nine attributes with context-dependent priorities, where data imbalance is a key challenge addressed by strategies like active learning, as confirmed by interviews and a follow-up survey.