Pando shows that when models give no or misleading explanations, gradient-based attribution and relevance patching improve accuracy in predicting held-out model decisions by 3-5 percentage points over black-box methods, while logit lens, sparse autoencoders, and circuit tracing provide no reliable 3
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
Pando: Do Interpretability Methods Work When Models Won't Explain Themselves?
Pando shows that when models give no or misleading explanations, gradient-based attribution and relevance patching improve accuracy in predicting held-out model decisions by 3-5 percentage points over black-box methods, while logit lens, sparse autoencoders, and circuit tracing provide no reliable 3