Latent reasoning faithfulness is a property of training stage and answer format, not of architecture or the final checkpoint alone.
Towards Faithful Model Explanation in NLP : A Survey
3 Pith papers cite this work, alongside 73 external citations. Polarity classification is still indexing.
3
Pith papers citing it
73
external citations · OpenAlex
citation-role summary
background 1
citation-polarity summary
years
2026 3roles
background 1polarities
support 1representative citing papers
Pruning attention layers in five LLMs across eight datasets maintains accuracy but degrades faithfulness and calibration.
A position paper argues that post-hoc XAI explanations are unfaithful and paradoxical, proposing a shift to expert-based verification and certification of AI systems.
citing papers explorer
-
Beyond Explainable AI (XAI): An Overdue Paradigm Shift and Post-XAI Research Directions
A position paper argues that post-hoc XAI explanations are unfaithful and paradoxical, proposing a shift to expert-based verification and certification of AI systems.