Latent reasoning faithfulness is a property of training stage and answer format, not of architecture or the final checkpoint alone.
Towards Faithful Model Explanation in NLP : A Survey
3 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 3roles
background 1polarities
support 1representative citing papers
Pruning attention layers in five LLMs across eight datasets maintains accuracy but degrades faithfulness and calibration.
A position paper argues that post-hoc XAI explanations are unfaithful and paradoxical, proposing a shift to expert-based verification and certification of AI systems.
citing papers explorer
-
Final Checkpoints Are Not Enough: Analyzing Latent Reasoning Faithfulness Along Training Trajectories
Latent reasoning faithfulness is a property of training stage and answer format, not of architecture or the final checkpoint alone.
-
Don't Go Breaking My LLM: The Impact of Pruning Attention Layers on Explanation Faithfulness and Confidence Calibration
Pruning attention layers in five LLMs across eight datasets maintains accuracy but degrades faithfulness and calibration.
-
Beyond Explainable AI (XAI): An Overdue Paradigm Shift and Post-XAI Research Directions
A position paper argues that post-hoc XAI explanations are unfaithful and paradoxical, proposing a shift to expert-based verification and certification of AI systems.