Introduces BonaFide benchmark of 3,066 ground-truth labeled CoTs showing most faithfulness metrics perform near chance with biases and poor scaling to longer chains.
On measuring faithfulness or self-consistency of natural language explanations
6 Pith papers cite this work, alongside 14 external citations. Polarity classification is still indexing.
years
2026 6representative citing papers
Proposes SCSuff metric for evaluating LLM explanation sufficiency via model-generated alternative inputs, showing explanations are typically insufficient and predictable from hidden states.
AttriCoT is a black-box algorithm that attributes causal importance to units in a specific CoT trace via a structural causal model estimated with linear forward passes.
Multi-agent debate in medical QA creates a consistency illusion by reducing answer contradictions while decreasing reasoning similarity; the Grounded Debate Protocol improves alignment with large effect sizes.
AtManRL learns an additive attention mask on CoT traces to produce a saliency reward that, when combined with outcome rewards in GRPO, trains LLMs to generate reasoning that genuinely influences final predictions.
citing papers explorer
-
Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth
Introduces BonaFide benchmark of 3,066 ground-truth labeled CoTs showing most faithfulness metrics perform near chance with biases and poor scaling to longer chains.
-
What LLMs explain is not what they believe: Evaluating explanation sufficiency under models' own input beliefs
Proposes SCSuff metric for evaluating LLM explanation sufficiency via model-generated alternative inputs, showing explanations are typically insufficient and predictable from hidden states.
-
Local Causal Attribution of Chain-of-Thought Reasoning
AttriCoT is a black-box algorithm that attributes causal importance to units in a specific CoT trace via a structural causal model estimated with linear forward passes.
-
The Consistency Illusion: How Multi-Agent Debate Hides Reasoning Misalignment
Multi-agent debate in medical QA creates a consistency illusion by reducing answer contradictions while decreasing reasoning similarity; the Grounded Debate Protocol improves alignment with large effect sizes.
-
AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency
AtManRL learns an additive attention mask on CoT traces to produce a saliency reward that, when combined with outcome rewards in GRPO, trains LLMs to generate reasoning that genuinely influences final predictions.
- Reasoning Dynamics and the Limits of Monitoring Modality Reliance in Vision-Language Models