Pith. sign in

REVIEW 3 cited by

Faithfulness Tests for Natural Language Explanations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.18029 v2 pith:GIED6HKK submitted 2023-05-29 cs.CL cs.AI

classification cs.CLcs.AI
keywords explanationsnlespredictionsreasonstestscounterfactualfaithfulnesslanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Explanations of neural models aim to reveal a model's decision-making process for its predictions. However, recent work shows that current methods giving explanations such as saliency maps or counterfactuals can be misleading, as they are prone to present reasons that are unfaithful to the model's inner workings. This work explores the challenging question of evaluating the faithfulness of natural language explanations (NLEs). To this end, we present two tests. First, we propose a counterfactual input editor for inserting reasons that lead to counterfactual predictions but are not reflected by the NLEs. Second, we reconstruct inputs from the reasons stated in the generated NLEs and check how often they lead to the same predictions. Our tests can evaluate emerging NLE models, proving a fundamental tool in the development of faithful NLEs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Training Large Language Models for Self-Explanation Faithfulness

    cs.LG 2026-07 conditional novelty 6.0 of 10

    RL fine-tuning with a counterfactual mention/influence reward raises LLM self-explanation faithfulness (Phi-CCT) from near zero to ~0.66 in-distribution for two 8B models, with partial transfer to held-out tasks.

  2. Rethinking Human Preference Evaluation of LLM Rationales

    cs.AI 2025-09 conditional novelty 6.0 of 10

    A fine-grained attribute-based evaluation of LLM rationales can explain human preferences and reveal model trade-offs that binary comparisons obscure.

  3. On the Risk of Misleading Reports: Diagnosing Textual Biases in Multimodal Clinical AI

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A new perturbation test shows that medical vision-language models rely more on clinical text than on images, with calibration errors growing when text conflicts with the image.

Pith tools