Pith. sign in

REVIEW 6 cited by

Towards Faithfully Interpretable NLP Systems: How should we define and evaluate faithfulness?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2004.03685 v3 pith:K57RXHY4 submitted 2020-04-07 cs.CL cs.LG

classification cs.CLcs.LG
keywords faithfulnessshouldcurrentevaluationinterpretationbinarycallcriteria
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With the growing popularity of deep-learning based NLP models, comes a need for interpretable systems. But what is interpretability, and what constitutes a high-quality interpretation? In this opinion piece we reflect on the current state of interpretability evaluation research. We call for more clearly differentiating between different desired criteria an interpretation should satisfy, and focus on the faithfulness criteria. We survey the literature with respect to faithfulness evaluation, and arrange the current approaches around three assumptions, providing an explicit form to how faithfulness is "defined" by the community. We provide concrete guidelines on how evaluation of interpretation methods should and should not be conducted. Finally, we claim that the current binary definition for faithfulness sets a potentially unrealistic bar for being considered faithful. We call for discarding the binary notion of faithfulness in favor of a more graded one, which we believe will be of greater practical utility.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning

    cs.CL 2025-06 conditional novelty 7.0 of 10

    A pre-RL fine-tuning intervention called verbalization fine-tuning makes language models explicitly acknowledge when prompt cues drive them to reward-hack, cutting undetected reward hacking from 88% to 6% after RL.

  2. Training Large Language Models for Self-Explanation Faithfulness

    cs.LG 2026-07 conditional novelty 6.0 of 10

    RL fine-tuning with a counterfactual mention/influence reward raises LLM self-explanation faithfulness (Phi-CCT) from near zero to ~0.66 in-distribution for two 8B models, with partial transfer to held-out tasks.

  3. Which Prompting Technique Should I Use? An Empirical Investigation of Prompting Techniques for Software Engineering Tasks

    cs.SE 2025-06 conditional novelty 6.0 of 10

    Across ten software engineering tasks and four LLMs, no prompting technique wins consistently; ES-KNN is best on many tasks, some techniques underperform the baseline, and USC is best for code QA and code generation.

  4. From "Thinking" to "Justifying": Aligning High-Stakes Explainability with Professional Communication Standards

    cs.AI 2026-01 conditional novelty 5.0 of 10

    Conclusion-first, structured justifications (SEF) outperform chain-of-thought by 5.3 points on four high-stakes yes/no tasks, and six rule-based structure metrics correlate with correctness (r=0.20–0.42).

  5. Explainable Knowledge Graph Retrieval-Augmented Generation (KG-RAG) with KG-SMILE

    cs.AI 2025-09 reject novelty 4.0 of 10

    KG-SMILE applies perturbation and linear regression to a knowledge graph to attribute which entities and relations drive a GraphRAG system's answers.

  6. RAG-PRISM: A Personalized, Rapid, and Immersive Skill Mastery Framework with Adaptive Retrieval-Augmented Tutoring

    cs.CY 2025-08 reject novelty 4.0 of 10

    A RAG-based adaptive tutoring framework for cybersecurity training is presented and evaluated on a small synthetic QA dataset, reporting high faithfulness and relevancy for GPT-4, but with inconsistent metrics and no ...

Pith tools