Pith. sign in

REVIEW 1 cited by

Measuring Association Between Labels and Free-Text Rationales

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.12762 v4 pith:2IF4MGVD submitted 2020-10-24 cs.CL

classification cs.CL
keywords rationalesfree-textmodelsfaithfulassociationextractivefaithfulnesslanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In interpretable NLP, we require faithful rationales that reflect the model's decision-making process for an explained instance. While prior work focuses on extractive rationales (a subset of the input words), we investigate their less-studied counterpart: free-text natural language rationales. We demonstrate that pipelines, existing models for faithful extractive rationalization on information-extraction style tasks, do not extend as reliably to "reasoning" tasks requiring free-text rationales. We turn to models that jointly predict and rationalize, a class of widely used high-performance models for free-text rationalization whose faithfulness is not yet established. We define label-rationale association as a necessary property for faithfulness: the internal mechanisms of the model producing the label and the rationale must be meaningfully correlated. We propose two measurements to test this property: robustness equivalence and feature importance agreement. We find that state-of-the-art T5-based joint models exhibit both properties for rationalizing commonsense question-answering and natural language inference, indicating their potential for producing faithful free-text rationales.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Self-Critique and Refinement for Faithful Natural Language Explanations

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A self-critique and refinement framework with word-level feedback cuts unfaithfulness rates in LLM explanations by about 19 points on average, but gains may partly reflect the word-presence evaluation metric.

Pith tools