Pith. sign in

REVIEW 2 cited by

WT5?! Training Text-to-Text Models to Explain their Predictions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2004.14546 v1 pith:UXGYLZWS submitted 2020-04-30 cs.CL cs.LG

classification cs.CLcs.LG
keywords modelsnaturalpredictiontrainexplanationlanguageneuraloutput
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Neural networks have recently achieved human-level performance on various challenging natural language processing (NLP) tasks, but it is notoriously difficult to understand why a neural network produced a particular prediction. In this paper, we leverage the text-to-text framework proposed by Raffel et al.(2019) to train language models to output a natural text explanation alongside their prediction. Crucially, this requires no modifications to the loss function or training and decoding procedures -- we simply train the model to output the explanation after generating the (natural text) prediction. We show that this approach not only obtains state-of-the-art results on explainability benchmarks, but also permits learning from a limited set of labeled explanations and transferring rationalization abilities across datasets. To facilitate reproducibility and future work, we release our code use to train the models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 104 citations worldwide. Full citation record

  1. Iterative Self-Improvement of Vision Language Models for Image Scoring and Self-Explanation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    An iterative self-training method using DPO on self-generated score-conditioned explanations improves both image scoring accuracy and score-explanation consistency in VLMs.

  2. Towards Transparent AI: A Survey on Explainable Large Language Models

    cs.CL 2025-06 conditional novelty 3.0 of 10

    A review that groups LLM explainability methods by transformer architecture and discusses their evaluation and applications.

Pith tools