Pith. sign in

REVIEW 3 cited by

Teach Me to Explain: A Review of Datasets for Explainable Natural Language Processing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2102.12060 v4 pith:2P4UOGHV submitted 2021-02-24 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords explanationsdatasetscollectingexnlpexplainableidentifyreviewtextual
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Explainable NLP (ExNLP) has increasingly focused on collecting human-annotated textual explanations. These explanations are used downstream in three ways: as data augmentation to improve performance on a predictive task, as supervision to train models to produce explanations for their predictions, and as a ground-truth to evaluate model-generated explanations. In this review, we identify 65 datasets with three predominant classes of textual explanations (highlights, free-text, and structured), organize the literature on annotating each type, identify strengths and shortcomings of existing collection methodologies, and give recommendations for collecting ExNLP datasets in the future.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From "Thinking" to "Justifying": Aligning High-Stakes Explainability with Professional Communication Standards

    cs.AI 2026-01 conditional novelty 5.0 of 10

    Conclusion-first, structured justifications (SEF) outperform chain-of-thought by 5.3 points on four high-stakes yes/no tasks, and six rule-based structure metrics correlate with correctness (r=0.20–0.42).

  2. Can human clinical rationales improve the performance and explainability of clinical text classification models?

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Adding 96,679 human rationale highlights improves cancer-site classification less than adding the same number of full pathology reports, and the explainability gain is small.

  3. Faithful and Robust LLM-Driven Theorem Proving for NLI Explanations

    cs.CL 2025-05 conditional novelty 5.0 of 10

    The proposed Faithful-Refiner, combining syntactic parsing, quantifier and consistency checks, logical-relation guidance, and detailed proof feedback, raises explanation refinement rates on three NLI benchmarks by lar...

Pith tools