Pith. sign in

REVIEW 1 cited by

Rationalization through Concepts

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2105.04837 v1 pith:2L5KL5SM submitted 2021-05-11 cs.CL cs.LG

classification cs.CLcs.LG
keywords conceptsconratdecisionsexplanationexplanationsfurtherhumaninterpretable
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Automated predictions require explanations to be interpretable by humans. One type of explanation is a rationale, i.e., a selection of input features such as relevant text snippets from which the model computes the outcome. However, a single overall selection does not provide a complete explanation, e.g., weighing several aspects for decisions. To this end, we present a novel self-interpretable model called ConRAT. Inspired by how human explanations for high-level decisions are often based on key concepts, ConRAT extracts a set of text snippets as concepts and infers which ones are described in the document. Then, it explains the outcome with a linear aggregation of concepts. Two regularizers drive ConRAT to build interpretable concepts. In addition, we propose two techniques to boost the rationale and predictive performance further. Experiments on both single- and multi-aspect sentiment classification tasks show that ConRAT is the first to generate concepts that align with human rationalization while using only the overall label. Further, it outperforms state-of-the-art methods trained on each aspect label independently.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Regulation of Language Models With Interpretability Will Likely Result In A Performance Trade-Off

    cs.LG 2024-12 conditional novelty 6.0 of 10

    Forcing an LLM to classify using only human-specified legal concepts costs about 7.34% accuracy, but can speed up human decision-making despite the loss.

Pith tools