Pith. sign in

REVIEW 1 cited by

Generating Hierarchical Explanations on Text Classification via Feature Interaction Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2004.02015 v3 pith:TTRGIEIP submitted 2020-04-04 cs.CL

classification cs.CL
keywords explanationstextexplanationfeaturegeneratinghierarchicalhumansinteractions
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Generating explanations for neural networks has become crucial for their applications in real-world with respect to reliability and trustworthiness. In natural language processing, existing methods usually provide important features which are words or phrases selected from an input text as an explanation, but ignore the interactions between them. It poses challenges for humans to interpret an explanation and connect it to model prediction. In this work, we build hierarchical explanations by detecting feature interactions. Such explanations visualize how words and phrases are combined at different levels of the hierarchy, which can help users understand the decision-making of black-box models. The proposed method is evaluated with three neural text classifiers (LSTM, CNN, and BERT) on two benchmark datasets, via both automatic and human evaluations. Experiments show the effectiveness of the proposed method in providing explanations that are both faithful to models and interpretable to humans.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Concept-Level Explainability for Auditing & Steering LLM Responses

    cs.CL 2025-05 conditional novelty 6.0 of 10

    ConceptX is a concept-level attribution method that ranks semantically rich prompt words by their effect on an LLM's response, and editing those words can shift sentiment and reduce harmful outputs.

Pith tools