Pith. sign in

REVIEW 1 cited by

Evaluating Explanation Without Ground Truth in Interpretable Machine Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1907.06831 v2 pith:FHWGEF32 submitted 2019-07-16 cs.LG cs.AIcs.HCstat.ML

classification cs.LGcs.AIcs.HCstat.ML
keywords evaluationexplanationsexplanationlearningmachinebenchmarkdifferentevaluating
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Interpretable Machine Learning (IML) has become increasingly important in many real-world applications, such as autonomous cars and medical diagnosis, where explanations are significantly preferred to help people better understand how machine learning systems work and further enhance their trust towards systems. However, due to the diversified scenarios and subjective nature of explanations, we rarely have the ground truth for benchmark evaluation in IML on the quality of generated explanations. Having a sense of explanation quality not only matters for assessing system boundaries, but also helps to realize the true benefits to human users in practical settings. To benchmark the evaluation in IML, in this article, we rigorously define the problem of evaluating explanations, and systematically review the existing efforts from state-of-the-arts. Specifically, we summarize three general aspects of explanation (i.e., generalizability, fidelity and persuasibility) with formal definitions, and respectively review the representative methodologies for each of them under different tasks. Further, a unified evaluation framework is designed according to the hierarchical needs from developers and end-users, which could be easily adopted for different scenarios in practice. In the end, open problems are discussed, and several limitations of current evaluation techniques are raised for future explorations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Deep Learning-based Multi Project InP Wafer Simulation for Unsupervised Surface Defect Detection

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A U-Net trained on CAD layouts and flawed wafer photos can generate defect-free synthetic wafer images that serve as a template for automated defect detection in InP multi-project wafer manufacturing.

Pith tools