Pith. sign in

REVIEW 3 cited by

Quantifying Interpretability and Trust in Machine Learning Systems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1901.08558 v1 pith:R6ELSALW submitted 2019-01-20 cs.LG stat.ML

classification cs.LGstat.ML
keywords interpretabilitytrustdecisionsmethodsmeasurehumanmetricpredictions
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Decisions by Machine Learning (ML) models have become ubiquitous. Trusting these decisions requires understanding how algorithms take them. Hence interpretability methods for ML are an active focus of research. A central problem in this context is that both the quality of interpretability methods as well as trust in ML predictions are difficult to measure. Yet evaluations, comparisons and improvements of trust and interpretability require quantifiable measures. Here we propose a quantitative measure for the quality of interpretability methods. Based on that we derive a quantitative measure of trust in ML decisions. Building on previous work we propose to measure intuitive understanding of algorithmic decisions using the information transfer rate at which humans replicate ML model predictions. We provide empirical evidence from crowdsourcing experiments that the proposed metric robustly differentiates interpretability methods. The proposed metric also demonstrates the value of interpretability for ML assisted human decision making: in our experiments providing explanations more than doubled productivity in annotation tasks. However unbiased human judgement is critical for doctors, judges, policy makers and others. Here we derive a trust metric that identifies when human decisions are overly biased towards ML predictions. Our results complement existing qualitative work on trust and interpretability by quantifiable measures that can serve as objectives for further improving methods in this field of research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Efficient Adaptation of Reinforcement Learning Agents to Sudden Environmental Change

    cs.LG 2025-05 conditional novelty 6.0 of 10

    The dissertation shows that efficient online adaptation to sudden environmental change requires exploration that prioritises diverse, task-agnostic data and world models that selectively preserve reusable knowledge.

  2. Interactive Multi-Objective Probabilistic Preference Learning with Soft and Hard Bounds

    cs.AI 2025-06 conditional novelty 5.0 of 10

    Active-MoSH interactively learns a decision maker's soft and hard bounds on objectives and actively samples Pareto-optimal points, with a sensitivity analysis module intended to build trust.

  3. Evaluating Explainability: A Framework for Systematic Assessment and Reporting of Explainable AI Features

    cs.AI 2025-06 conditional novelty 5.0 of 10

    A structured evaluation framework and scorecard for explainable AI is demonstrated on heatmaps for detecting breast lesions in synthetic mammograms.

Pith tools