Pith. sign in

REVIEW 2 cited by

To Trust Or Not To Trust A Classifier

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1805.11783 v2 pith:W6OZAJKP submitted 2018-05-30 stat.ML cs.LG

classification stat.MLcs.LG
keywords classifierscoretrusthighmanyconfidenceexampleshould
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Knowing when a classifier's prediction can be trusted is useful in many applications and critical for safely using AI. While the bulk of the effort in machine learning research has been towards improving classifier performance, understanding when a classifier's predictions should and should not be trusted has received far less attention. The standard approach is to use the classifier's discriminant or confidence score; however, we show there exists an alternative that is more effective in many situations. We propose a new score, called the trust score, which measures the agreement between the classifier and a modified nearest-neighbor classifier on the testing example. We show empirically that high (low) trust scores produce surprisingly high precision at identifying correctly (incorrectly) classified examples, consistently outperforming the classifier's confidence score as well as many other baselines. Further, under some mild distributional assumptions, we show that if the trust score for an example is high (low), the classifier will likely agree (disagree) with the Bayes-optimal classifier. Our guarantees consist of non-asymptotic rates of statistical consistency under various nonparametric settings and build on recent developments in topological data analysis.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Density estimation in representation space to predict model uncertainty

    cs.LG 2019-08 conditional novelty 4.0 of 10

    A learned classifier over k-nearest-neighbor statistics in a pretrained network's representation space predicts misclassifications and detects out-of-distribution images without training on out-of-distribution examples.

  2. Energy-Aware Deep Learning on Resource-Constrained Hardware

    cs.LG 2025-05 conditional novelty 1.0 of 10

    A survey of energy-aware deep learning methods for resource-constrained devices, covering energy-aware design, adaptive inference, on-device training, and scheduling on energy-harvesting systems.

Pith tools