Pith. sign in

REVIEW 1 cited by

Failure Detection in Medical Image Classification: A Reality Check and Benchmarking Testbed

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.14094 v2 pith:CK275V4B submitted 2022-05-27 cs.AI

classification cs.AI
keywords detectionclassificationfailureimagingmedicalmethodsbenchmarkingcheck
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Failure detection in automated image classification is a critical safeguard for clinical deployment. Detected failure cases can be referred to human assessment, ensuring patient safety in computer-aided clinical decision making. Despite its paramount importance, there is insufficient evidence about the ability of state-of-the-art confidence scoring methods to detect test-time failures of classification models in the context of medical imaging. This paper provides a reality check, establishing the performance of in-domain misclassification detection methods, benchmarking 9 widely used confidence scores on 6 medical imaging datasets with different imaging modalities, in multiclass and binary classification settings. Our experiments show that the problem of failure detection is far from being solved. We found that none of the benchmarked advanced methods proposed in the computer vision and machine learning literature can consistently outperform a simple softmax baseline, demonstrating that improved out-of-distribution detection or model calibration do not necessarily translate to improved in-domain misclassification detection. Our developed testbed facilitates future work in this important area

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beyond scalar losses: calibrating segmentation models via gradient vector field surgery

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Scaling region-based segmentation loss gradients by per-voxel prediction error improves calibration metrics on 2D/3D medical segmentation datasets with little or no loss in Dice score.

Pith tools