Pith. sign in

REVIEW 2 cited by

Benchmarking Bayesian Deep Learning on Diabetic Retinopathy Detection Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.12717 v1 pith:37OXMVG4 submitted 2022-11-23 stat.ML cs.AIcs.CVcs.LG

classification stat.MLcs.AIcs.CVcs.LG
keywords deeplearningtasksbayesianmethodsbenchmarkpredictivereal-world
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Bayesian deep learning seeks to equip deep neural networks with the ability to precisely quantify their predictive uncertainty, and has promised to make deep learning more reliable for safety-critical real-world applications. Yet, existing Bayesian deep learning methods fall short of this promise; new methods continue to be evaluated on unrealistic test beds that do not reflect the complexities of downstream real-world tasks that would benefit most from reliable uncertainty quantification. We propose the RETINA Benchmark, a set of real-world tasks that accurately reflect such complexities and are designed to assess the reliability of predictive models in safety-critical scenarios. Specifically, we curate two publicly available datasets of high-resolution human retina images exhibiting varying degrees of diabetic retinopathy, a medical condition that can lead to blindness, and use them to design a suite of automated diagnosis tasks that require reliable predictive uncertainty quantification. We use these tasks to benchmark well-established and state-of-the-art Bayesian deep learning methods on task-specific evaluation metrics. We provide an easy-to-use codebase for fast and easy benchmarking following reproducibility and software design principles. We provide implementations of all methods included in the benchmark as well as results computed over 100 TPU days, 20 GPU days, 400 hyperparameter configurations, and evaluation on at least 6 random seeds each.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools

    cs.LG 2025-05 conditional novelty 5.0 of 10

    The paper proposes adding the entropy of an LLM's final answer to the entropy of an external tool's output as an uncertainty score for tool-calling QA systems, and shows it predicts answer correctness on synthetic and...

  2. Entropy-Based Non-Invasive Reliability Monitoring of Convolutional Neural Networks

    cs.CV 2025-08 reject novelty 3.0 of 10

    A CNN's activation entropy separates clean from FGSM-attacked image batches in a small VGG-16 test, but fitted binning, tiny samples, and contradictory numbers weaken the claim.

Pith tools