Pith. sign in

REVIEW 3 major objections 6 minor 14 references

Uncertainty-aware last-layer heads can improve safety operating points for referable diabetic retinopathy screening, but only internal gains that fail under dataset shift are not enough for trustworthy claims.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 10:18 UTC pith:CKTDZM7T

load-bearing objection Useful empirical safety hygiene paper: last-layer uncertainty helps on APTOS selective referral, fails under APTOS→DDR transfer, and the author is honest about both. the 3 major comments →

arxiv 2607.02569 v1 pith:CKTDZM7T submitted 2026-06-30 cs.CV cs.LG

Uncertainty-Aware Last-Layer Adaptation of RETFound for Referable Diabetic Retinopathy Screening Under Dataset Shift

classification cs.CV cs.LG
keywords diabetic retinopathy screeningRETFounduncertainty estimationlast-layer adaptationselective referraldataset shiftBayesian neural networksSNGP
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper asks a practical safety question about frozen RETFound retinal features plus small last-layer heads for referable diabetic retinopathy screening: can uncertainty-aware heads produce safer operating points than a deterministic cached-feature baseline, and do those gains survive a second dataset? On APTOS, uncertainty-aware heads improved high-sensitivity and selective-referral behavior; the strongest selective-referral result deferred about 20% of cases and drove accepted-case false negatives to zero while keeping high accepted-case specificity. The paper also shows that ordinary threshold tuning can cut false negatives at large false-positive cost, so that reduction is not unique to Bayesian modeling. On DDR, native Bayesian heads moved in the same qualitative direction but with weaker tradeoffs, and an APTOS-trained SNGP-style checkpoint transferred poorly and failed as an external selective-referral signal. The central message is that safety-coverage tradeoffs and second-dataset validation under shift are required before claiming trustworthy retinal screening behavior.

Core claim

Uncertainty-aware last-layer heads on frozen RETFound features can improve internal safety-oriented operating points for referable diabetic retinopathy screening, especially selective referral, but false-negative reduction is not unique to Bayesian modeling and uncertainty signals that look useful in-distribution can fail under external dataset shift; therefore trustworthy claims require explicit safety-coverage evaluation and second-dataset validation.

What carries the argument

Uncertainty-aware last-layer adaptation on frozen 1024-d RETFound features: cached softmax, temperature scaling, variational Bayesian and diagonal Laplace last-layer heads, and an SNGP-style random-feature head, evaluated by full-coverage metrics, threshold sweeps, and selective-referral coverage/referral tradeoffs on APTOS with native and transfer checks on DDR.

Load-bearing premise

The load-bearing premise is that frozen RETFound features plus small last-layer heads are a fair and sufficient testbed for screening-oriented uncertainty claims under shift, so observed internal gains and transfer failures can be attributed mainly to last-layer uncertainty rather than to representation mismatch or a single external dataset.

What would settle it

A controlled re-run that fine-tunes or re-fits the same uncertainty families on multiple independent external retinal cohorts and still finds that selective-referral ranking and zero-accepted-false-negative operating points collapse under shift would falsify the claim that last-layer uncertainty plus safety-coverage evaluation is enough to support trustworthy screening claims in this setting.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This paper evaluates uncertainty-aware last-layer adaptation of frozen RETFound features for binary referable diabetic retinopathy screening on APTOS 2019 and DDR. It compares a cached-feature softmax head, temperature scaling, variational Bayesian last-layer heads, a diagonal Laplace approximation, and an SNGP-style cached-feature head. The central claim is carefully mixed: on APTOS, uncertainty-aware heads improve high-sensitivity and selective-referral operating points (notably ~20% deferral with zero accepted-case false negatives and high accepted specificity for SNGP entropy), but false-negative reduction is not unique to Bayesian modeling because threshold lowering can achieve the same at high false-positive cost; on DDR, native Bayesian heads qualitatively reproduce the direction with weaker tradeoffs, while an APTOS-trained SNGP checkpoint fails under direct transfer. The authors conclude that trustworthy screening claims require explicit safety–coverage evaluation and second-dataset validation under shift.

Significance. If the results hold, the paper is a useful empirical contribution to safety-centered evaluation of foundation-model adaptation in medical imaging. Its main value is not a new SOTA method but the disciplined framing: threshold-sweep controls, selective-referral metrics, and an explicit negative transfer result that prevent overclaiming uniqueness of Bayesian last-layer uncertainty. The public repository, locked summary tables, and reproducible asset scripts are genuine strengths. The work is incremental methodologically (frozen features + standard last-layer uncertainty families) but addresses a practical gap—how screening systems should be judged when false negatives and human-review burden matter more than aggregate accuracy. That framing is relevant to clinical ML evaluation even if the specific last-layer recipes are not novel.

major comments (3)
  1. Results §5.3 and Methods §3.4: the strongest APTOS claim (SNGP predictive entropy at ~80% coverage → zero accepted-case FN with specificity 0.9146) depends on an uncertainty ranking and coverage chosen on the same validation set used for model/checkpoint selection (Tables 3–4: sens vs val-loss checkpoints). Without a pre-specified signal+coverage protocol or nested holdout confirming ranking quality, the headline selective-referral gain can be partly selection-driven. Given that the paper itself shows ranking collapse under shift (Table 6, Fig. 5: APTOS→DDR SNGP still leaves hundreds of accepted FN at the same coverage), the internal zero-FN result needs a fixed evaluation protocol or held-out confirmation to support the safety claim as stated.
  2. Methods §3.2–3.3 and transfer protocol §4.3 / Table 7: attributing the APTOS→DDR SNGP failure primarily to last-layer uncertainty under shift is under-supported. The experiment freezes RETFound_mae natureCFP 1024-d features and restores an APTOS-fitted diagonal precision state without refit. Representation mismatch, domain-specific random-feature geometry, or the non-refit precision proxy could dominate the collapse (sens 0.2701, FN 616). The paper should either (i) re-fit the last-layer precision/head on DDR as a control, or (ii) clearly limit the claim to “this cached-feature SNGP checkpoint does not transfer without retraining,” rather than treating the result as general evidence about uncertainty robustness under shift.
  3. Results Tables 3–6: key safety metrics (FN, sensitivity, selective-referral accepted-case FN) are reported as point estimates on single fixed splits (APTOS val/test n=366; DDR val/test ~1878–1879) with no multi-seed variance, bootstrap intervals, or McNemar-style paired tests. On APTOS, moving from 17 FN (softmax) to 3 FN (SNGP sens) or to 0 accepted FN under deferral is load-bearing for the internal-safety narrative; without uncertainty on those counts, it is hard to judge whether the gains are stable or split-specific. At minimum, report variability over seeds or resampling for the primary safety operating points.
minor comments (6)
  1. Abstract and §5.1: “deferred approximately 20 percent of cases” is clear, but the corresponding APTOS selective-referral table is not numbered alongside the DDR Table 6; adding an APTOS selective-referral table (coverage, FN, Spec, Acc) would make the strongest result easier to verify.
  2. Methods §3.3: variational Bayesian and Laplace heads are described only at a high level (diagonal Gaussian posteriors, KL term, diagonal Laplace). Specify prior/posterior parameterization, number of MC samples at inference, and KL weight so that the comparison is reproducible from the text alone, not only the repository.
  3. Figure 1–6 captions are informative, but several figures are referenced without axis units or exact operating-point labels in the text; ensure each figure states the dataset split (val vs test) used.
  4. Related Work: FusionFM [Zou et al., 2025] is cited as arXiv:2508.11721; if that preprint post-dates or is concurrent with this work, clarify the relationship so readers do not infer dependence.
  5. Table 1 / §3.1: binary mapping (grades 0–1 non-referable; 2–4 referable) is clinically standard but should briefly note any disagreement with alternative referable definitions used in prior DR screening papers for comparability.
  6. §6.5 Limitations correctly notes single external dataset and frozen backbone; consider also stating that temperature scaling cannot change ranking-based confusion matrices at fixed threshold (already in §3.3) earlier when interpreting Softmax+Temp rows in Tables 3–4.

Circularity Check

0 steps flagged

Empirical last-layer comparison on public datasets; no derivation reduces to its inputs by construction.

full rationale

This paper is a safety-centered empirical evaluation of cached-feature last-layer heads (softmax, temperature scaling, variational Bayes, diagonal Laplace, SNGP-style) on frozen RETFound features for binary referable DR, with APTOS internal results and DDR second-dataset / transfer checks. There is no first-principles derivation chain, uniqueness theorem, or self-citation load-bearing premise: references are external (Zhou et al. RETFound, Guo temperature scaling, Blundell/Ritter/Daxberger Bayesian last layers, Liu SNGP, selective classification literature). Reported operating points are measured performance under stated selection policies (val-loss vs max-sens vs balanced; fixed ~80% selective-referral coverage; threshold sweeps), not quantities forced by fitting a parameter and renaming it a prediction. The paper explicitly controls against overclaiming (threshold tuning also drives FN to zero at high FP cost; APTOS→DDR SNGP is a negative transfer result). Mild model-selection on validation is ordinary experimental practice and does not make accepted-case FN/specificity equal the selection objective by construction. Score 0 with empty steps is the correct finding.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 2 invented entities

The work is empirical ML evaluation. Load-bearing premises are standard domain choices (binary referable mapping, frozen foundation features, last-layer uncertainty families) plus free choices of operating points and coverage levels. No new physical entities are postulated; invented entities are only named model variants used as experimental conditions.

free parameters (5)
  • decision threshold for positive class
    Swept and selected for balanced accuracy or lowest FN; directly determines reported FN/FP operating points on both datasets.
  • selective-referral coverage (~80% accepted)
    Chosen reporting operating point for deferred-case analyses; accepted-case metrics are conditional on this coverage.
  • Bayesian last-layer variational posterior / KL weighting
    Diagonal Gaussian variational head hyperparameters control uncertainty quality and the sensitivity-selected checkpoints.
  • SNGP random Fourier features and diagonal precision proxy
    RFF dimension/scale and precision estimate define the SNGP-style head whose internal success and transfer failure are central.
  • temperature scaling parameter
    Fitted post hoc on validation logits for calibration comparison without changing ranking at fixed threshold.
axioms (5)
  • domain assumption Grades 0–1 are non-referable and grades 2–4 are referable DR for binary screening evaluation.
    Section 3.1 mapping; defines the clinical label the entire study optimizes.
  • domain assumption Frozen RETFound_mae natureCFP 1024-d features are an adequate representation for last-layer screening comparisons.
    Section 3.2; isolates last-layer uncertainty but assumes the frozen encoder does not dominate failure modes.
  • domain assumption Selective referral metrics on accepted cases are the appropriate safety lens, distinct from full-coverage thresholding.
    Sections 3.4–3.5 and Discussion 6.4; frames what counts as a successful safety result.
  • standard math Standard approximate Bayesian and distance-aware last-layer constructions (VI, diagonal Laplace, SNGP-style RFF) are valid uncertainty estimators for this setting.
    Methods 3.3 citing Blundell, Ritter/Daxberger, Liu et al.; used without new theory.
  • ad hoc to paper APTOS-trained SNGP precision state evaluated on DDR without refit is a meaningful transfer test of shift robustness.
    Section 4.3 protocol choice; a negative result under this protocol is treated as evidence against unvalidated robustness claims.
invented entities (2)
  • SNGP-style cached-feature head (not full end-to-end SNGP RETFound) no independent evidence
    purpose: Provide a distance-aware last-layer uncertainty baseline on frozen features.
    Paper-specific implementation variant; independent evidence is limited to the reported APTOS/DDR experiments.
  • Sensitivity-selected last-layer checkpoints (Bayes/Laplace/SNGP) no independent evidence
    purpose: Define high-sensitivity operating points for safety comparison.
    Selection policy invented for the study’s safety narrative rather than a standard pretrained artifact.

pith-pipeline@v1.1.0-grok45 · 13048 in / 3458 out tokens · 28373 ms · 2026-07-12T10:18:46.018395+00:00 · methodology

0 comments
read the original abstract

This paper presents a safety-centered empirical evaluation of uncertainty-aware last-layer adaptation for referable diabetic retinopathy screening using RETFound, a self-supervised vision-transformer retinal foundation model used here as a frozen feature encoder, and the public APTOS 2019 and DDR diabetic retinopathy fundus image datasets. We compare a cached-feature softmax head, post-hoc temperature scaling, variational Bayesian last-layer heads, a diagonal Laplace last-layer approximation, and an SNGP-style cached-feature head. On APTOS, uncertainty-aware operating points improved sensitivity and selective-referral behavior. The strongest APTOS selective-referral result deferred approximately 20 percent of cases and reduced accepted-case false negatives to zero while preserving high accepted-case specificity. However, threshold tuning also reduced false negatives at high false-positive cost, so false-negative reduction alone was not unique to Bayesian modeling. On DDR, native Bayesian heads qualitatively reproduced the APTOS direction but with weaker tradeoffs, while the APTOS-trained SNGP checkpoint transferred poorly and failed to provide useful external selective-referral behavior. These results highlight the value of safety-centered evaluation beyond aggregate accuracy: uncertainty-aware last-layer heads can improve internal safety-oriented operating points, but trustworthy retinal screening claims require explicit safety-coverage evaluation and second-dataset validation under shift.

Figures

Figures reproduced from arXiv: 2607.02569 by Karim Mardhani.

Figure 1
Figure 1. Figure 1: APTOS full-coverage false-negative and false-positive counts for key internal models. [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: DDR full-coverage false-negative and false-positive counts. The APTOS-trained SNGP [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: DDR sensitivity and specificity comparison for native DDR heads and the APTOS-to [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: DDR threshold-sweep comparison across balanced and low-false-negative operating [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Accepted-case false negatives at approximately 80% coverage for DDR Bayesian and [PITH_FULL_IMAGE:figures/full_fig_p008_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: APTOS internal validation versus APTOS-to-DDR transfer for the sensitivity-selected [PITH_FULL_IMAGE:figures/full_fig_p009_6.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

14 extracted references · 1 linked inside Pith

  1. [1]

    Nature , year =

    A Foundation Model for Generalizable Disease Detection from Retinal Images , author =. Nature , year =

  2. [2]

    arXiv preprint arXiv:2508.11721 , year =

    FusionFM: Fusing Eye-specific Foundational Models for Optimized Ophthalmic Diagnosis , author =. arXiv preprint arXiv:2508.11721 , year =

  3. [3]

    2019 , howpublished =

    APTOS 2019 Blindness Detection , author =. 2019 , howpublished =

  4. [4]

    2019 , volume =

    Li, Tao and Gao, Yutong and Wang, Kang and Guo, Song and Liu, Hailin and Kang, Hong , journal =. 2019 , volume =

  5. [5]

    1996 , doi =

    Bayesian Learning for Neural Networks , author =. 1996 , doi =

  6. [6]

    Proceedings of the 32nd International Conference on Machine Learning , year =

    Weight Uncertainty in Neural Networks , author =. Proceedings of the 32nd International Conference on Machine Learning , year =

  7. [7]

    International Conference on Learning Representations , year =

    A Scalable Laplace Approximation for Neural Networks , author =. International Conference on Learning Representations , year =

  8. [8]

    Advances in Neural Information Processing Systems , year =

    Laplace Redux -- Effortless Bayesian Deep Learning , author =. Advances in Neural Information Processing Systems , year =

  9. [9]

    Advances in Neural Information Processing Systems , year =

    Simple and Principled Uncertainty Estimation with Deterministic Deep Learning via Distance Awareness , author =. Advances in Neural Information Processing Systems , year =

  10. [10]

    Proceedings of the 34th International Conference on Machine Learning , year =

    On Calibration of Modern Neural Networks , author =. Proceedings of the 34th International Conference on Machine Learning , year =

  11. [11]

    Journal of Machine Learning Research , year =

    On the Foundations of Noise-Free Selective Classification , author =. Journal of Machine Learning Research , year =

  12. [12]

    Advances in Neural Information Processing Systems Workshops , year =

    Selective Classification for Deep Neural Networks , author =. Advances in Neural Information Processing Systems Workshops , year =

  13. [13]

    JAMA , year =

    Development and Validation of a Deep Learning Algorithm for Detection of Diabetic Retinopathy in Retinal Fundus Photographs , author =. JAMA , year =

  14. [14]

    JAMA , year =

    Development and Validation of a Deep Learning System for Diabetic Retinopathy and Related Eye Diseases Using Retinal Images from Multiethnic Populations with Diabetes , author =. JAMA , year =