Pith. sign in

REVIEW 4 major objections 6 minor 20 references

Adaptive Class Learning to Screen Diabetic Disorders in Fundus Images of Eye

T0 review · 4 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read This paper claims a two-stage fine-tuning scheme, CELD, lets a fundus classifier add Glaucoma as a third class without forgetting Healthy and Diabetic Retinopathy knowledge, reaching an overall accuracy of 0.9100 on pooled public datasets.

desk verdict Two-stage fine-tuning dressed as incremental learning; the forgetting claim isn't tested, but the perturbation analysis shows some clinical thought. read the letter →

arxiv 2501.12048 v1 pith:PY6EQ6BF submitted 2025-01-21 cs.CV cs.AI

classification cs.CVcs.AI
keywords fundusimagesclassextensionlimiteddatacatastrophicforgettingdiabeticretinopathyglaucomaclassificationDenseNet121explainability
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper proposes a two-stage training scheme, Class Extension with Limited Data (CELD), for classifying retinal fundus images. A DenseNet121 network is first trained to separate Healthy from Diabetic Retinopathy (DR) images, and its optimized weights are then used to initialize a second network that is fine-tuned on an expanded dataset also containing Glaucoma images. The authors claim this lets the model pick up the new class without catastrophic forgetting of the original two classes, even when labeled data for the new class is scarce. On pooled public fundus datasets the three-class CELD model reaches an overall accuracy of 0.9100, outperforming direct three-class training of the same backbone models, and perturbation experiments show the model leans on the green channel and the optic disc.

What carries the argument

The load-bearing mechanism is sequential fine-tuning with parameter reuse: a DenseNet121 classifier $C_S$ is trained on the two-class source domain, and its optimized weights initialize a three-class classifier $C_T$ that is fine-tuned on $\tau_T$ with cross-entropy loss. DenseNet121's dense connectivity supplies strong gradient flow and feature reuse, which the paper argues reduces overfitting on imbalanced data. The framework's formal condition $\tau_S \subset \tau_T$ is what lets the target task keep the source task's images in the training set, so the model is never required to reconstruct old knowledge from memory alone. The explainability component uses six controlled perturbations—reducing the green channel, randomly removing green segments, reducing contrast, adding Gaussian noise, edge sharpening, and optic disc occlusion—to probe which image features drive the model's decisions.

What would settle it

Train the second stage on Glaucoma images only, or on a target set with the Healthy and DR images removed, then measure Healthy and DR accuracy on the original test set; if accuracy drops well below the reported 0.8729 two-class level, the claim that CELD prevents catastrophic forgetting is contradicted.

Watch

Extended reading notes

Core claim

The central claim is that a classifier can be extended from two classes (Healthy, DR) to three classes (Healthy, DR, Glaucoma) by fine-tuning from the source classifier's weights on an expanded dataset, rather than retraining from scratch. The paper formalizes this with source data $\tau_S$ and target data $\tau_T$ satisfying $\tau_S \subset \tau_T$, trains DenseNet121 on the two-class source task, then reuses the optimized weights $\omega_S^*$ to initialize classifier $C_T$ and minimizes cross-entropy on the three-class target task. The reported result is an overall accuracy of 0.9100 with noticeably higher F1-scores for DR and Glaucoma than direct three-class training with SeResNet101, DenseNet121, or ViT. The perturbation analysis is part of the same contribution: degrading the green channel or occluding the optic disc degrades DR and Glaucoma classification, indicating which input features the model relies on.

Load-bearing premise

The framework's claim of preventing catastrophic forgetting is only tested when the expanded training set still contains all the original Healthy and DR images; if the intended use case is adding a class without access to the old data, the forgetting claim is not actually demonstrated.

Editorial extensions

If this is right

  • A screening model can grow from two to three disease classes by fine-tuning on a small batch of new-class images, without retraining from scratch.
  • The 0.9100 accuracy on pooled public data suggests the approach is usable for DR and Glaucoma screening where labeled Glaucoma images are scarce.
  • If the perturbation findings hold, the model's decisions are driven by clinically sensible features: the green channel for DR lesions and the optic disc for Glaucoma.
  • DenseNet121 is a suitable backbone for this two-stage extension, outperforming SeResNet101 and ViT in the paper's comparisons.
  • The same weight-transfer recipe could be applied as new ocular disease classes are added over time in a deployed screening system.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A stronger test of the forgetting claim would be to fine-tune on Glaucoma images alone, or on a target set that excludes most Healthy and DR images; the current protocol keeps all source-task data in $\tau_T$, so catastrophic forgetting is not truly measured.
  • The perturbation results could be turned into a prospective clinical hypothesis: a fundus screening tool that emphasizes green-channel contrast and optic-disc structure should transfer to other camera types better than one using full-color texture only.
  • The same two-stage extension pattern could be tried for other scarce ocular conditions, such as age-related macular degeneration, with the caveat that the paper's Glaucoma F1-score of 0.6667 remains the weakest, so rare-class performance will likely still lag.
  • A direct extension would be to add a replay buffer or regularization penalty when old data cannot be stored, which would turn CELD into a genuinely incremental learner rather than a two-stage fine-tuner.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper proposes a two-stage framework, CELD, for retinal fundus image classification. A DenseNet121 classifier is first trained to distinguish Healthy from Diabetic Retinopathy (DR) images, and is then fine-tuned on an expanded dataset that also includes Glaucoma images, pooled from Messidor2, Chaksu, and LES-AV. The authors report an overall three-class accuracy of 0.9100 and a perturbation-based explainability analysis. The central claim is that CELD enables class extension with limited new data while preventing catastrophic forgetting.

Significance. If the claimed behavior were established, the framework would be relevant to incremental medical image classification under data scarcity. The use of public datasets and the inclusion of perturbation-based explainability are positive aspects. However, the core claim is not supported by the experimental design: the second stage fine-tunes on the full expanded dataset containing all old-class data, so the setup does not test catastrophic forgetting. In addition, all accuracy numbers come from a single train/validation/test split with no confidence intervals, multiple runs, or per-class test-set sizes, and the backbone is selected on the same pooled data used for final evaluation. As it stands, the contribution is not established.

major comments (4)
  1. [2.1, Eq. (4)] The experimental protocol cannot support the claim that CELD 'prevents catastrophic forgetting of previous learning, while leveraging existing knowledge to learn new classes in the presence of limited data.' Since the manuscript states τS ⊂ τT and Eq. (4) minimizes the loss over all N samples of τT, the second stage retrains on every Healthy and DR image from the source task. The model is therefore never updated from new-class data alone, and no forgetting can occur by construction; no scenario is tested in which old-class data is unavailable. To support the claim, the second stage should be trained only on Glaucoma samples (or a small replay set), and the paper should report per-class accuracy on the original two classes before and after the second stage, together with a standard fine-tuning baseline. Without such experiments, the main contribution is unsupported.
  2. [3.4, Fig. 3 and Table 2] All reported accuracies, including the headline 0.9100, come from a single train/validation/test split with no confidence intervals and no multiple runs. With only 199 Glaucoma images in the pooled dataset, the test split contains roughly 20 Glaucoma samples, so the reported Glaucoma F1-score of 0.6667 has very wide uncertainty. The paper should report per-class test-set sizes, confidence intervals or standard deviations over multiple seeds and splits, and ideally a bootstrap or permutation-based assessment of the accuracy difference between CELD and the baselines.
  3. [2.2 and 3.4] DenseNet121 was selected as the backbone because it achieved the highest two-class accuracy on the same pooled dataset that is later used to report the CELD result. This is model selection on the evaluation distribution and can inflate the final reported accuracy. Architecture selection should be performed on a separate validation split, or the entire comparison should be repeated with a fixed, pre-specified backbone, in order to avoid circularity in the comparison with SeResNet101 and ViT.
  4. [3.4, perturbation analysis] The perturbation experiments are described qualitatively: confusion matrices and F1-scores are shown, but there are no error bars, significance tests, or repeated perturbation instantiations. Statements such as 'significantly decreased performance' and 'highly depends on the optic disc' are therefore not statistically supported. At minimum, the authors should report the mean and standard deviation of each metric over multiple perturbation runs and state the number of test images used for each class and perturbation type.
minor comments (6)
  1. [3.3] The F1-score equation contains a typo: 'Precsion' should be 'Precision'.
  2. [3.4] The three-class results are presented only in Fig. 3 without a corresponding numerical table; a table with precision, recall, and F1 per class would improve readability and reproducibility.
  3. [3.1] The split description gives only percentages; the paper should report the exact number of training, validation, and test images per class and per source dataset, since the class imbalance is substantial.
  4. [1] The statement that CELD is 'unlike transfer learning' is not substantiated, because the proposed procedure is itself a fine-tuning method; the distinction should be clarified.
  5. [References] Reference [8] has a typo in the title: 'Rethinking ImageNet pre-training' should be 'Rethinking ImageNet Pre-training'.
  6. [general] No code or trained models are released, which limits reproducibility; the authors should state availability or provide a public implementation.

Circularity Check

0 steps flagged · score 0.0 of 10

No derivation-level circularity; the paper is an empirical fine-tuning study and its equations do not reduce a claimed prediction to an input. The central forgetting claim is an experimental validity issue rather than a circular one.

full rationale

CELD is presented as two-stage fine-tuning (Eqs. 1-5), with the second-stage loss minimized over the expanded dataset τT. Although τS⊂τT means old-class data is re-used in the fine-tuning loss, the paper does not derive a numerical prediction from this equation; it reports measured accuracies (0.9100) on a held-out split. The claim that CELD 'prevents catastrophic forgetting' is therefore an empirical assertion, not a quantity that reduces by construction to the training objective. The only mild concern is that DenseNet121 was selected on the same pooled dataset used in the final evaluation (Sec. 2.2 and Sec. 3.4), which is a model-selection/validation issue, not an equation-level circularity. No load-bearing self-citation chain is present: the cited works by the same group (Refs. 1-2) support background statements only. The paper's main weakness—never testing the setting where old-class data is unavailable during fine-tuning—is a correctness/validity limitation that should be raised separately, not scored as derivation-level circularity.

Assumptions & free parameters 4 free parameters · 3 assumptions · 0 invented entities

The central claims rest on two unstated supporting assumptions: that the dataset pooling is legitimate and that the expanded fine-tuning setup is a meaningful test of catastrophic forgetting. Neither is verified. The free parameters are mostly standard hyperparameters, but the backbone selection on the same evaluation data is a notable model-selection bias.

free parameters (4)
  • initial learning rate = 1e-5
    Set in Section 3.2 without sensitivity analysis; the final accuracy may depend on this choice.
  • batch size = 8
    Set in Section 3.2 without justification or sensitivity analysis.
  • backbone architecture = DenseNet121
    Chosen because it had the highest accuracy among SeResNet101, DenseNet121, and ViT on the same pooled dataset, as reported in Section 3.4.
  • early stopping criterion = not specified
    Mentioned in Section 3.2 but the patience and validation metric are not stated, making the training procedure underdetermined.
assumptions (3)
  • domain assumption The pooled fundus images from Messidor2, Chaksu, and LES-AV are label-consistent and can be treated as one domain.
    Section 3.1 pools images from different cameras, populations, and acquisition protocols. Label noise or domain shift would directly affect the reported accuracy.
  • ad hoc to paper Fine-tuning on τT, where τS is a subset, constitutes incremental class learning and tests catastrophic forgetting.
    Section 2.1 defines τS ⊂ τT and then fine-tunes on the full τT, so old-class data is always present. This makes the setup unsuitable for detecting forgetting, yet the paper claims anti-forgetting benefits.
  • domain assumption The ground truth labels in the three public datasets are correct and uniform.
    Section 3.1 relies on labels from Messidor2, Chaksu, and LES-AV without verifying label quality or grading criteria across datasets.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Adaptive Class Learning to Screen Diabetic Disorders in Fundus Images of Eye." pith.science (2026). https://pith.science/paper/PY6EQ6BF

@misc{pith2026250112048,
  author       = {Pith},
  title        = {Pith review of: Adaptive Class Learning to Screen Diabetic Disorders in Fundus Images of Eye},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PY6EQ6BF}},
  note         = {Machine review of arXiv:2501.12048}
}
read the original abstract

The prevalence of ocular illnesses is growing globally, presenting a substantial public health challenge. Early detection and timely intervention are crucial for averting visual impairment and enhancing patient prognosis. This research introduces a new framework called Class Extension with Limited Data (CELD) to train a classifier to categorize retinal fundus images. The classifier is initially trained to identify relevant features concerning Healthy and Diabetic Retinopathy (DR) classes and later fine-tuned to adapt to the task of classifying the input images into three classes: Healthy, DR, and Glaucoma. This strategy allows the model to gradually enhance its classification capabilities, which is beneficial in situations where there are only a limited number of labeled datasets available. Perturbation methods are also used to identify the input image characteristics responsible for influencing the models decision-making process. We achieve an overall accuracy of 91% on publicly available datasets.

Figures

Figures reproduced from arXiv: 2501.12048 by the authors.

Figure 1
Figure 1. The basic workflow of the proposed CELD-Framework humans have an inherent ability to learn new skill over the time without for￾getting prior knowledge. In this work, the proposed CELD framework exploits this notion of natural learning ability by retaining the knowledge acquired from previously learned classes to enable the network to adapt to new class. This re￾duces the requirement for extensive datasets for each i… view at source ↗
Figure 2
Figure 2. The original image and its perturbated versions for each of the classes: DR, Healthy and Glaucoma. 3.1 Dataset A total of 3,111 retinal color fundus images were obtained from three publicly available datasets: Messidor2 3 , Chaksu [11], and LES-AV [17]. The Messidor2 dataset has 1,744 macula-centered RGB images. There are 1017 images belong￾ing to the healthy class and 727 images belonging to the DR category. The Ch… view at source ↗
Figure 3
Figure 3. Quantitative Result for 3 Class Classification [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Quantitative Result over CELD framework with Data Perturbation [PITH_FULL_IMAGE:figures/full_fig_p011_4.png]
Figure 5
Figure 5. Figure 5: Confusion matrix over CELD framework with Data Perturbation. The matrix shows performance of CELD (a) With no perturbation, (b) Reduce green (RG) , (c) Random green removal (RGR), (d) Reducing image contrast (RC), (e) Gaussian noise (GN), (f) Edge sharpening (ES), (g) …

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

20 extracted references · 18 canonical work pages

  1. [1]

    In: Proceedings of the 43rd Annual International Conference of the IEEE Engineering in Medicine & Biology Society (EMBC)

    Basu, S., Mitra, S.: Segmentation in diabetic retinopathy using deeply-supervised multiscalar attention. In: Proceedings of the 43rd Annual International Conference of the IEEE Engineering in Medicine & Biology Society (EMBC). pp. 2614–2617. IEEE (2021)

  2. [2]

    In: 2020 IEEE symposium series on computational intelligence (SSCI)

    Basu, S., Mitra, S., Saha, N.: Deep learning for screening covid-19 using chest x-ray images. In: 2020 IEEE symposium series on computational intelligence (SSCI). pp. 2521–2527. IEEE (2020)

  3. [3]

    The Lancet Global Health1, e339–e349 (2013)

    Bourne, R.R., Stevens, G.A., et al.: Causes of vision loss worldwide, 1990–2010: A systematic analysis. The Lancet Global Health1, e339–e349 (2013)

  4. [4]

    arXiv preprint arXiv:2010.11929 (2020)

    Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al.: An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929 (2020)

  5. [5]

    In- formatics and Health1, 57–69 (2024)

    Fatima, M., Pachauri, P., et al.: Enhancing retinal disease diagnosis through AI: Evaluating performance, ethical considerations, and clinical implementation. In- formatics and Health1, 57–69 (2024)

  6. [6]

    Computers in Biology and Medicine152, 106391 (2023) 14 S

    Garcea, F., Serra, A., et al.: Data augmentation for medical imaging: A systematic literature review. Computers in Biology and Medicine152, 106391 (2023) 14 S. Dey et al

  7. [7]

    In: Proceedings of the IEEE 5th International Conference on Cybernetics, Cognition and Machine Learning Applications (ICC- CMLA)

    Grover, K.S., Kapoor, N.: Detection of glaucoma and diabetic retinopathy using fundus images and deep learning. In: Proceedings of the IEEE 5th International Conference on Cybernetics, Cognition and Machine Learning Applications (ICC- CMLA). pp. 407–412. IEEE (2023)

  8. [8]

    : Rethinking ImageNet pre-training

    He, K., Girshick, R., et al. : Rethinking ImageNet pre-training. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 4918–4927 (2019)

Show all 20 references
  1. [9]

    In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (June 2018)

    Hu, J., Shen, L., Sun, G.: Squeeze-and-excitation networks. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (June 2018)

  2. [10]

    In: Proceed- ingsoftheIEEEconferenceonComputerVisionandPatternRecognition(CVPR)

    Huang, G., Liu, Z.,et al.: Densely connected convolutional networks. In: Proceed- ingsoftheIEEEconferenceonComputerVisionandPatternRecognition(CVPR). pp. 4700–4708 (2017)

  3. [11]

    Scientific Data10, 70 (2023)

    Kumar, J.H., Seelamantula, C.S., et al.: Cháks.u: A glaucoma specific fundus image database. Scientific Data10, 70 (2023)

  4. [12]

    Nature521, 436–444 (2015)

    LeCun, Y., Bengio, Y., Hinton, G.: Deep learning. Nature521, 436–444 (2015)

  5. [13]

    In: International Conference on Learning Representations (2019),https://openreview.net/forum? id=Bkg6RiCqY7

    Loshchilov, I., Hutter, F.: Decoupled weight decay regularization. In: International Conference on Learning Representations (2019),https://openreview.net/forum? id=Bkg6RiCqY7

  6. [14]

    Computers and Electrical Engineering 117, 109243 (2024)

    Madarapu, S., Ari, S., et al.: A multi-resolution convolutional attention network for efficient diabetic retinopathy classification. Computers and Electrical Engineering 117, 109243 (2024)

  7. [15]

    World Journal of Clinical Cases11, 3736 (2023)

    Morya, A.K., Ramesh, P.V., et al.: Diabetes more than retinopathy, it’s effect on the anterior segment of eye. World Journal of Clinical Cases11, 3736 (2023)

  8. [16]

    Expert Systems with Applications 249, 123418 (2024)

    Navaneethan,R.,Devarajan,H.:Enhancingdiabeticretinopathydetectionthrough preprocessing and feature extraction with MGA-CSG algorithm. Expert Systems with Applications 249, 123418 (2024)

  9. [17]

    In: Medical Image Computing and Computer Assisted Intervention–MICCAI 2018: 21st International Confer- ence, Granada, Spain, September 16-20, 2018, Proceedings, Part II 11

    Orlando, J.I., Barbosa Breda, J., et al.: Towards a glaucoma risk index based on simulated hemodynamics from fundus images. In: Medical Image Computing and Computer Assisted Intervention–MICCAI 2018: 21st International Confer- ence, Granada, Spain, September 16-20, 2018, Proce...

  10. [18]

    Healthcare Analytics 5, 100303 (2024)

    Shamrat, F.J.M., Shakil, R., et al.: An advanced deep neural network for fundus image analysis and enhancing diabetic retinopathy detection. Healthcare Analytics 5, 100303 (2024)

  11. [19]

    : Three types of incremental learning

    Van de Ven, G.M., Tuytelaars, T., et al. : Three types of incremental learning. Nature Machine Intelligence4, 1185–1197 (2022)

  12. [20]

    Medical Image Analysis11, 555–566 (2007)

    Walter, T., Massin, P., et al.: Automatic detection of microaneurysms in color fundus images. Medical Image Analysis11, 555–566 (2007)

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.