{"id":"67028cfd-c105-4013-b665-c829ae4b431d","arxiv_id":"2607.17691","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"A routine stacking of HSV luminosity correction and CLAHE reports higher contrast metrics than HE/AHE baselines, but an undefined metric reference and a vessel-map-as-disease-label evaluation break the supporting evidence.","lead":"This paper stacks two standard image-processing steps — HSV-based lighting correction and CLAHE contrast enhancement — for retinal photos and reports better PSNR/SSIM/CNR than histogram baselines on three public datasets. The claimed 'disease detection' is a 50-pixel brightness-area threshold scored against vessel-segmentation annotations, and the metric reference is never defined, so the central claims are not yet supported.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Quantitative claims rest on ground-truth labels the cited datasets do not contain; both PSNR/SSIM/CNR and disease-detection metrics are unauditable.","rationale":"The reader's weakest_assumption identifies the same core problem: the evaluation uses datasets whose annotations are vessel segmentations, not disease labels, and the full-reference metrics lack a defined reference. My stress-test converges on this as the single most load-bearing concern because it invalidates both the enhancement claim and the screening claim. The paper provides no way to audit the numbers: if the ground truth is not what the authors claim, then the reported PSNR/SSIM/CNR and accuracy/sensitivity/specificity/AUC are not meaningful. I considered whether the internal contradiction in the method description (Figure 1 vs. §2.3) might be more load-bearing, but that is a presentation error that a revision could fix. The ground-truth mismatch is structural and cannot be patched without redoing the experiments on appropriate datasets. Therefore the reader's REJECT verdict stands; my concern does not change it.","tokens_in":7846,"tokens_out":3433,"duration_ms":37386,"concrete_test":"Download the DRIVE, STARE, and CHASEDB1 ground-truth files and enumerate all annotation types. If the only labels are vessel/fovea segmentation masks and no image-level disease diagnoses or reference enhanced images exist, attempt to reproduce Table 4 using these masks as labels and Table 2 using any plausible reference (original image, CLAHE-only output, or manual gold standard). Failure to reproduce the reported values with any defined reference confirms the metrics are undefined.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's two central claims—enhancement superiority (PSNR/SSIM/CNR in Table 2) and screening-level disease detection (Table 4)—both depend on ground-truth data that DRIVE, STARE, and CHASEDB1 do not provide. Section 3.2 states disease detection is evaluated 'against expert ophthalmologist annotations from DRIVE, STARE, and CHASEDB1,' but the only expert annotations in these benchmarks are pixel-level vessel segmentations (and DRIVE fovea masks); there are no image-level disease labels. Thus the reported 87.4% accuracy and 0.869 AUC cannot be a disease-detection result unless the authors invented or substituted a different label source, which is not disclosed. Table 2 reports PSNR, SSIM, and CNR—full-reference metrics—yet the reference image is never defined. DRIVE/STARE/CHASEDB1 contain no ground-truth enhanced images; the only possible references are the original inputs or the vessel masks, and neither would yield the stated numbers (e.g., PSNR vs. the original would be infinite for the no-op baseline and lower for any enhancement). Without a defined reference or genuine disease labels, the quantitative validation is not reproducible and the conclusions in §4.4 and §5 ('suitable for clinical screening') are unsupported. This is the load-bearing flaw: it undermines both headline contributions, not just a peripheral metric.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a two-stage retinal fundus image enhancement pipeline: HSV-based luminosity correction followed by CLAHE applied to the V (or luminance) channel, together with a binary Otsu thresholding step that flags 'hyper-reflective regions' as a disease-detection mechanism. The authors report PSNR/SSIM/CNR improvements over HE, AHE, and CLAHE-only baselines on DRIVE, cross-dataset generalization to STARE and CHASEDB1, and disease-detection accuracy/sensitivity/specificity/AUC around 84–90%. They claim the method is statistically superior, requires no training data or GPU, and is suitable for clinical screening at 0.14 s per image.","tokens_in":8065,"tokens_out":4060,"duration_ms":44967,"significance":"If the reported results were valid, the contribution would be practically significant: a simple, fast, training-free enhancement and screening pipeline could be deployed in resource-limited settings. The paper also makes a good-faith effort to report standard deviations, statistical tests, and cross-dataset consistency. However, the quantitative validation has two foundational flaws: full-reference metrics (PSNR, SSIM, CNR) are computed without any defined reference image, and the 'disease detection' evaluation uses datasets that contain only vessel-segmentation annotations, not image-level disease labels. These flaws affect both headline claims and make the reported numbers unauditable. The paper does not provide a reproducible evaluation protocol, so its central conclusions are unsupported.","major_comments":[{"comment":"PSNR, SSIM, and CNR are full-reference metrics, yet the reference image is never defined. DRIVE, STARE, and CHASEDB1 provide original fundus photographs and pixel-level vessel segmentations, but no 'ground-truth enhanced' image. If the reference were the original input, the no-op baseline would have infinite PSNR and SSIM=1, which is inconsistent with the HE value of 21.4 dB / 0.74. Without a stated reference, all quantitative enhancement results — including the Wilcoxon significance tests — are not reproducible and the headline claim of superiority is unverifiable.","section":"§4.2, Table 2"},{"comment":"The disease-detection evaluation is described as using 'expert ophthalmologist annotations from DRIVE, STARE, and CHASEDB1,' but these benchmarks contain pixel-level vessel and fovea segmentations, not image-level disease labels. The reported 87.4% accuracy and 0.869 AUC therefore do not measure disease detection; the 50-pixel hyper-reflective area rule can at best be compared against vessel maps. The conclusion in §4.4 and §5 that the method is 'suitable for clinical screening' is unsupported by the data. This is a category mismatch, not a minor phrasing issue.","section":"§3.2, Table 4"},{"comment":"The method description and Figure 1 are internally inconsistent. The text states Stage 1 is HSV decomposition with Gaussian luminance gain and no gamma correction, and Stage 2 is CLAHE on the V channel. The Figure 1 caption refers to 'gamma correction' and 'CLAHE applied to the L* channel in L*a*b* space.' This is a direct contradiction that prevents readers from knowing the actual pipeline. Since the paper's central claim is a specific enhancement method, this inconsistency is load-bearing for reproducibility.","section":"§2.4 and Figure 1"},{"comment":"The 50-pixel area threshold is said to be 'set on DRIVE training partition to exclude noise artefacts,' but DRIVE has no disease labels, so this threshold cannot have been selected to optimize disease classification. In addition, AUC with DeLong's method is claimed for what appears to be a single fixed threshold; no ROC construction or threshold sweep is described. The sensitivity/specificity estimates in Table 4 are therefore not statistically grounded. The 'hyper-reflective area' proxy is presented as a validated biomarker without any clinical evidence.","section":"§2.5, §3.1, Table 4"}],"minor_comments":[{"comment":"The abstract states experiments are conducted on 'the publicly available DRIVE dataset,' while the full text uses DRIVE, STARE, and CHASEDB1. Please update the abstract to reflect the multi-dataset design.","section":"Abstract and §2"},{"comment":"CNR is never defined. Please provide the formula used for contrast-to-noise ratio, including which regions are considered signal and noise.","section":"§4.2"},{"comment":"The Gaussian smoothing in Eq. (1) is not normalized; please clarify how G(x,y) is scaled before division to avoid division by very small values near the image border.","section":"§2.3, Eq. (1)"},{"comment":"Reference [5] (Son et al., J Digit Imaging 2019) is titled 'Towards accurate segmentation of retinal vessels and the optic disc,' not an enhancement method. The listed PSNR/SSIM for this reference appear to be a citation mismatch. Also, the deep-learning comparisons are not on identical image subsets, as acknowledged, so the 'comparable' statement should be softened.","section":"Table 5"},{"comment":"Wilcoxon signed-rank statistics (W, p) are reported only for the CLAHE-only comparison; the other comparisons lack effect sizes or test statistics. Multiple-comparison correction is not discussed.","section":"§4.2"}],"recommendation":"reject","confidential_remarks":"The two loading-bearing flaws — undefined reference for full-reference metrics and the use of vessel-segmentation datasets as disease ground truth — are not correctable by a revision within the current scope. To salvage the work, the authors would need new data (image-level disease labels) and a fundamentally different evaluation protocol (e.g., no-reference metrics or a reader study), which goes beyond a standard revision. I also note the figure/text inconsistency suggests the manuscript may not have been carefully checked; however, that is secondary to the validation issues."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nYou should know two things before reading this one: the numbers in Tables 2 and 3 are unauditable because PSNR/SSIM/CNR need a reference image and none is defined, and the disease-detection claim in Table 4 rests on ground truth the datasets do not contain. DRIVE, STARE, and CHASEDB1 are vessel-segmentation benchmarks; they carry pixel-level vessel and fovea annotations, not image-level disease labels. The paper calls these 'expert ophthalmologist annotations' and reports 87.4% accuracy and 0.869 AUC as if they were disease labels. That is a category mismatch, not a methodological quibble.\n\nWhat is worth crediting: the pipeline is simple and cheap—HSV luminosity normalization followed by V-channel CLAHE—and 0.14 s per image on CPU is plausible. Fixing hyperparameters on DRIVE and transferring them to STARE and CHASEDB1 is the right instinct, and the Wilcoxon testing is appropriate. The paper is clearly written apart from a serious internal contradiction: Figure 1 describes L*a*b* and gamma correction, while the text and equations use HSV and Gaussian division in the V channel. The references are real and mostly on point, but the novelty is modest; luminosity correction and CLAHE both appear in the cited literature, and stacking them is an engineering choice, not a new principle.\n\nSoft spots in proportion: the undefined reference metric is load-bearing. Without a stated reference, the headline PSNR=29.3 dB, SSIM=0.91, CNR=3.12 cannot be reproduced or interpreted. The disease-detection section has the same structural problem, and the 50-pixel threshold is tuned on DRIVE training data. Table 5 compares numbers from other papers with different protocols, which is not apples-to-apples. No code or data is shipped. These are not peripheral issues; both central claims collapse without them.\n\nWho is it for: a reader looking for a lightweight CPU preprocessing recipe might find the pipeline plausible, but the evidence as written does not support it. I would not send this version to peer review. If the authors define the reference, replace the ground truth with genuine disease labels, and reconcile the method description, it could become a small, honest engineering note.\n\nRecommendation: desk reject in current form; invite a revision along those lines.","headline":"Both headline claims rest on undefined or mismatched ground truth—the enhancement metrics have no reference image, and the 'disease' labels are actually vessel segmentations—so the paper's quantitative results are not auditable.","tokens_in":8690,"tokens_out":5169,"would_cite":false,"duration_ms":50597,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"HSV luminosity normalization before CLAHE beats standard enhancement and flags vascular disease via a simple brightness rule, with no training data.","keywords":["CLAHE","HSV color space","luminosity correction","retinal fundus imaging","contrast enhancement","disease screening","PSNR/SSIM evaluation"],"falsifier":"Take a set of healthy retinal images with no disease and dense vessel patterns; if the 50-pixel hyper-reflective-area rule flags them as diseased, the screening threshold is measuring vessel density, and the reported AUC against true clinical labels will not reproduce.","tokens_in":7559,"feed_emoji":"👁️","tokens_out":7743,"duration_ms":64047,"temperature":0.7,"pith_summary":"CLAHE alone cannot remove the smooth, non-uniform illumination that makes retinal images hard to read, the paper argues, so it first divides the Value channel of an HSV decomposition by a large-kernel Gaussian-smoothed version of itself. On three public benchmarks, this two-stage pipeline reports PSNR 29.3 dB, SSIM 0.91, and CNR 3.12, exceeding histogram equalization, adaptive histogram equalization, and CLAHE alone with statistical significance. The same fixed parameters transfer across datasets without retuning. A thresholded version of the enhanced image, kept only when bright connected regions exceed 50 pixels, is then claimed to flag vascular disease at 87.4 percent accuracy, 84.3 percent sensitivity, and 90.1 percent specificity in 0.14 seconds per image on a CPU. The contribution is a training-free, GPU-free screening pipeline; the caveat is that the ground-truth maps used for screening may be vessel segmentations rather than disease labels.","feed_headline":"Luminosity fix plus CLAHE beats standard retinal enhancement","feed_subtitle":"A no-training pipeline hits 87 percent screening accuracy at 0.14 seconds per image on public benchmarks.","key_machinery":"The central mechanism is the two-stage pipeline. Stage 1 constructs a luminance gain surface G(x,y) by convolving the HSV Value channel with a Gaussian kernel (σ=60 px) and divides the channel by it, V_corrected = V/G, normalizing the illumination field while preserving relative intensity variations. Stage 2 applies CLAHE (clip limit 0.01, 8×8 tile grid) to the corrected V channel, recombines with unchanged H and S, and converts back to RGB. The identity that does the work is the division by the smoothed luminance: it flattens the background, making a single global threshold meaningful across images with different optics. Screening then uses a variance-minimizing global threshold to binarize","core_discovery":"The discovery is that replacing the luminance channel with a Gaussian-smoothed gain surface before applying contrast-limited adaptive histogram equalization (CLAHE) removes slowly varying illumination, letting local contrast enhancement amplify vessel and lesion detail without amplifying noise. The paper reports statistically significant gains on the primary test set—PSNR rising from 21.4 dB for HE and 23.1 dB for AHE to 29.3 dB, SSIM from 0.74/0.79 to 0.91, CNR from 1.82/2.10 to 3.12—with fixed parameters carrying over to two additional datasets. It further claims that a global variance-minimizing threshold on the enhanced image, followed by connected-component labeling and a 50-pixel brigh","pith_inferences":["If the 50-pixel bright-area rule is actually capturing vessel density rather than pathology, the reported accuracy and AUC are measuring the wrong thing; a direct test is to run the same rule on healthy retinas with dense vasculature.","The paper's comparison with deep learning uses different experimental protocols, so a like-for-like benchmark on the same images and reference would be needed before claiming parity.","The full-reference metrics (PSNR/SSIM/CNR) require a reference image that the paper never names; re-evaluating with a public reference or a no-reference metric would settle whether the 29.3 dB gain is meaningful."],"forward_implications":["As a training-free, GPU-free pipeline running in 0.14 s per image, it could slot into low-cost screening workflows where deep-learning alternatives are unavailable.","Fixed parameters carrying over to three benchmarks suggests the pipeline is robust to camera optics and population differences, a practical advantage for deployment.","The 87.4 percent accuracy, if it reproduces under true disease labels, would catch most positives while keeping referrals low, making it a plausible first-pass filter.","The ordering—luminosity correction before local contrast enhancement—is a general recipe that could improve other CLAHE-based medical image pipelines."],"fun_headline_variants":["Luminosity-adaptive CLAHE beats HE and AHE on retinal images","Retinal enhancement: CLAHE with luminosity fix wins on all metrics","CLAHE after luminosity correction improves fundus contrast and fidelity","No-training retinal enhancement beats standard methods in 0.14s","Adaptive contrast boost on corrected brightness tops retinal benchmarks"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The screening claim holds only if the expert annotation maps used as ground truth are disease labels, but they are pixel-level vessel segmentations, and the enhancement metrics need a reference image the paper never provides.","fun_headline_variants_meta":{"raw":{"variants":["Luminosity-adaptive CLAHE beats HE and AHE on retinal images","Retinal enhancement: CLAHE with luminosity fix wins on all metrics","CLAHE after luminosity correction improves fundus contrast and fidelity","No-training retinal enhancement beats standard methods in 0.14s","Adaptive contrast boost on corrected brightness tops retinal benchmarks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000926,"raw_usage":{"total_tokens":3875,"prompt_tokens":888,"completion_tokens":2987,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":632,"completion_tokens_details":{"reasoning_tokens":2899}},"tokens_in":632,"tokens_out":2987,"duration_ms":19513,"temperature":1.0,"reasoning_tokens":2899,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T17:16:13.584528+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a set of healthy retinal images with no disease and dense vessel patterns; if the 50-pixel hyper-reflective-area rule flags them as diseased, the screening threshold is measuring vessel density, and the reported AUC against true clinical labels will not reproduce.","supporting_citations":[],"review_version":1}