{"id":"8c178f63-966e-4ba2-b750-3c55eeb5e485","arxiv_id":"2506.02312","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"DEFFA-Unet, a dual-encoder attention U-Net with custom augmentation, is presented for retinal vessel segmentation, but its claimed state-of-the-art performance is not consistently supported by the reported numbers.","lead":"A new dual-encoder U-Net with attention fusion and custom data balancing is proposed for retinal vessel segmentation. The paper claims state-of-the-art generalization, but its own reported tables do not consistently show superiority, and the augmentation method may leak target data.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The cross-domain generalization claim in Tables 5-6 depends on SOTA-CSA color statistics whose reference dataset is unspecified in Section 2.1.2; if those statistics include the target datasets, the reported gains are leakage rather than generalization.","rationale":"I agree with the reader's REJECT verdict but locate the load-bearing issue differently. The domain-invariance assumption could be wrong, but the paper includes an ablation (w/o D-I in Table 6) that at least tests it; the more urgent problem is the unspecified \"SOTA dataset\" in the augmentation. A generalization comparison cannot be validated if target statistics may have been used to generate training data. This is not an accusation; the paper is simply underspecified at the exact point where the headline result lives. The same ambiguity also helps explain the internal contradiction between the abstract's \"significantly outperforms\" claim and the Discussion's disclaimer, and the absence of code or error bars makes the claim uncheckable. Because the ambiguity is concrete and directly targets the central generalization claim, the rejection should stand. If the authors clarify that \"SOTA\" means source-only statistics and reproduce the numbers, a conditional acceptance would be feasible.","tokens_in":22144,"tokens_out":3991,"duration_ms":38034,"concrete_test":"Rerun the Table 5 protocol changing only the reference dataset used in Eqs. (4)-(5): (a) SOTA-CSA statistics from the DRIVE training split only; (b) statistics from each target dataset (STARE/CHASE/HRF); (c) augmentation disabled. If (b) reproduces the reported STARE/CHASE/HRF AUC/MCC gains over (a), the reported generalization advantage is explained by target-statistics leakage; if (a) matches the paper's numbers, publicly release the dataset list to document that no target statistics entered training. One run with fixed seeds and identical hyperparameters is sufficient to disambiguate.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is the significant cross-validation generalization advantage. The direct evidence is Tables 5 and 6, which compare models trained on one dataset and tested on others. Those tables are only interpretable if the SOTA-CSA augmentation in Section 2.1.2 uses statistics from the training source domain alone. Equations (4)-(6) compute the mean and standard deviation \"within SOTA dataset\" and apply them to the source dataset, but the paper never states which datasets constitute \"SOTA.\" In Table 5 the model is trained on DRIVE and tested on STARE, CHASE, and HRF; in Table 6 it is trained on IOSTAR and tested on STARE, CHASE, and DRIVE. If the \"SOTA dataset\" is (or includes) the target/test dataset, the augmentation injects target-domain color statistics into training. Section 3.3.2 claims no fine-tuning with target datasets, but using target statistics for augmentation is a different, unspecified form of information transfer. This confound matters: Table 6 shows SOTA-CSA contributes a large part of the cross-domain margin (STARE AUC 0.9415 without CSA vs. 0.9834 with CSA), and no code or error bars are provided. The Discussion even disclaims superiority, so the abstract's \"significantly outperforms\" is not supported unless this ambiguity is resolved.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes DEFFA-Unet, a dual-encoder U-Net variant for retinal vessel segmentation, together with two data-processing methods: JESB (Jaccard-based synthetic balancing) and SOTA-CSA (color-statistics-based augmentation), and a high-frequency domain-invariant input branch. The authors report experiments on DRIVE, STARE, CHASE, HRF, and IOSTAR, claiming improved segmentation accuracy and, especially, cross-domain generalization. The paper contains model descriptions, quantitative comparisons with prior methods, cross-domain evaluations, and an ablation study.","tokens_in":22421,"tokens_out":2989,"duration_ms":26136,"significance":"If the claims were substantiated, a lightweight dual-encoder network with domain-invariant preprocessing and targeted augmentation would constitute a practically useful contribution to retinal vessel segmentation, and the cross-domain generalization results would be of interest to the medical imaging community. The paper does provide a well-structured architecture description and a reasonably broad experimental setup. However, the reported evidence is internally inconsistent and does not support the stated conclusions; the proposed method is not numerically superior to several compared methods on the primary metric in the paper's own tables, and the cross-domain augmentation source is ambiguous. The potential significance is therefore not realized in the current manuscript.","major_comments":[{"comment":"The claim that 'our proposed method achieved the best scores in all reported metrics' is directly contradicted by the paper's own tables: on DRIVE the proposed method's DSC is 0.8247 versus 0.8310 for CSGNet and 0.8303 for DPFNet; on STARE the proposed DSC is 0.8111 versus 0.8516 for CSUNet and 0.8493 for CSGNet; on CHASE the proposed DSC is 0.8090 versus 0.8302 for DPFNet; on IOSTAR the proposed DSC is 0.8182 versus 0.8516 for SCS-Net. The proposed method also does not achieve the best recall in several rows. The manuscript provides no error bars, confidence intervals, or significance tests for any of these comparisons, so the abstract's and Section 1's claims of 'significantly outperforms' are not supported by the presented evidence.","section":"Section 3.3.1, Tables 3 and 4"},{"comment":"The 'SOTA dataset' D used to compute the color statistics μ_C and σ_C is never defined. The cross-domain experiments in Tables 5 and 6 train on one dataset (e.g., DRIVE or IOSTAR) and test on others. If the 'SOTA dataset' includes or overlaps with the target test sets, then SOTA-CSA injects target-domain color statistics into the training data, and the reported cross-domain gains are a form of leakage rather than evidence of generalization. The paper must unambiguously state that the statistics are computed from the training source domain only, and preferably verify this by ablating the augmentation source. Without this clarification, the central cross-domain generalization claim is uninterpretable.","section":"Section 2.1.2, Equations (4)-(6)"},{"comment":"The core assumption underlying the second encoder, that high-frequency components of retinal images are domain-invariant and preserve vessel structure across datasets, is asserted without proof or empirical validation. The entire generalization benefit is attributed to this branch, and the 'w/o D-I' ablation in Table 6 lacks error bars or repeated runs. The authors should provide evidence for this assumption, for instance by analyzing the distribution of high-frequency features across domains or by comparing the proposed high-frequency preprocessing to the original FDA-based approach with a controlled experiment.","section":"Section 2.1.3, Equations (7)-(10)"},{"comment":"The Discussion states: 'It is crucial to clarify that this paper does not claim superiority of the proposed method over others.' This directly contradicts the Abstract ('the proposed method significantly outperforms the compared methods in cross-validation model generalization'), the Introduction, and Section 3.3.1, which claim the best scores in all reported metrics. This is not a minor wording issue: it makes the paper's central claim incoherent, and it also suggests that the authors themselves do not stand behind the headline results. The manuscript needs to be rewritten so that the claims match the actual numerical evidence.","section":"Section 4 (Discussion)"}],"minor_comments":[{"comment":"There is a typo 'CHSE' in Section 2.1.1, presumably meaning CHASE; also the dataset overview in Table 1 lists CHASE as having 28 total images with 'R:28, H:0', but the text in Section 3.1 says CHASEDB1 was divided into 20 training and 8 testing images, which is inconsistent with the total count of 28 and the healthy/retinopathy split.","section":"Table 1"},{"comment":"The header 'Trained on DRIVE (D) Leave-1 (D, S, C, H)' is confusing and appears to mix two separate training protocols; this makes it difficult to interpret the baseline comparisons and should be clarified.","section":"Table 5"},{"comment":"In Table 5, the method labeled 'HGC-Net' is cited as [45], but the reference list associates [45] with SCS-Net and [46] with the self-supervision boosted method; the citation numbering appears inconsistent.","section":"References"},{"comment":"The MCC value '0.68774' is reported with an extra digit compared to the other entries; this looks like a transcription error.","section":"Table 5"},{"comment":"The abbreviation 'w/o D-I' is not defined in the table caption; the text later explains it as 'without domain-invariant feature guidance', but the caption should state this.","section":"Table 6"}],"recommendation":"reject","confidential_remarks":"The manuscript contains internal contradictions that, in my view, cannot be resolved through minor revision: the claimed 'best scores' are refuted by the paper's own tables, and the Discussion explicitly disclaims the superiority claim. The SOTA-CSA leakage ambiguity is fixable in principle, but the deeper issue is that the reported evidence does not support the paper's central generalization claim. I would recommend rejection on those grounds."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a legitimate engineering paper whose architecture and ablations are worth a look, but the evidence as presented does not support the stated claims of state-of-the-art superiority or cross-domain generalization. The central problem is not the design but the evaluation.\n\nWhat is new: DEFFA-Unet is a new combination of known components — a second encoder on high-frequency domain-invariant inputs, a channel/spatial feature filtering fusion module, an attention-based skip-connection replacement, plus two data-side tricks: JESB (a KMeans-SMOTE variant using Jaccard distance on masks) and SOTA-CSA (normalization-based color augmentation). The ablation study shows each module adds something on DRIVE, which is credible. The comparison set is broad, including several recent dual-encoder and cascaded methods, and the parameter footprint is modest. That part is fine engineering.\n\nWhere it goes soft. First, Section 3.3.1 says 'our proposed method achieved the best scores in all reported metrics,' but Tables 3 and 4 show the method trails CSGNet on DRIVE DSC, CSUNet on STARE DSC, DPFNet on CHASE DSC, and SCS-Net on IOSTAR DSC. The abstract also claims 'significantly outperforms' in cross-validation generalization, while the Discussion says 'this paper does not claim superiority.' You can't have both. That is a real contradiction, not a wording quibble.\n\nSecond, and more load-bearing: SOTA-CSA. Equations (4)-(6) compute color statistics from 'SOTA dataset D' and apply them to the source training set, but the paper never states which datasets D includes. In Tables 5-6 the models are trained on one dataset and tested on others. If D includes the target/test datasets, then the augmentation has injected target-domain color statistics into training, and the cross-domain gains, particularly the large AUC jump on STARE with SOTA-CSA in Table 6, are leakage rather than generalization. The paper needs to state explicitly that D is the source training set only, and ideally show the sensitivity. Without that, the central generalization claim is unverifiable.\n\nThird, no error bars or significance tests, no code, and free parameters (alpha_loss, alpha_enhance, window_size, alpha_blend) are not sensitivity-analyzed. That is minor-to-moderate for this subfield, but combined with the above it means the headline numbers are not reliable.\n\nThe assumption that high-frequency image components are domain-invariant is asserted, not proven, but that is a hypothesis the dual-encoder design could test; an ablation removing it is already present (Prop. w/o D-I), so it's at least partially addressed.\n\nWho it's for: anyone working on retinal vessel segmentation or domain generalization augmentation. The architecture is plausible and could be salvaged with a clearer evaluation protocol. I would not cite it in its current form, but I'd accept it for peer review and ask for a major revision that fixes the claims, specifies the SOTA dataset, adds error bars, and releases code.","headline":"A plausible engineering combination, but the headline generalization claim is compromised by an unspecified augmentation reference set and contradicted by the paper's own tables.","tokens_in":23018,"tokens_out":3873,"would_cite":false,"duration_ms":31611,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A dual-encoder U-Net that adds an enhanced high-frequency input stream claims better retinal vessel segmentation and stronger cross-dataset generalization than single-encoder baselines.","keywords":["retinal vessel segmentation","dual encoder","domain-invariant features","attention mechanism","feature fusion","model generalization","data balancing","data augmentation"],"falsifier":"A direct check is to compare the pixel-value distributions of the enhanced high-frequency maps at vessel locations across DRIVE and STARE: if those distributions are no closer to each other than the raw color distributions are, the premise that these components are domain-invariant fails.","tokens_in":21919,"feed_emoji":"👁️","tokens_out":11012,"duration_ms":93556,"temperature":0.7,"pith_summary":"The paper proposes a retinal-vessel segmentation network, DEFFA-Unet, built on the idea that the high-frequency part of a fundus image carries the vessel structure and stays roughly unchanged when the imaging device or domain changes. To exploit this, it adds a second encoder that consumes an enhanced high-frequency map of the same input while the first encoder reads the raw image, and it replaces ordinary skip connections with attention-guided fusion. The authors claim that this dual encoding, together with a clustering-based data-balancing scheme and a color-statistics augmentation, outperforms U-Net, Attention U-Net, UNet++ and prior state-of-the-art models on four fundus benchmarks and one SLO benchmark, and that it generalizes across datasets better than multi-domain methods without target-domain fine-tuning. If true, the practical payoff is a vessel segmenter that transfers from one hospital's imaging setup to another without new annotations.","feed_headline":"Dual-encoder U-Net tops five retinal vessel segmentation tests","feed_subtitle":"A second encoder fed high-frequency maps boosts accuracy and cross-dataset transfer.","key_machinery":"The load-bearing mechanism is the domain-invariant input stream. Each raw image $I$ is converted to a high-frequency map by subtracting a local average, $H(x,y)=I(x,y)-L_{\\mathrm{avg}}(x,y)$, then amplifying it by a global factor $G$ computed from the variance of $H$. Encoder-2 consumes $I_{\\mathrm{enhanced}}=G\\cdot H$, while Encoder-1 consumes the raw image. Before the bottleneck, the Feature Filtering Fusion (FFF) module applies channel attention to the domain-specific features (\"what to emphasize\") and spatial attention to the domain-invariant features (\"where to look\"), then convolves the concatenation. Replacing each skip connection, the Feature Reconstructing Fusion (FRF) module concatenates low-level features from both encoders, splits them into dilated-context and local paths, and uses fused high-level features to produce attention weights for the reconstructed low-level map. JESB balances training sets by clustering binary vessel masks with Jaccard distance and synthesizing samples for small clusters, while SOTA-CSA augments by mixing each image with the color statistics of reference datasets.","core_discovery":"The paper's discovery claim is that the high-frequency part of a fundus image is domain-invariant, so a second encoder fed an enhanced version of that signal can give a U-Net both richer vessel features and substantially better generalization. DEFFA-Unet is this dual-encoder design: Encoder-1 reads the raw image, Encoder-2 reads $I_{\\mathrm{enhanced}}=G\\cdot H$, where $H$ is the image minus its local average and $G$ is a variance-derived enhancement factor. A Feature Filtering Fusion (FFF) module applies channel attention to the domain-specific stream and spatial attention to the domain-invariant stream before the bottleneck, and a Feature Reconstructing Fusion (FRF) module replaces skip connections with attention-weighted reconstruction from both low- and high-level features. With Jaccard-distance-based synthetic balancing and color-statistics augmentation, the paper reports top or near-top accuracy, recall, specificity, precision, Dice, IoU, AUC and MCC on DRIVE, CHASEDB1, STARE, HRF and IOSTAR, and reports that leave-one-out cross-domain and cross-modality tests beat or match multi-domain methods without target-domain fine-tuning.","pith_inferences":["Beyond the paper: if high-frequency domain invariance is a general property of curvilinear structures, the same second-stream design should transfer to other medical targets such as corneal nerves or coronary vessels; this is untested here.","Beyond the paper: the reported gains bundle the dual encoder, both fusion modules, and the augmentation schemes together; isolating each contribution with a noise-fed second encoder would show which ingredient actually carries the generalization.","Beyond the paper: because SOTA-CSA only needs color statistics from a reference set, using unlabeled statistics from the target device would turn the method into a lightweight unsupervised adaptation step."],"forward_implications":["A vessel segmenter trained on one fundus dataset could be deployed on unseen datasets without target-domain fine-tuning, reducing the need for new annotations for every scanner or patient population.","The attention-guided replacement of skip connections is designed to cut false positives, the failure mode the paper says is most costly in clinical screening.","The JESB balancing and SOTA-CSA augmentation recipes use only cheap image-level operations and could be dropped into other segmentation pipelines.","The model matches or exceeds multi-domain methods while using a small parameter footprint (2.85 MB in the reported tables), so the reported generalization gain is not bought by model scale."],"supporting_citations":[{"why":"The encoder-decoder baseline whose architecture and skip connections the paper extends and evaluates against.","marker":"[10]"},{"why":"Nested skip-connection baseline re-trained under the same protocol for comparison.","marker":"[13]"},{"why":"Supplies the FOV masks for CHASE and STARE and is the main cross-domain comparison method.","marker":"[20]"},{"why":"The dual-encoding U-Net variant that motivates the two-encoder design and is compared on all datasets.","marker":"[22]"},{"why":"An alternative dual-path progressive fusion network used as a comparison in the segmentation and generalization tables.","marker":"[23]"},{"why":"The clustering-plus-synthesis method JESB adapts, replacing its distance with Jaccard distance.","marker":"[34]"},{"why":"The normalization-guided augmentation approach that SOTA-CSA adapts into color-space mixing.","marker":"[35]"},{"why":"Source of the idea that high-frequency components carry transferable vessel structure across domains.","marker":"[38]"},{"why":"The attention-gate approach that motivates the Feature Reconstructing Fusion module.","marker":"[42]"},{"why":"A leave-one-out domain-generalization method compared against on STARE, CHASE, and HRF.","marker":"[36]"}],"fun_headline_variants":["Dual-encoder U-Net wins five retinal vessel segmentation benchmarks","High-frequency second encoder sharpens vessel segmentation across datasets","Attention-guided dual-encoder U-Net beats state-of-the-art on vessel tests","Domain-invariant high-frequency input boosts U-Net vessel segmentation","DEFFA-Unet: dual encoders filter features for top retinal vessel results"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the high-frequency components of a fundus image carry the vessel structure and stay roughly constant when the acquisition device changes, and the paper asserts this without directly measuring that stability.","fun_headline_variants_meta":{"raw":{"variants":["Dual-encoder U-Net wins five retinal vessel segmentation benchmarks","High-frequency second encoder sharpens vessel segmentation across datasets","Attention-guided dual-encoder U-Net beats state-of-the-art on vessel tests","Domain-invariant high-frequency input boosts U-Net vessel segmentation","DEFFA-Unet: dual encoders filter features for top retinal vessel results"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000253,"raw_usage":{"total_tokens":1595,"prompt_tokens":1005,"completion_tokens":590,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":621,"completion_tokens_details":{"reasoning_tokens":511}},"tokens_in":621,"tokens_out":590,"duration_ms":5281,"temperature":1.0,"reasoning_tokens":511,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:26:26.852883+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct check is to compare the pixel-value distributions of the enhanced high-frequency maps at vessel locations across DRIVE and STARE: if those distributions are no closer to each other than the raw color distributions are, the premise that these components are domain-invariant fails.","supporting_citations":[],"review_version":1}