{"id":"1a3f1acc-6bec-4420-b39a-c68ede9d431d","arxiv_id":"2506.07925","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A literature survey of U-Net variants for remote sensing change detection that compiles reported results from prior papers without running any new experiments.","lead":"This preprint reviews U-Net based deep learning architectures for detecting changes in satellite images, summarizing reported performance from prior papers. It is a survey without new experiments, and its comparative conclusions are weakened by mixing results from different datasets and metrics.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's central comparative claim is unsupported because Table 4 ranks U-Net variants using scores from different datasets, sensors, resolutions, and metrics, without any controlled common-evaluation protocol.","rationale":"I read the paper in good faith. It is a literature survey that aims to help practitioners choose U-Net variants for remote sensing change detection. The central claim, as stated in the abstract and Section 1, is that a comparison of U-Net variations provides guidance on which design choices improve change detection. For that claim to hold, the performance evidence must support comparisons across variants. The paper does not provide new experiments; its quantitative evidence is Table 4, which aggregates numbers from separate papers evaluated on separate datasets with separate metrics. Because the scores are not normalized, not paired, and lack variance estimates, the conclusion that NDR-U-Net is the best (mIoU 0.92) or that Siamese Swin-U-Net has a strong F1 is an artifact of the source studies rather than a result of this paper's analysis. The inconsistency between the abstract's '18 variations' and the body's '8 variants' further undermines the comprehensiveness claim, though the primary flaw remains the uncontrolled comparison. I agree with the reader's weakest assumption and with the REJECT verdict, so no adjustment is needed. I would not manufacture an additional objection: the bibliography may be a useful starting point, but the paper as a comparative study does not substantiate its conclusions.","tokens_in":9791,"tokens_out":3021,"duration_ms":37163,"concrete_test":"Construct a matrix from Tables 3 and 4 with rows as the eight variants and columns as dataset, sensor, resolution, metric, and reported score; then identify all pairs that share the same dataset and the same metric. If, as the tables currently suggest, no such directly comparable pair exists, the claims that NDR-U-Net 'achieves the highest Mean IoU' and that the variants are 'top contenders' cannot be supported by the presented data. As an additional check, run the eight variants on a common benchmark such as LEVIR-CD with identical preprocessing, optimizer, patch size, and evaluation script, and compare F1/mIoU with confidence intervals; if the ranking changes materially, the Table 4 ordering is an artifact of non-comparable reporting.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's stated purpose is to provide a comparative assessment of U-Net variants for change detection, and its only empirical evidence is Table 4. That table places side-by-side: Siamese Swin-U-Net F1=94.67 on CDD (Google Earth, 0.03-1 m), STCD-EffV2T U-Net mIoU=0.87 on OSCD/Sentinel-2 (Iran), T-U-Net 'Pixel Accuracy higher accuracy' on LEVIR-CD/WHU-CD/DSIIN-CD, Ensemble U-Net-ResNet OA=89.20% on HRS/WorldView-2 (Beijing), Optimised U-Net OA=92.10% on Vaihingen/LiDAR, Bilateral Attention U-Net mIoU=0.845 on Gaofen-2 (Seoul), NDR-U-Net mIoU=0.92 on GF-2 (Europe), and HARNU-Net mIoU=0.88 on GF-2 (New Zealand/Texas). These numbers differ in dataset difficulty, class balance, resolution, sensor, and metric; none share a common evaluation protocol or report error bars. From these rows the text concludes that 'NDR-U-Net achieves the highest Mean IoU (0.92)' and that Optimized U-Net and NDR-U-Net are 'top contenders,' but a higher number from one study is not evidence of a better model when nothing is controlled. The abstract also promises 18 variants and a 'comprehensive analysis of 34 papers,' yet Section 5 states 'we study 8 variants' and Table 2 lists 8; no inclusion/exclusion criteria are given for the 34 papers. Because the contribution is entirely the comparison, this incomparability is load-bearing: without a shared benchmark or explicit control for dataset and protocol, the ranking collapses into a list of reported numbers.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript presents a narrative review of U-Net architecture variants applied to remote sensing change detection. It surveys 34 papers, categorizes eight U-Net variants (Siamese Swin-U-Net, STCD-EffV2T U-Net, T-U-Net, Ensemble U-Net-ResNet, Optimised U-Net, Bilateral Attention U-Net, NDR-U-Net, HARNU-Net), and tabulates their reported performance scores. The stated goal is to help practitioners choose among U-Net variants for change detection, with conclusions such as 'NDR-U-Net achieves the highest Mean IoU (0.92)' and that Optimised U-Net and NDR-U-Net are 'top contenders.' The paper also discusses limitations of standard U-Net and qualitatively maps each variant to mitigation strategies for remote sensing change detection challenges.","tokens_in":10145,"tokens_out":2027,"duration_ms":24648,"significance":"If the comparative assessment were valid, it would serve as a useful entry point for practitioners selecting U-Net backbones for change detection. The paper compiles a relevant set of recent architectures and identifies plausible design themes (attention mechanisms, transfer learning, skip-connection modifications, transformer blocks). It also makes the worthwhile observation that standard U-Net's limited receptive field and coarse skip connections are bottlenecks for remote sensing change detection. However, the paper's central contribution is a ranking of variants based on incomparable published scores, and the lack of any controlled evaluation or explicit inclusion criteria undermines this ranking. As a review, the paper would require a systematic methodology and a transparent statement of what can and cannot be concluded from heterogeneous literature values, neither of which is present.","major_comments":[{"comment":"Table 4 is the sole empirical basis for the comparative claim, but it juxtaposes performance scores measured on different datasets, sensors, resolutions, and metrics. For example, Siamese Swin-U-Net reports F1=94.67 on CDD, STCD-EffV2T U-Net reports Mean IoU=0.87 on OSCD/Sentinel-2, and NDR-U-Net reports Mean IoU=0.92 on GF-2. These numbers do not control for dataset difficulty, class balance, evaluation protocol, or even the same metric. Without a common benchmark or error bars, the statement that 'NDR-U-Net achieves the highest Mean IoU (0.92)' and the designation of 'top contenders' are not justified. This incomparability is load-bearing: if it is not addressed, the paper's central conclusion collapses into a list of reported values.","section":"Section 4.1, Table 4"},{"comment":"The abstract claims 'a comparison and analysis of 18 different U-Net variations,' but Section 5 states 'we study 8 variants,' and Tables 2, 3, and 4 list only 8 architectures. The discrepancy is never explained, and no inclusion or exclusion criteria are given for the 34 papers surveyed. Consequently, the reader cannot assess whether the selected variants and papers are representative, which is a second load-bearing weakness for a review whose stated purpose is to provide a comparative assessment.","section":"Abstract and Section 5"},{"comment":"Several architecture descriptions are inconsistent with the cited sources or with the paper's own table. For instance, Table 2 lists T-U-Net as 'Triple branch architecture,' but the text describes it as incorporating a transformer module in the decoder; both T-U-Net and Optimised U-Net are cited to the same reference [29], which appears to be the Optimised U-Net paper only. Additionally, the description of NDR-U-Net is hedged as 'it likely introduces modifications in the encoder path,' which is speculative rather than a verified account of the cited method. These inconsistencies impair the accuracy of the qualitative comparison that the paper does provide.","section":"Section 3.2, Table 2"}],"minor_comments":[{"comment":"The abstract contains grammatical errors and typos, such as 'this paper fill the gap' and 'ever changing' (missing hyphen), which should be corrected.","section":"Abstract"},{"comment":"Figure 1 is referenced as showing a 'substantial increase' in U-Net variant publications, but the figure appears to be a line chart without axes labels or a data source; the claim is not substantiated by the displayed content.","section":"Figure 1"},{"comment":"The description of Stack U-Net (reference [13]) is vague ('Stack U-Net incorporates multi-resolution feature extraction'), and the reference list entry lacks full publication details, making it difficult to locate the original work.","section":"Section 3.1"},{"comment":"The sentence 'Optimized U-Net shows promising overall accuracy (92.10%) for land cover classification' is presented without context on the dataset or the number of classes, which would be needed for any meaningful interpretation.","section":"Section 4.1"},{"comment":"The discussion of techniques such as min-max normalization and data augmentation is not tied to any specific variant or experimental result, making these paragraphs loosely connected to the paper's comparative theme.","section":"Section 4.2"}],"recommendation":"reject","confidential_remarks":"The paper is a narrative review with a strong claim of comparison that it cannot support. The central flaw—ranking incomparable published scores—is unlikely to be fixable without new experiments on a shared benchmark, which is beyond the scope of the current manuscript. There are also potential citation mismatches (e.g., [29] used for two different architectures) that would require careful verification even in a revised version. The topic is timely, but the execution is not at the standard expected for a journal publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this paper is a survey that promises 18 U-Net variants and a comparative assessment, but it details only 8 and ranks them with numbers pulled from different datasets, sensors, resolutions, and metrics. The central comparison table can't support the conclusions drawn from it.\n\nWhat the paper does well: it offers a reasonable taxonomy of recent U-Net modifications for remote sensing change detection—attention-based, transfer learning, skip-connection changes, encoder changes. The discussion of standard U-Net limitations is mostly accurate, and the bibliography is a useful starting point for someone new to the area. The effort to organize a scattered literature is genuine.\n\nWhere it falls apart: the abstract promises a comprehensive analysis of 34 papers and a comparison of 18 U-Net variations, but Section 5 says \"we study 8 variants\" and Table 2 lists 8. No inclusion/exclusion criteria are given for the 34 papers. Table 4 is the core evidence, and it mixes F1-score (94.67), Mean IoU (0.87, 0.845, 0.92, 0.88), Overall Accuracy (89.20%, 92.10%), and one qualitative \"Higher accuracy\" for T-U-Net. These come from CDD, OSCD, LEVIR-CD/WHU-CD/DSIIN-CD, HRS/WorldView-2, Vaihingen/LiDAR, and Gaofen-2/GF-2—different sensors, ground resolutions, class balances, and evaluation protocols. No error bars, no normalized benchmark. From that table the text concludes that NDR-U-Net achieves the highest Mean IoU and that Optimized U-Net and NDR-U-Net are top contenders. That conclusion simply does not follow. A higher number from one study is not evidence of a better model when nothing is controlled.\n\nThis is not a minor weakness; the comparison is the paper's stated purpose. If the rank ordering is removed, the paper becomes a simple categorized list of variants with commentary, which is fine but much more modest. The categorization is salvageable; the quantitative comparison is not.\n\nWho this is for: a practitioner skimming for a list of recent U-Net variants and their reported scores might get a few pointers, but they would have to go back to the original papers to learn anything trustworthy. As a survey it does not meet the bar for a reliable comparative study.\n\nMy recommendation: desk reject in current form. The authors would need to add a selection methodology, either run their own controlled experiments or explicitly state that no cross-paper comparison is possible, correct the abstract/discussion mismatch, and remove the unsupported rankings. If they revise it into a properly scoped taxonomy without pseudo-quantitative rankings, it could be a useful contribution for practitioners.","headline":"A survey that promises more than it delivers: the categorization is useful, but the quantitative comparison mixes incomparable numbers and the abstract's 18-variant claim is never backed up.","tokens_in":10666,"tokens_out":2111,"would_cite":false,"duration_ms":25570,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper compiles reported results for 18 U-Net variants and argues that design choices such as Siamese branches, transformers, and attention gates address specific U-Net limitations in satellite change detection.","keywords":["remote sensing","change detection","U-Net","satellite imagery","image segmentation","attention mechanism","Swin transformer","deep learning"],"falsifier":"Run NDR-U-Net, Siamese Swin-U-Net, and Optimised U-Net on one shared benchmark, such as LEVIR-CD, with identical preprocessing, training protocol, and metric computation; the paper's ranking stands only if their relative order under these controlled scores matches Table 4.","tokens_in":9566,"feed_emoji":"🛰️","tokens_out":5517,"duration_ms":58855,"temperature":0.7,"pith_summary":"This paper surveys the growing family of U-Net-based architectures for detecting landscape changes in satellite imagery. It catalogs 18 variants, groups their modifications into attention mechanisms, transformer blocks, transfer learning, and skip-connection refinements, and assembles reported performance scores from 34 papers. The aim is to tell practitioners which design choices matter for remote sensing change detection, where images from different times and sensors must be compared. The headline result is a ranking in which NDR-U-Net, Siamese Swin-U-Net, and Optimised U-Net lead on different metrics. A sympathetic reader would take the paper as a structured map of the design space, not as a controlled benchmark.","feed_headline":"NDR-U-Net tops comparison of 18 U-Net variants for change detection","feed_subtitle":"A review of 34 papers maps which design changes—transformers, attention, skip connections—support satellite change detection.","key_machinery":"The central object is the U-Net variant taxonomy, organized around five modification strategies: attention gates, transformer blocks, transfer-learning encoders, encoder or skip-connection redesigns, and hierarchical nested blocks. The mechanism that carries the argument is the mapping from U-Net components to known failure modes: the encoder to receptive field, the decoder to global context, and skip connections to feature fusion. Each variant is read as an intervention on one component, and reported metrics (F1, Mean IoU, Overall Accuracy) are used as evidence for which interventions help.","core_discovery":"This paper claims that the standard U-Net, inherited from medical segmentation, stumbles on satellite change detection for three reasons: limited receptive field, weak global context, and skip connections that merge semantically distant features. It then shows how 18 variants target these weaknesses: Swin transformers and attention modules capture long-range dependencies, pre-trained encoders like EfficientNetV2 bring transfer learning, bottleneck and bilateral attention refine skip connections, and nested dense residual blocks add multi-scale features. Based on assembled scores, the paper identifies NDR-U-Net (Mean IoU 0.92) as the leading segmentation result and Siamese Swin-U-Net (F1 94.67) as the leading change detection classifier. The intended contribution is a practical map: choose a variant by matching its modification to the dominant failure mode of the task.","pith_inferences":["Beyond the paper: a controlled re-run of the top variants on one benchmark would be needed before treating the Table 4 ranking as a leaderboard, since the paper compiles scores from different studies rather than running a single experiment.","Beyond the paper: the taxonomy suggests a natural next test, combining a Swin-transformer encoder with an attention-gated skip connection, because those two modifications target the receptive-field and skip-connection limitations separately.","Beyond the paper: because the datasets span optical, LiDAR, and different resolutions, the paper's data hint that sensor-specific adapter modules may matter more than any single architecture choice, though the paper does not test this."],"forward_implications":["If the assembled scores are taken at face value, NDR-U-Net is the variant to try first for segmentation-heavy change detection, since it reports the highest Mean IoU (0.92).","Siamese Swin-U-Net's F1 of 94.67 makes it the leading candidate when the task is binary land-cover change classification.","The paper's mapping from U-Net components to failure modes implies that skip-connection fixes (bottlenecks, bilateral attention) and long-range modules (transformers) address different problems, so neither alone covers all cases.","Practitioners with small datasets should favor transfer-learning variants such as STCD-EffV2T U-Net, since pre-trained encoders are the paper's stated remedy for limited training data."],"supporting_citations":[{"why":"Supplies the baseline U-Net architecture from which all variants depart.","marker":"[8]"},{"why":"Supplies the Siamese Swin-U-Net variant and its reported F1 score of 94.67.","marker":"[26]"},{"why":"Supplies the NDR-U-Net variant and its reported Mean IoU of 0.92.","marker":"[32]"},{"why":"Supplies the Bilateral Attention U-Net and its reported Mean IoU of 0.845.","marker":"[27]"},{"why":"Supplies the Optimised U-Net and its reported overall accuracy of 92.10%.","marker":"[29]"},{"why":"Supplies the STCD-EffV2T U-Net and its transfer-learning backbone with reported Mean IoU of 0.87.","marker":"[28]"},{"why":"Supplies the HARNU-Net variant and its hierarchical attention mechanism with reported Mean IoU of 0.88.","marker":"[10]"}],"fun_headline_variants":["NDR-U-Net wins U-Net shootout for satellite change detection","Study ranks 18 U-Net tweaks: NDR-U-Net leads change detection","Which U-Net variant detects satellite changes best? NDR-U-Net","34-paper review crowns NDR-U-Net for satellite change detection","NDR-U-Net beats 17 U-Net variants in change detection test"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's ranking assumes that F1, Mean IoU, and Overall Accuracy scores taken from different publications, measured on different datasets and sensors, can be compared directly as if from a single benchmark.","fun_headline_variants_meta":{"raw":{"variants":["NDR-U-Net wins U-Net shootout for satellite change detection","Study ranks 18 U-Net tweaks: NDR-U-Net leads change detection","Which U-Net variant detects satellite changes best? NDR-U-Net","34-paper review crowns NDR-U-Net for satellite change detection","NDR-U-Net beats 17 U-Net variants in change detection test"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000303,"raw_usage":{"total_tokens":1710,"prompt_tokens":882,"completion_tokens":828,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":498,"completion_tokens_details":{"reasoning_tokens":724}},"tokens_in":498,"tokens_out":828,"duration_ms":8161,"temperature":1.0,"reasoning_tokens":724,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:21:12.188888+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run NDR-U-Net, Siamese Swin-U-Net, and Optimised U-Net on one shared benchmark, such as LEVIR-CD, with identical preprocessing, training protocol, and metric computation; the paper's ranking stands only if their relative order under these controlled scores matches Table 4.","supporting_citations":[{"cited_title":"SA-UNet: Spatial Attention U-Net for Retinal Vessel Segmentation,","cited_arxiv_id":null,"evidence_quote":"Supplies the NDR-U-Net variant and its reported Mean IoU of 0.92."},{"cited_title":"It allows the network to focus on informative features","cited_arxiv_id":null,"evidence_quote":"Supplies the Bilateral Attention U-Net and its reported Mean IoU of 0.845."},{"cited_title":"Concurrent Spatial and Channel Squeeze & Excitation in Fully Convolutional Networks,","cited_arxiv_id":null,"evidence_quote":"Supplies the Optimised U-Net and its reported overall accuracy of 92.10%."},{"cited_title":"USE-Net: Incorporating Squeeze- and-Excitation Blocks into U-Net for Prostate Zonal Segmentation of Multi-Institutional MRI Datasets,","cited_arxiv_id":null,"evidence_quote":"Supplies the STCD-EffV2T U-Net and its transfer-learning backbone with reported Mean IoU of 0.87."},{"cited_title":"Change Detection Based on Deep Siamese Convolutional Network for Optical Aerial Images,","cited_arxiv_id":null,"evidence_quote":"Supplies the HARNU-Net variant and its hierarchical attention mechanism with reported Mean IoU of 0.88."}],"review_version":1}