{"id":"605a3527-db9e-4b44-b9cd-cc0289dff7f6","arxiv_id":"2411.17088","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"ΩSFormer, a dual-branch Transformer that fuses satellite imagery and DEM, is claimed to reach 0.976 mIoU on a new terraced-field dataset, but the cross-scale setting in the paper's own table is 0.883 mIoU.","lead":"The paper introduces a new dual-modal dataset and a Transformer network for detecting terraced fields from satellite images and elevation data, reporting very high segmentation accuracy. The main catch: the headline accuracy is for same-resolution data, while the paper's own cross-scale experiment is much worse, contradicting the abstract's claim.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Abstract reports 0.976 mIoU for cross-scale dual-modal data, but Table 4 shows the cross-scale (2m/12.5m/2m) result is 0.883; the headline number belongs to same-scale 2m/2m/2m, directly contradicting the central cross-scale claim.","rationale":"The reader's verdict is REJECT with high correctness risk. My independent read identifies a more direct load-bearing flaw than the one highlighted as the weakest assumption. The reader's weakest_assumption focuses on label leakage from annotating over a 3D DEM model; that is a plausible and serious concern, but it requires inferring how much information the DEM branch can extract from the label construction process. The Table 4 contradiction is stronger because it is visible within the paper itself: the abstract's headline 0.976 mIoU is explicitly attributed to cross-scale dual-modal data, while Section 4.3 reports the cross-scale result as 0.883 mIoU. This is not a subtle statistical issue or a matter of alternative interpretation; the same metric under the stated condition differs by 0.093, and the higher value is assigned to a different experimental condition. Since the central novelty is cross-scale super-resolution fusion, the absence of a valid cross-scale result means the main contribution is unsupported as reported. The reader's rationale also noticed this discrepancy, but the formal weakest_assumption field points to label leakage, so my agreement is partial rather than full. My additional concern about label leakage does not change the verdict; it reinforces it. The recommended action remains rejection, or at minimum major revision with corrected reporting and independent validation. Since the reader already rejected the paper and my concern points in the same direction, the verdict should remain unchanged.","tokens_in":22226,"tokens_out":2718,"duration_ms":25136,"concrete_test":"Using the released DMRVD and ΩSFormer code, rerun the three configurations from Table 4 with identical training hyperparameters and data splits, reporting mIoU on the same test set with at least three seeds. Then check whether the 0.976 value reproduces under the cross-scale 2m/12.5m/2m configuration or only under the same-scale 2m/2m/2m configuration. If 0.976 appears only under same-scale, the abstract's cross-scale claim is contradicted and the manuscript must be corrected; if 0.976 reproduces under cross-scale, the discrepancy with Table 4 must be explained.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that ΩSFormer achieves cross-scale super-resolution fusion, with the abstract stating 'For cross-scale dual-modal datasets, ΩSFormer achieved the best performance with mIOU and OA values of 0.976 and 0.981 respectively.' Section 4.3 and Table 4 define the cross-scale setting as RGB 2m, DEM 12.5m, label 2m and report mIoU 0.883 for that configuration. The 0.976 mIoU appears in Table 4 only for the same-scale high-resolution setting (2m/2m/2m). Thus the headline result in the abstract is not the cross-scale result; it is the same-scale result. This is an internal inconsistency, not a matter of disagreement with external consensus. The claimed improvements over single-modal and dual-modal baselines (0.165, 0.297, 0.128) are likewise tied to the 0.976 mIoU in Table 2, so unless Table 2 was run under the cross-scale condition, those improvements do not demonstrate cross-scale capability either. The paper's own cross-scale experiment shows a substantial drop (0.976 → 0.883), which undercuts the claim that the Ω-like architecture 'fully integrates' cross-scale features to achieve high accuracy. The label-leakage concern raised by the reader is also serious and may inflate DEM gains, but the Table 4 discrepancy is more immediately load-bearing because it can be established from the manuscript alone, without needing additional assumptions about how labels were produced.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes ΩSFormer, a dual-modal Transformer network for terraced-field vectorization extraction (TFVE) from high-resolution RGB imagery and DEM data, together with a new dual-modal raster-vector dataset (DMRVD) covering nine counties in China. The authors claim that ΩSFormer is the first model to address cross-modal and cross-scale super-resolution fusion for TFVE, and report very high accuracy: mIoU 0.976 and OA 0.981, with improvements of 0.165, 0.297, and 0.128 over the best single-modal imagery, single-modal DEM, and dual-modal baselines, respectively. The manuscript includes comparisons with U-Net, U-Net++, DeepLabV3+, HRNet, HRFormer, SegFormer, W-Net, W-Net++, and SAM, an ablation of STSRO, and cross-scale experiments with three resolution combinations.","tokens_in":22546,"tokens_out":4442,"duration_ms":37268,"significance":"If the central claims were correct, the paper would contribute a useful dataset and architecture for terraced-field extraction, an application relevant to soil and water conservation. The release of DMRVD and code is a strength, and the comparison against multiple baselines is a positive feature. However, the paper's headline claim of cross-scale dual-modal performance is internally contradicted by its own Table 4, and the label-production procedure entangles the DEM input with the ground truth. These issues undermine the main contribution as stated, so the current manuscript cannot be accepted as a reliable demonstration of cross-scale super-resolution TFVE.","major_comments":[{"comment":"The abstract states that 'For cross-scale dual-modal datasets, ΩSFormer achieved the best performance with mIOU and OA values of 0.976 and 0.981 respectively,' but Table 4 defines the cross-scale configuration as RGB 2m, DEM 12.5m, and label 2m, and reports mIoU 0.883 for that row. The 0.976 value belongs to the same-scale 2m/2m/2m configuration. This is an internal contradiction in the paper's central claim about cross-scale performance.","section":"Abstract vs. Table 4"},{"comment":"The paper's own cross-scale experiment produces the worst accuracy among the three configurations: 0.883 mIoU, below both same-scale 2m (0.976) and same-scale 12.5m (0.921), and the text explicitly says that 'the semantic segmentation results of cross-scale are the worst.' The improvements of 0.165, 0.297, and 0.128 reported in the abstract are computed from Table 2's 0.976 result, which is the same-scale high-resolution result, not the cross-scale result. No baseline comparisons under the cross-scale condition are provided, so the claim that ΩSFormer is best under cross-scale input is unsupported.","section":"Section 4.3 and Table 4"},{"comment":"The ground-truth labels were produced by overlaying high-resolution imagery on a 3D model and visually interpreting boundaries, where the 3D model is constructed from the DEM that is later used as the second input modality. Because the DEM is entangled with label construction, the reported dual-modal gains over single-modal imagery may reflect label leakage rather than learned terrain understanding. The manuscript does not analyze or control for this dependency.","section":"Section 2.2, label production"},{"comment":"The quality of DMRVD is validated by the accuracy that the authors' own model achieves on it ('The quality of DMRVD can be evaluated based on the degree of accuracy achieved in the classification process'), which is circular. An independent assessment of label quality—for example, inter-annotator agreement or comparison with field survey data—is needed to substantiate the dataset contribution.","section":"Section 4.4.1"},{"comment":"All accuracy values are point estimates from a single training run with no error bars, repeated trials, or statistical significance tests. Consequently, the reported improvements (e.g., 0.128 over the best dual-modal baseline, 0.035 from STSRO in Table 3) cannot be distinguished from run-to-run variation, and the superiority claims are not statistically grounded.","section":"Tables 2–4"}],"minor_comments":[{"comment":"SegFormer is cited as (Yuan et al. 2021a) in Section 3 and in the comparison network list, but the Introduction correctly attributes SegFormer to Xie et al. (2021); the citation key should be corrected to avoid ambiguity with HRFormer.","section":"References and Table 2"},{"comment":"Equation (10) and the surrounding text contain garbled subscripts and superscripts (e.g., the indices in the discrete contour vibration equation) that make the formula difficult to read; the notation should be cleaned up.","section":"Section 3.5, Eq. (10)"},{"comment":"The comparison with SAM lacks any description of how SAM was prompted or adapted for binary terraced-field segmentation, making the comparison impossible to reproduce.","section":"Table 2, SAM row"},{"comment":"The metric is spelled inconsistently as both 'mIOU' and 'mIoU'; please standardize the notation throughout the manuscript.","section":"General notation"}],"recommendation":"reject","confidential_remarks":"The internal contradiction between the abstract's cross-scale claim and Table 4 is decisive: the headline result is a same-scale result, and the paper's own cross-scale experiment is the worst configuration. Combined with the label-entanglement issue in Section 2.2, I do not see a revision path within the scope of this manuscript. A future resubmission would need to either reframe the contribution as same-scale dual-modal fusion with independent label-quality assessment, or provide entirely new cross-scale experiments and baseline comparisons."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, you should know two things about this one. The DMRVD dataset is genuinely new and open: 19,330 image-DEM pairs, 371k vector polygons, nine counties across China. That is a real resource for terraced field mapping and worth having. The ΩSFormer network, on the other hand, is an engineering stack of known modules—HRFormer-style high-res transformers, OCNet-like object context, CVNet for contour vectorization—and the paper's central claim about cross-scale fusion isn't supported by its own numbers.\n\nHere's the problem. The abstract says cross-scale dual-modal datasets achieved mIoU 0.976 and OA 0.981, and that this beats single-modal and dual-modal baselines by 0.165, 0.297, 0.128. Table 4 tells a different story: the cross-scale experiment (2 m RGB, 12.5 m DEM, 2 m label) gives 0.883 mIoU. The 0.976 belongs to the same-scale 2m/2m/2m condition. So the headline result is not the cross-scale result; the cross-scale result is substantially worse. The claimed improvements are tied to the same-scale number, so they don't evidence cross-scale capability either. That's an internal contradiction in the paper's main selling point, not an interpretation difference.\n\nThere's also a label-leakage worry. Labels were made by annotating over a 3D model that displays the DEM (Section 2.2), so DEM information is baked into the ground truth. That makes it hard to interpret the apparent gain from adding DEM. Add no error bars, a circular dataset-quality argument (Section 4.4.1 validates the dataset by the authors' own model's accuracy), and no quantitative vectorization metrics—the vectorization module is shown only qualitatively.\n\nThe net is: the dataset is worth something, the network is a plausible engineering contribution, but the paper as written overclaims. The good news is that the problems are correctable in revision: rewrite the abstract to match Table 4, run the cross-scale comparisons with an unbiased label protocol or at least disclose the entanglement, add uncertainty and vector boundary metrics, and reframe the claim from 'cross-scale super-resolution works' to 'same-scale works and cross-scale is a harder open problem.'\n\nI'd send it to peer review—the dataset justifies referee time—but with the expectation of major revision. The reader's REJECT is fair for the current version; I'd soften it to 'major revision required' because the underlying resource is real and the central conflict can be fixed without new collection.","headline":"Useful new dataset, but the paper's central cross-scale claim contradicts its own Table 4; fix that before believing the 0.976 number.","tokens_in":80,"tokens_out":3443,"would_cite":true,"duration_ms":68546,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"ΩSFormer fuses high-resolution imagery with low-resolution DEM to extract terraced-field vectors at 0.976 mIoU.","keywords":["terraced field extraction","vectorization","dual-modal fusion","super-resolution transformer","semantic segmentation","digital elevation model","soil and water conservation"],"falsifier":"Retrain on labels created without any DEM or 3D terrain model, then compare dual-modal and RGB-only mIoU; if the 0.165-point margin shrinks to near zero, the claimed terrain-modal contribution is mostly label leakage rather than learned generalization.","tokens_in":21963,"feed_emoji":"🌾","tokens_out":8626,"duration_ms":71807,"temperature":0.7,"pith_summary":"The paper claims that terraced fields—stepped farmlands built for soil and water conservation—can be extracted from remote sensing data as ready-to-use vector polygons with near-perfect agreement, reaching an mIoU of 0.976 and an overall accuracy of 0.981 on its same-scale 2 m dual-modal test set. To do this, it proposes ΩSFormer, a dual-branch Transformer whose Ω-like shape keeps high-resolution features alive through the encoder and fuses spectral imagery with digital elevation data at multiple scales. The paper also introduces DMRVD, a 22 441 km² dataset of 19 330 RGB images, 19 330 DEMs, and 371 150 hand-vectorized terraced-field polygons across nine counties in four Chinese provinces, which it says is the first dataset built specifically for deep-learning-based terraced-field vectorization extraction. If the reported results hold, automated monitoring of soil erosion, carbon sequestration, and agricultural land use would no longer need to rely on slow manual delineation.","feed_headline":"New dual-modal AI maps terraced fields at 97.6% mIoU","feed_subtitle":"Fusing RGB images with low-res elevation data, ΩSFormer outputs clean vector polygons for soil-conservation monitoring.","key_machinery":"The carrying mechanism is the Ω-like network shape formed by the super-resolution Transformer connection module (SRTCM): at each encoder stage the original high-resolution features are fused with progressively downsampled features, making the network wide at the top and narrow at the bottom, with the two modality branches arranged symmetrically in a dual-modal, dual-branch (D2MB) input scheme. Multi-head self-attention on non-overlapping windows preserves local edge detail while 3×3 depthwise convolutions enable cross-window exchange. The spatial topological semantic relationship optimization (STSRO) then augments each pixel's representation by comparing it with whole object regions, and the vectorization extraction module (VEM) reparameterizes a contour vibration neural network to iteratively evolve segmentation boundaries into smooth vector polygons.","core_discovery":"On the paper's own terms, the central discovery is that a dual-modal, dual-branch super-resolution Transformer can make terraced-field extraction nearly indistinguishable from human annotation. The network fuses high-resolution RGB imagery with lower-resolution DEM data by continuously re-injecting original high-resolution features into each downsampling stage, so that fine edge information is never lost and the DEM is effectively super-resolved by the imagery. A spatial-topological refinement stage then sharpens boundary pixels using object-region relationships, and a contour-vibration network iteratively deforms the segmentation boundary into smooth vector outlines. Against U-Net, U-Net++, DeepLabV3+, HRNet, W-Net, W-Net++, HRFormer, SegFormer, and SAM, the paper reports the best mIoU of 0.976 and OA of 0.981, with mIoU improvements of 0.165, 0.297, and 0.128 over the best single-modal imagery, single-modal DEM, and dual-modal baselines. The abstract presents these as cross-scale dual-modal results; the paper's Table 4 reports the same 0.976/0.981 values for the same-scale 2 m dual-modal configuration.","pith_inferences":["Editorial inference: since the labels were created by visual interpretation of imagery overlaid on a 3D terrain model built from the DEM, the DEM's measured contribution may include label leakage; an unbiased test needs labels produced independently of elevation data.","Editorial inference: the reported cross-scale advantage should be read with care, because the same paper's Table 4 shows the 2 m/12.5 m cross-scale configuration at mIoU 0.883, below the same-scale 2 m result of 0.976; a genuine super-resolution fusion mechanism would be expected to narrow that gap.","Editorial inference: a natural extension is to test ΩSFormer on terraced-field types and geographies outside the nine Chinese counties, for example Mediterranean or Andean terraces, to see whether the near-perfect accuracy reflects dataset-specific regularities or general terrain understanding."],"forward_implications":["If correct, terraced-field mapping for soil-erosion monitoring can be automated across large, heterogeneous regions at near-human accuracy.","The vector output eliminates the pixelation and storage bloat of raster masks, giving infinitely scalable, directly analyzable field boundaries.","High-resolution imagery can act as a super-resolution prior for coarse DEMs, recovering terrain detail that the DEM alone cannot provide.","Separate dual-branch encoding is a better fusion strategy than stacking modalities into extra channels, because each branch can correct its own alignment errors before fusion.","DMRVD gives the research community a reusable open benchmark for terraced-field vectorization, with 371 150 labeled polygons across diverse terrain."],"supporting_citations":[{"why":"Supplies the classic U-Net baseline whose 0.718 single-modal mIoU defines the lower bound the proposal must beat.","marker":"Ronneberger et al. 2015"},{"why":"Supplies DeepLabV3+, the backbone that prior terraced-field extraction works and one of the strongest baselines in the comparison table.","marker":"Chen et al. 2018a"},{"why":"Provides HRFormer, the high-resolution Transformer whose windowed self-attention and stage design ΩSFormer adapts; Fig. 4 is modified from it.","marker":"Yuan et al. 2021a"},{"why":"Supplies the contour vibration neural network that VEM reparameterizes to iterate terraced-field boundaries into vectors.","marker":"Xu et al. 2022"},{"why":"Supplies the Lovász-softmax loss that smooths segmentation boundaries during training.","marker":"Berman et al. 2018"},{"why":"Supplies the initialization strategy for the active contour used by VEM.","marker":"Cheng et al. 2019"},{"why":"Motivates the cross-scale internal-recursion super-resolution idea used to explain DEM enhancement by imagery.","marker":"Shocher et al. 2018"},{"why":"Supplies SAM as a state-of-the-art baseline that ΩSFormer must outperform (0.794 RGB mIoU).","marker":"Kirillov et al. 2023"}],"fun_headline_variants":["ΩSFormer fuses RGB and DEM to vectorize terraced fields at 97.6% mIoU","Dual-modal transformer super-resolves DEM to 97.6% mIoU on terraced fields","Cross-scale dual-modal AI maps terraced fields with 97.6% mIoU","ΩSFormer super-resolves DEM for 97.6% mIoU terraced field extraction"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the DEM is an independent input, when in fact the ground-truth labels were created by looking at imagery on a 3D terrain surface derived from that same DEM, so the DEM's value may be baked into the labels rather than learned by the network.","fun_headline_variants_meta":{"raw":{"variants":["ΩSFormer fuses RGB and DEM to vectorize terraced fields at 97.6% mIoU","Dual-modal transformer super-resolves DEM to 97.6% mIoU on terraced fields","Cross-scale dual-modal AI maps terraced fields with 97.6% mIoU","ΩSFormer super-resolves DEM for 97.6% mIoU terraced field extraction"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000732,"raw_usage":{"total_tokens":3366,"prompt_tokens":1126,"completion_tokens":2240,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":742,"completion_tokens_details":{"reasoning_tokens":2135}},"tokens_in":742,"tokens_out":2240,"duration_ms":15299,"temperature":1.0,"reasoning_tokens":2135,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T12:34:40.899407+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Retrain on labels created without any DEM or 3D terrain model, then compare dual-modal and RGB-only mIoU; if the 0.165-point margin shrinks to near zero, the claimed terrain-modal contribution is mostly label leakage rather than learned generalization.","supporting_citations":[],"review_version":1}