{"id":"6e020211-b34a-4cdc-abb7-0b95962287b0","arxiv_id":"1909.01868","paper_version":3,"verdict":"REJECT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"New-application report: CNN and ConvLSTM networks trained on StaMPS labels select persistent scatterer pixels in Sentinel-1 interferograms, with reported validation accuracy of 93.50% for the LSTM variant.","lead":"The authors train convolutional and convolutional-LSTM networks to pick 'persistent scatterer' pixels from stacks of satellite radar images, the pixels used to measure ground deformation. They report that the LSTM version is faster and selects denser scatterer maps than the standard StaMPS algorithm, but the comparison uses training labels from StaMPS itself.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central 'improved classification' claim is evaluated only against StaMPS-derived training labels plus the authors' own STIP proxy, so higher accuracy and density do not establish that CLSTM-ISS selects more true PS pixels than StaMPS.","rationale":"The reader's REJECT verdict is appropriate. My stress-test pass found the same load-bearing weakness: the paper never validates the central 'improved classification' claim against an independent PS ground truth. The network is trained on StaMPS labels, so validation accuracy is a measure of faithfulness to StaMPS, not of selection quality. Section 4 admits that true labels are lacking and therefore switches to R-index and land-cover interpretation; Section 5 adds STIP from the authors' own prior work. All of these are heuristics for where PS pixels 'should' be, not measurements of actual phase stability or deformation accuracy. Because CLSTM-ISS is more permissive in the test scene, selecting roughly five times as many PS pixels as StaMPS, its higher density and STIP counts could be an artifact of a lower effective threshold rather than of better classification. The velocity and time-series comparisons in Section 5 are visual and qualitative, and they rely on the same selected pixels, so they cannot break the circularity. The paper's contribution—a fast learned surrogate for StaMPS—may have practical value, and if the authors released code, weights, and an independent benchmark, the verdict could change. But as written, the strongest claim is unsupported.","tokens_in":19512,"tokens_out":4037,"duration_ms":41579,"concrete_test":"Use known stable scatterers in the Kathmandu scene as independent ground truth: for example, corner reflectors deployed around the 2015 Nepal earthquake or GPS/levelling benchmarks visible in Sentinel-1. Compute precision, recall, and F1 for the PS masks produced by StaMPS, CNN-ISS, and CLSTM-ISS against these known points within a small spatial tolerance. If CLSTM-ISS does not show higher F1 than StaMPS on the independent targets, Sections 4 and 5's R-index and STIP arguments do not support the central claim. If no field ground truth is available, an equivalent test is a synthetic interferometric stack with known phase-stable pixels and realistic noise, evaluating ROC curves for all three methods.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The decisive weakness is the evaluation of the central claim. Training labels are generated by StaMPS (Section 3), so the 93.50% validation accuracy reported in Table 5 measures the network's agreement with StaMPS labels, not correctness against true PS phase stability; it cannot support 'improved classification' over StaMPS. The paper itself acknowledges the lack of true labels in Section 4: 'with the lack of true labels, using these metrics for quality evaluation would provide an incorrect estimate of the classifier performance.' The subsequent evaluation then substitutes (i) a heuristic combination of R-index and a land-cover image to infer where PS pixels 'should' occur, and (ii) Section 5's STIP measure taken from Narayan et al. 2018a/b by the same authors. The interpretation that 'more PS in city/lengthening and fewer in forest/river' is better assumes these proxies are valid PS ground truth. CLSTM-ISS also selects roughly five times as many PS pixels as StaMPS (192,177 vs 38,286), so the higher 'reliable PS density' could be caused by an over-permissive threshold or false positives if the proxy is wrong. The velocity and time-series comparisons in Section 5 are qualitative visual pattern matching and do not resolve this. Thus the claim that CLSTM-ISS improved PS selection over StaMPS is load-bearing and unsupported by an independent benchmark.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes two deep-learning architectures, CNN-ISS and CLSTM-ISS, for pixel-wise classification of persistent scatterer (PS) and non-PS pixels in multi-temporal InSAR interferograms. The networks are trained on roughly 10,000 100-by-100 interferometric image patches from three Sentinel-1 study sites, using StaMPS-derived PS labels as ground truth, and tested on an unseen Kathmandu dataset. The authors report that CLSTM-ISS achieves 93.50% validation accuracy versus 89.21% for CNN-ISS, produces a higher PS density than StaMPS, and runs in minutes compared with 108 minutes for StaMPS. The claimed superiority over StaMPS is based on a qualitative comparison using R-index and land-cover maps, and on a reliability analysis using the STIP metric introduced in the authors' prior work.","tokens_in":19797,"tokens_out":5513,"duration_ms":59278,"significance":"If the central claim were established, the contribution would be practically valuable: a trained network that selects PS pixels in near real time could remove a major computational bottleneck in MT-InSAR processing and support time-critical deformation monitoring. The paper also addresses a real problem, uses real-world multi-site training data, and explicitly acknowledges the absence of true labels. However, the evaluation does not support the central claim of improved classification over StaMPS: the headline validation accuracies measure agreement with the same StaMPS-generated labels used for training, while the external checks are either heuristic (R-index plus land-cover) or based on the authors' own STIP metric with an unvalidated threshold. The reported gains in PS density and STIP counts could plausibly be explained by an over-permissive classifier rather than by better phase-stability selection. No code, trained weights, or repeated experimental runs are provided, limiting reproducibility.","major_comments":[{"comment":"The reported validation accuracy of 93.50% for CLSTM-ISS measures agreement with StaMPS-generated training labels, not correctness of PS selection. Since the same labels define the training target, high validation accuracy cannot support the claim that CLSTM-ISS improves classification over StaMPS. The paper itself acknowledges this in Section 4: 'with the lack of true labels, using these metrics for quality evaluation would provide an incorrect estimate of the classifier performance.' An independent ground truth or a strictly separate reference method is required before any claim of improvement can be made.","section":"Section 4, Table 5"},{"comment":"The qualitative evaluation uses a combination of the R-index and a classified land-cover image as the reference for where PS pixels 'should' occur. This is a heuristic proxy, not an independent validation of phase stability. The assumption that more PS pixels in man-made and lengthening areas and fewer in forest and river areas necessarily indicates better classification is not quantitatively justified, and no uncertainty or sensitivity analysis is given. Consequently, statements such as 'CLSTM-ISS outperformed the other two methods' (Section 4, after Table 4) are not supported by the evidence presented.","section":"Section 4, Figures 6-8 and Table 4"},{"comment":"The STIP reliability check relies on a metric introduced by the same authors in Narayan et al. (2018a, 2018b), and the STIP>35 threshold is asserted without independent validation for the Kathmandu data. Given that CLSTM-ISS selects roughly five times as many PS pixels as StaMPS (192,177 versus 38,286 in Table 5), the larger absolute number of STIP>35 pixels (186,435 versus 35,413) could simply reflect a more permissive selection threshold. Without an analysis of false positives against known non-PS targets, the higher STIP count does not establish that CLSTM-ISS selects more true PS pixels.","section":"Section 5, STIP analysis"},{"comment":"The velocity maps and time-series displacement comparisons are evaluated qualitatively by visual pattern matching. The similarity of CLSTM-ISS to StaMPS is not a meaningful benchmark because StaMPS is the source of the training labels. No quantitative metric (e.g., RMS difference against independent deformation measurements), no error bars, and no repeated experimental runs are reported. This leaves the central claim that CLSTM-ISS improves 'reliable PS density' without a rigorous, independent basis.","section":"Section 5, Figures 9-11"}],"minor_comments":[{"comment":"The text states that 'random sampling was used to select test samples (images) from the training data,' which conflicts with the description of the Kathmandu dataset as an unseen test set in Table 1 and the following paragraph. Please clarify whether random sampling refers to validation samples only.","section":"Section 3"},{"comment":"The filter counts in Table 2 appear inconsistent: layer '(conv+BN)4+relu' is listed with 32 filters but an output dimension of 64 channels, and similar inconsistencies appear in Table 3 (including a duplicated row label '(convlstm+BN)2+relu'). Please verify the architecture tables and the corresponding text.","section":"Table 2"},{"comment":"Equation (11) is garbled in the manuscript, making it impossible to verify the ConvLSTM gate equations. A clean, correctly typeset version of the equations is needed.","section":"Equation (11)"},{"comment":"The description of the f1-loss and 'probabilistic' accuracy is confusing: stating that a true non-PS pixel with predicted probability 0.4 counts as 0.6 false positive and 0.4 true negative does not match standard definitions of accuracy or loss. Please define the loss exactly.","section":"Section 4.3"},{"comment":"Table 6 appears to contain no visible entries in the manuscript, although the text gives the key numbers (92.49%, 80.10%, 97.01%). The numerical values should be presented in the table itself.","section":"Table 6 and Figure 10"},{"comment":"The statement that STIP>35 is 'a threshold generally used to define a coherent PS pixel' is made without a citation. Please provide a reference or supporting analysis for this threshold.","section":"Section 5, STIP threshold"},{"comment":"The observation that more than 95% of pixels are non-PS is repeated nearly verbatim in Section 3 and Section 4.3; consider keeping it in one place.","section":"Sections 3 and 4.3"}],"recommendation":"reject","confidential_remarks":"The central evaluation is circular: training labels come from StaMPS, and the external reliability metric (STIP) originates from the authors' own prior papers. The absence of an independent benchmark is not a cosmetic issue; it directly undermines the headline claim that CLSTM-ISS 'improved the classification of PS and non-PS pixels compared to those of StaMPS.' I would only reconsider if the authors can supply an independent validation (e.g., synthetic data with known PS locations, corner reflectors, or comparison against another published PS-selection algorithm) and reframe the claims accordingly. The paper also cites little of the existing deep-learning-for-InSAR literature, despite the presence of related work."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my read of arXiv:1909.01868. The genuinely new thing is using CNN and ConvLSTM semantic segmentation for PS pixel selection in MT-InSAR, trained on real interferometric stacks from Delhi, Ahmedabad, Nainital, and tested on Kathmandu. The authors report a real practical gain: once trained, selection takes minutes instead of hours. CLSTM-ISS output is visually denser in urban areas and cleaner in vegetation. The reasoning about avoiding pooling layers is thoughtful. The literature gap appears genuine.\n\nThe soft spot is evaluation of the central claim. The 93.50% validation accuracy is computed against StaMPS-generated labels, the same labels used for training, so it measures agreement with StaMPS, not correctness or superiority. The paper itself says true labels don't exist and that such metrics would give an incorrect estimate, then substitutes a combination of R-index, land-cover class, and the STIP metric from the authors' own earlier papers. That heuristic may be reasonable but is not independent ground truth. Since CLSTM-ISS selects about five times more PS pixels than StaMPS (192k vs 38k), higher \"reliable PS density\" could partly reflect false positives. Velocity and time-series comparisons are qualitative visual pattern matching. So the claim of improvement over StaMPS is not supported in this version.\n\nThat said, this is not a sloppy paper. The authors are transparent about the label problem. The contribution is better framed as a fast learned approximation of StaMPS, with plausibly better density in some terrain classes, rather than a proven improvement. With open code, weights, and a check against a geodetic benchmark or a separate MT-InSAR processor, this could be a solid paper. The computational efficiency argument alone is worth serious attention. The intended readers are InSAR method developers and near-real-time monitoring groups; an ML-centric audience gets less.\n\nI'd send it to peer review, not desk reject. A competent referee can require the missing independent validation and a more careful claim. The paper deserves referee time, but acceptance should hinge on whether the improvement claim survives with real ground truth.","headline":"New application with a real speedup, but the headline accuracy is circular and the improvement claim needs independent validation before it can be believed.","tokens_in":20308,"tokens_out":2637,"would_cite":false,"duration_ms":25107,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A spatio-temporal deep network selects more reliable radar persistent scatterers than the standard StaMPS method.","keywords":["persistent scatterer interferometry","PS pixel selection","deep learning","convolutional LSTM","semantic segmentation","interferometric phase stack","Sentinel-1","deformation monitoring"],"falsifier":"Run CLSTM-ISS on a fresh non-urban site with corner reflectors or GPS-validated displacement, and compare phase noise and velocity error on the pixels CLSTM-ISS adds beyond StaMPS; if the added pixels are not more phase-stable, the improvement claim collapses.","tokens_in":19311,"feed_emoji":"🛰️","tokens_out":8049,"duration_ms":74490,"temperature":0.7,"pith_summary":"The paper claims that a deep learning network can replace the slow, iterative PS pixel-selection step in multi-temporal InSAR with a fast pixel-wise classification. Two architectures are proposed, CNN-ISS and CLSTM-ISS, trained on roughly 10,000 real interferometric image patches labelled by the StaMPS algorithm, an open-access persistent-scatterer processing chain. On the unseen Kathmandu test site, CLSTM-ISS reports 93.50% validation accuracy, selects more PS pixels in urban and lengthening areas, selects fewer in forest and water, and completes selection in 8.2 minutes compared with 108.2 minutes for StaMPS. The central claim is that CLSTM-ISS, by learning both spatial and temporal phase behaviour, produces a higher density of reliable PS pixels than either StaMPS or the spatial-only CNN-ISS.","feed_headline":"Deep learning picks reliable radar pixels in minutes","feed_subtitle":"On unseen Kathmandu data, CLSTM-ISS scores 93.5 percent accuracy and runs in 8 minutes versus 108 for StaMPS.","key_machinery":"The load-bearing component is the convolutional LSTM (convlstm) cell, which replaces the matrix multiplications inside an LSTM with convolutions over image neighbourhoods. This lets the network carry a cell state through the ten interferograms and learn spatial phase patterns and their temporal coherence together. The full CLSTM-ISS stacks two convlstm layers, one convolution layer, dropout, and a fully connected output, and it is trained with an f1-loss in which the PS class receives a weight of 200 against 1 for non-PS. Pooling layers are deliberately omitted because PS pixels are isolated and pooling would bias learning toward spatially correlated nuisance phase components.","core_discovery":"The central discovery claimed is that a convolutional long short-term memory network, CLSTM-ISS, which treats the interferogram stack as an image time series, classifies PS and non-PS pixels better than the StaMPS algorithm and better than the spatial-only CNN-ISS. On the unseen Kathmandu test set, CLSTM-ISS achieves 93.50% validation accuracy versus 89.21% for CNN-ISS, selects 192,177 PS pixels versus 38,286 for StaMPS, and 97.01% of its PS pixels pass the STIP greater-than-35 reliability threshold versus 92.49% for StaMPS and 80.10% for CNN-ISS. In area-wise terms it detects the highest PS density in man-made areas (52.9%) and lengthening areas (48.1%) and the lowest in forest and vegetation (3.4%). The authors interpret this as CLSTM-ISS learning the true spatio-temporal coherence of scatterers rather than simply reproducing its StaMPS training labels.","pith_inferences":["The reliability comparison rests on STIP, a coherence measure introduced by the same group; an independent check using corner reflectors or GPS-validated deformation is needed before generalising the \"more reliable\" claim.","The paper demonstrates ten-interferogram stacks from Sentinel-1; generalisation to other stack lengths, sensors, or orbital geometries is plausible but not demonstrated, and would need its own training data.","The clean separation of forest, water, and uncropped land suggests the same spatio-temporal architecture could segment other decorrelated terrain types, such as snow, cropland with seasonal cycles, or wetlands, if labelled stacks were available.","If the speed advantage holds, the practical bottleneck in MT-InSAR would shift from PS selection to phase unwrapping and time-series inversion, which the paper does not replace."],"forward_implications":["A trained CLSTM-ISS could cut PS selection from hours or days to minutes, making near-real-time deformation monitoring feasible for repeated satellite acquisitions.","The higher density of reliable PS pixels should improve phase unwrapping and produce clearer velocity maps, since CLSTM-ISS retains the StaMPS-like displacement pattern while adding coherent points.","Because CLSTM-ISS includes 78.44% of the pixels StaMPS selects, it is unlikely to lose the information current processing chains rely on, while adding new coherent pixels in man-made and lengthening terrain.","The method is not tied to StaMPS's proprietary logic: since labels came from an open-access algorithm and the input is a standard interferometric stack, the same training scheme can be adapted as better training labels become available."],"supporting_citations":[{"why":"Supplies the StaMPS phase-stability algorithm whose pixel labels are used as training targets.","marker":"Hooper et al. 2007"},{"why":"Defines the amplitude-dispersion PS selection baseline that the paper compares against as the conventional method.","marker":"Ferretti et al. 2001"},{"why":"Provides the SqueeSAR/SHP alternative PS-selection benchmark mentioned for context.","marker":"Ferretti et al. 2011"},{"why":"Gives the criterion that only the highest-temporal-coherence pixel in a cluster should be kept, used to interpret CNN-ISS clustering.","marker":"Agram and Zebker 2007"},{"why":"Defines STIP, the similar-time-series-interferometric-pixel measure used as the reliability threshold in evaluation.","marker":"Narayan et al. 2018a"},{"why":"Further develops the STIP/coherence measure that supplies the greater-than-35 STIP threshold for reliable PS pixels.","marker":"Narayan et al. 2018b"},{"why":"Supports the R-index and land-cover expectation that PS pixels concentrate on man-made objects and slopes facing the satellite.","marker":"Notti et al. 2011"},{"why":"Provides the previously reported Kathmandu subsidence pattern used to check the velocity maps produced from each PS selection.","marker":"Krishnan et al. 2018"}],"fun_headline_variants":["AI radar pixel picker hits 93.5%, beats StaMPS","Deep learning speeds radar pixel selection to 93.5%","CLSTM-ISS: radar pixels at 93.5% accuracy, faster","93.5% accurate radar pixels: AI outdoes classic","Faster radar pixel selection: AI scores 93.5%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison rests on assuming that a pixel is good if it lies where the slope-orientation R-index and land-cover map say PS pixels should occur and if more than 35 similar-time-series neighbours surround it; if those proxies are wrong, the claim that CLSTM-ISS outperforms StaMPS is unsupported.","fun_headline_variants_meta":{"raw":{"variants":["AI radar pixel picker hits 93.5%, beats StaMPS","Deep learning speeds radar pixel selection to 93.5%","CLSTM-ISS: radar pixels at 93.5% accuracy, faster","93.5% accurate radar pixels: AI outdoes classic","Faster radar pixel selection: AI scores 93.5%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000563,"raw_usage":{"total_tokens":2744,"prompt_tokens":1089,"completion_tokens":1655,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":705,"completion_tokens_details":{"reasoning_tokens":1561}},"tokens_in":705,"tokens_out":1655,"duration_ms":14480,"temperature":1.0,"reasoning_tokens":1561,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T05:05:55.970812+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run CLSTM-ISS on a fresh non-urban site with corner reflectors or GPS-validated displacement, and compare phase noise and velocity error on the pixels CLSTM-ISS adds beyond StaMPS; if the added pixels are not more phase-stable, the improvement claim collapses.","supporting_citations":[],"review_version":1}