REVIEW 1 major objections 26 references
A diffusion model trained on Sentinel-2B flood scenes removes clouds while preserving water body continuity and spectral signatures needed for inundation mapping.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.3
2026-06-28 15:14 UTC pith:5AYDJFUB
load-bearing objection Applies a masked diffusion transformer to cloud removal in Sentinel-2 flood scenes but the abstract supplies no metrics or baselines to back the performance claims. the 1 major comments →
Deep Learning for Remote Sensing to Improve Flood Inundation Mapping
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
A cloud-removal framework based on Denoising Diffusion Probabilistic Models and the Masked Diffusion Transformer architecture, trained on multispectral Sentinel-2B flood scenes with realistic cloud patterns, generates cloud-free image realizations that preserve both visual fidelity and hydrological consistency, providing improved continuity of water bodies and preservation of spectral signatures critical for water detection indices.
What carries the argument
Masked Diffusion Transformer, which applies self-attention across wider spatial context and uses masked token modeling to reconstruct cloud-obscured regions in multispectral flood imagery.
Load-bearing premise
Reconstructions from the model trained on Sentinel-2B scenes with realistic cloud patterns will preserve hydrological consistency and spectral signatures in real unseen cloud-obscured flood events.
What would settle it
Quantitative comparison of water detection indices computed on the model's outputs against the same indices from actual cloud-free Sentinel-2 acquisitions of identical flood events, or against independent in-situ flood extent measurements.
If this is right
- Reconstructed images maintain continuity of water bodies across cloud-obscured areas.
- Spectral signatures required for standard water detection indices remain intact.
- Continuous optical observations become available even during peak cloud cover in extreme precipitation.
- The generated images support more reliable inputs for flood-related decision making and disaster risk management.
Where Pith is reading between the lines
- The same training strategy could be applied to other optical sensors to test whether the hydrological consistency holds beyond Sentinel-2B.
- Integration with radar-based flood products could be tested to see if the diffusion outputs improve multi-sensor fusion during persistent cloud cover.
- The framework might enable retrospective reconstruction of historical flood events from partially clouded archives.
- Operational pipelines could use the model outputs to reduce gaps in near-real-time inundation maps for early warning systems.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript introduces a cloud-removal framework for optical flood imagery based on Denoising Diffusion Probabilistic Models using the Masked Diffusion Transformer architecture. The model is trained on multispectral Sentinel-2B flood scenes containing realistic cloud patterns and is claimed to generate cloud-free realizations that preserve visual fidelity and hydrological consistency. Evaluation is described using standard image-quality metrics together with flood-specific hydrological measures, with the central assertion that the method improves continuity of water bodies and preserves spectral signatures needed for water-detection indices.
Significance. If the quantitative results and validation details support the claims, the work would offer a generative-modeling alternative to temporal compositing or interpolation for cloud removal in flood monitoring. This could be relevant for maintaining continuous optical observations during extreme precipitation events, with potential utility for disaster-risk applications. The architecture choices (self-attention and masked token modeling) are standard extensions of diffusion models and do not introduce obvious internal inconsistencies.
major comments (1)
- [Abstract] Abstract: the assertion that the model 'demonstrates improved continuity of water bodies and preservation of spectral signatures critical for water detection indices' is unsupported by any quantitative metrics, baseline comparisons, error bars, or validation details. This absence directly undermines evaluation of the central claim that the approach is 'robust and physically consistent.'
Simulated Author's Rebuttal
We thank the referee for the constructive feedback. We address the single major comment below and agree that the abstract requires revision for clarity and precision.
read point-by-point responses
-
Referee: [Abstract] Abstract: the assertion that the model 'demonstrates improved continuity of water bodies and preservation of spectral signatures critical for water detection indices' is unsupported by any quantitative metrics, baseline comparisons, error bars, or validation details. This absence directly undermines evaluation of the central claim that the approach is 'robust and physically consistent.'
Authors: We acknowledge the referee's point that the abstract presents a high-level claim without explicit quantitative anchors. The full manuscript (Section 4) reports the supporting evaluation: standard metrics (PSNR, SSIM, LPIPS) with baseline comparisons to temporal interpolation and GAN-based methods, plus flood-specific measures (water-body continuity via connected-component analysis and spectral fidelity via NDWI/NDWI correlation), all with error bars across multiple test scenes. However, we agree the abstract should not stand alone without clearer linkage. We will revise the abstract to either include concise quantitative highlights or rephrase the claim to explicitly reference the evaluation results. revision: yes
Circularity Check
No significant circularity detected
full rationale
The paper presents a standard supervised training pipeline for a Masked Diffusion Transformer on external Sentinel-2B multispectral scenes containing realistic cloud patterns. Performance is assessed with independent image-quality metrics and flood-specific hydrological measures. No equations, fitted parameters renamed as predictions, self-citation load-bearing steps, or ansatz smuggling appear in the provided abstract or described method. The central claim rests on empirical generalization rather than reducing to its own inputs by construction.
Axiom & Free-Parameter Ledger
Cite this review
Pith. "Pith review of Deep Learning for Remote Sensing to Improve Flood Inundation Mapping." pith.science (2026). https://pith.science/paper/5AYDJFUB
@misc{pith2026260602310,
author = {Pith},
title = {Pith review of: Deep Learning for Remote Sensing to Improve Flood Inundation Mapping},
year = {2026},
howpublished = {\url{https://pith.science/paper/5AYDJFUB}},
note = {Machine review of arXiv:2606.02310}
}
read the original abstract
Flooding is the most pervasive natural disaster worldwide. Timely and accurate flood inundation mapping are essential for informing disaster risk management. Optical satellite missions provide high-resolution, multispectral observations critical for flood detection and inundation mapping. However, their operational utility is severely constrained by cloud cover during extreme precipitation events. Conventional cloud-removal techniques based on temporal compositing or interpolation often fail to capture inundation dynamics. In this study, we introduce a cloud-removal framework for flood imagery based on Denoising Diffusion Probabilistic Models, leveraging the Masked Diffusion Transformer architecture. The proposed approach exploits self-attention mechanisms to capture wider spatial context and employs masked token modeling to explicitly learn the reconstruction of cloud-obscured regions. Trained on multispectral Sentinel-2B flood scenes with realistic cloud patterns, the model generates cloud-free image realizations that preserve both visual fidelity and hydrological consistency. Reconstruction performance is evaluated using standard image quality metrics alongside flood-specific hydrological measures, demonstrating improved continuity of water bodies and preservation of spectral signatures critical for water detection indices. The results indicate that diffusion-based generative modeling offers a robust and physically consistent alternative for cloud removal in optical flood monitoring, enabling more reliable, continuous observations to support disaster risk management and flood-related decision making.
Figures
Reference graph
Works this paper leans on
-
[1]
Mobility and Resilience : A Global Assessment of Flood Impacts on Road Transportation Networks,
Y . He, J. E. Maruyama Rentschler, P. Avner, J. Gao, X. Yue, and J. Radke, “Mobility and Resilience : A Global Assessment of Flood Impacts on Road Transportation Networks,”Policy Research Working Paper Series, May 2022. [Online]. Available: https: //ideas.repec.org//p/wbk/wbrwps/10049.html
2022
-
[2]
High-resolution mapping of global surface water and its long-term changes,
J.-F. Pekel, A. Cottam, N. Gorelick, and A. S. Belward, “High-resolution mapping of global surface water and its long-term changes,”Nature, vol. 540, no. 7633, pp. 418–422, 2016
2016
-
[3]
Global flood extent segmentation in optical satellite images,
E. Portal ´es-Juli`a, G. Mateo-Garc ´ıa, C. Purcell, and L. G ´omez-Chova, “Global flood extent segmentation in optical satellite images,”Scientific reports, vol. 13, no. 1, p. 20316, 2023
2023
-
[4]
Effective- ness of sentinel-1 and sentinel-2 for flood detection as- sessment in europe,
A. Tarpanelli, A. C. Mondini, and S. Camici, “Effective- ness of sentinel-1 and sentinel-2 for flood detection as- sessment in europe,”Natural Hazards and Earth System Sciences, vol. 22, no. 8, pp. 2473–2489, 2022
2022
-
[5]
Intro- ducing a new index for flood mapping using sentinel-2 imagery (sfmi),
H. Farhadi, H. Ebadi, A. Kiani, and A. Asgary, “Intro- ducing a new index for flood mapping using sentinel-2 imagery (sfmi),”Computers & Geosciences, vol. 194, p. 105742, 2025
2025
-
[6]
Overview of sentinel-2,
F. Spoto, O. Sy, P. Laberinti, P. Martimort, V . Fernandez, O. Colin, B. Hoersch, and A. Meygret, “Overview of sentinel-2,” in2012 IEEE international geoscience and remote sensing symposium. IEEE, 2012, pp. 1707–1710
2012
-
[7]
Sentinel-2: Esa’s optical high-resolution mission for gmes operational services,
M. Drusch, U. Del Bello, S. Carlier, O. Colin, V . Fernan- dez, F. Gascon, B. Hoersch, C. Isola, P. Laberinti, P. Mar- timortet al., “Sentinel-2: Esa’s optical high-resolution mission for gmes operational services,”Remote sensing of Environment, vol. 120, pp. 25–36, 2012
2012
-
[8]
Cloud cover throughout the agricultural growing season: Impacts on passive optical earth obser- vations,
A. K. Whitcraft, E. F. Vermote, I. Becker-Reshef, and C. O. Justice, “Cloud cover throughout the agricultural growing season: Impacts on passive optical earth obser- vations,”Remote sensing of Environment, vol. 156, pp. 438–447, 2015
2015
-
[9]
Sentinel 1 evolution: Sentinel-1c and-1d mod- els,
R. Torres, S. Lokas, G. Di Cosimo, D. Geudtner, and D. Bibby, “Sentinel 1 evolution: Sentinel-1c and-1d mod- els,” in2017 IEEE International Geoscience and Remote Sensing Symposium (IGARSS). IEEE, 2017, pp. 5549– 5550
2017
-
[10]
A speckle filter for sentinel-1 sar ground range detected data based on residual convolutional neu- ral networks,
A. Sebastianelli, M. P. Del Rosso, S. L. Ullo, and P. Gamba, “A speckle filter for sentinel-1 sar ground range detected data based on residual convolutional neu- ral networks,”IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 15, pp. 5086–5101, 2022
2022
-
[11]
Sensitivity of sentinel-1 backscatter to characteristics of buildings,
K. Koppel, K. Zalite, K. V oormansik, and T. Jagdhuber, “Sensitivity of sentinel-1 backscatter to characteristics of buildings,”International Journal of Remote Sensing, vol. 38, no. 22, pp. 6298–6318, 2017
2017
-
[12]
A method for compositing polar modis satellite images to remove cloud cover for landfast sea-ice detection,
A. D. Fraser, R. A. Massom, and K. J. Michael, “A method for compositing polar modis satellite images to remove cloud cover for landfast sea-ice detection,”IEEE transactions on geoscience and remote sensing, vol. 47, no. 9, pp. 3272–3282, 2009
2009
-
[13]
Automatic mosaicking of satellite imagery con- sidering the clouds,
Y . Kang, L. Pan, Q. Chen, T. Zhang, S. Zhang, and Z. Liu, “Automatic mosaicking of satellite imagery con- sidering the clouds,”ISPRS Annals of the Photogramme- try, Remote Sensing and Spatial Information Sciences, vol. 3, pp. 415–421, 2016
2016
-
[14]
Multi- temporal landsat data automatic cloud removal using poisson blending,
C. Hu, L.-Z. Huo, Z. Zhang, and P. Tang, “Multi- temporal landsat data automatic cloud removal using poisson blending,”IEEE Access, vol. 8, pp. 46 151– 46 161, 2020
2020
-
[15]
Missing information reconstruction of remote sensing data: A technical review,
H. Shen, X. Li, Q. Cheng, C. Zeng, G. Yang, H. Li, and L. Zhang, “Missing information reconstruction of remote sensing data: A technical review,”IEEE Geoscience and Remote Sensing Magazine, vol. 3, no. 3, pp. 61–85, 2015
2015
-
[16]
Generative deep learning models for cloud removal in satellite imagery: A comparative review of gans and diffusion methods,
S. Edirisinghe, B. Schoen-Phelan, and S. Hensman, “Generative deep learning models for cloud removal in satellite imagery: A comparative review of gans and diffusion methods,”ISPRS Open Journal of Photogram- metry and Remote Sensing, p. 100110, 2025
2025
-
[17]
Cloud removal in remote sensing images using genera- tive adversarial networks and sar-to-optical image trans- lation,
F. N. Darbaghshahi, M. R. Mohammadi, and M. Soryani, “Cloud removal in remote sensing images using genera- tive adversarial networks and sar-to-optical image trans- lation,”IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–9, 2021
2021
-
[18]
Cloud removal using multimodal gan with adversarial consistency loss,
Y . Zhao, S. Shen, J. Hu, Y . Li, and J. Pan, “Cloud removal using multimodal gan with adversarial consistency loss,” IEEE Geoscience and Remote Sensing Letters, vol. 19, pp. 1–5, 2021
2021
-
[19]
An in-depth review and analysis of mode collapse in generative adversarial net- works,
F. L. Barsha and W. Eberle, “An in-depth review and analysis of mode collapse in generative adversarial net- works,”Machine Learning, vol. 114, no. 6, p. 141, 2025
2025
-
[20]
Cloud-gan: Cloud removal for sentinel-2 imagery using a cyclic consistent genera- tive adversarial networks,
P. Singh and N. Komodakis, “Cloud-gan: Cloud removal for sentinel-2 imagery using a cyclic consistent genera- tive adversarial networks,” inIGARSS 2018-2018 IEEE International Geoscience and Remote Sensing Sympo- sium. IEEE, 2018, pp. 1772–1775
2018
-
[21]
Denoising diffusion probabilistic models,
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,”Advances in neural information processing systems, vol. 33, pp. 6840–6851, 2020
2020
-
[22]
Masked diffusion transformer is a strong image synthesizer,
S. Gao, P. Zhou, M.-M. Cheng, and S. Yan, “Masked diffusion transformer is a strong image synthesizer,” in Proceedings of the IEEE/CVF international conference on computer vision, 2023, pp. 23 164–23 173
2023
-
[23]
Unraveling the 2021 central tennessee flood event using a hierarchical multi-model inundation modeling frame- work,
S. Gangrade, G. R. Ghimire, S.-C. Kao, M. Morales- Hern´andez, A. A. Tavakoly, J. L. Gutenson, K. H. Spar- row, G. K. Darkwah, A. J. Kalyanapu, and M. L. Follum, “Unraveling the 2021 central tennessee flood event using a hierarchical multi-model inundation modeling frame- work,”Journal of Hydrology, vol. 625, p. 130157, 2023
2021
-
[24]
Cloudsen12, a global dataset for semantic un- derstanding of cloud and cloud shadow in sentinel-2,
C. Aybar, L. Ysuhuaylas, J. Loja, K. Gonzales, F. Her- rera, L. Bautista, R. Yali, A. Flores, L. Diaz, N. Cuenca et al., “Cloudsen12, a global dataset for semantic un- derstanding of cloud and cloud shadow in sentinel-2,” Scientific data, vol. 9, no. 1, p. 782, 2022
2022
-
[25]
Nas-unet: Neural architecture search for medical image segmentation,
Y . Weng, T. Zhou, Y . Li, and X. Qiu, “Nas-unet: Neural architecture search for medical image segmentation,” IEEE access, vol. 7, pp. 44 247–44 257, 2019
2019
-
[26]
A vit- based multiscale feature fusion approach for remote sens- ing image segmentation,
W. Wang, C. Tang, X. Wang, and B. Zheng, “A vit- based multiscale feature fusion approach for remote sens- ing image segmentation,”IEEE Geoscience and Remote Sensing Letters, vol. 19, pp. 1–5, 2022
2022
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.