REVIEW 3 major objections 6 minor 79 references
Improving Satellite Imagery Masking using Multi-task and Transfer Learning
T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read A single multi-task network can replace the multi-source satellite masking pipeline, predicting water, cloud, cloud shadow, snow/ice, and terrain shadow from six HLS bands with higher accuracy and a 30x faster downstream sediment pipeline.
desk verdict Useful engineering result: a single multi-task network that predicts all five masks from six HLS bands, with honest discussion of its own label circularity; worth serious review, but the headline accuracy numbers only measure agreement with DSWx. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the multi-task segmentation model: one backbone $f_\theta$ maps a $512\times512\times6$ HLS patch to a shared feature map $z_i$, and five small heads $g^m_{\phi_m}$ each map $z_i$ to a per-pixel probability mask, trained with the summed binary cross-entropy loss $\mathcal{L}=\frac{1}{N}\sum_i\sum_m \mathcal{L}_{\text{bce}}(\hat{y}^m_i, y^m_i)$. This machinery carries both claims at once: sharing the backbone is what makes five masks nearly free at inference, and sharing the representation is what improves rare classes such as snow/ice and cloud shadow by letting the model use context from easier classes like water and cloud. The second piece of machinery is transfer learning, specifically pre-training the backbone on Satlas, a large remote-sensing dataset, before fine-tuning on DSWx labels.
What would settle it
A decisive test would be to score the Swin-T/Satlas multi-task model against a new, globally distributed set of manually labeled HLS pixels for all five masks: if water F1 fell below DeepWaterMap's 82.21% or cloud-shadow F1 fell below LANA's 57.53% on that independent set, the DSWx-reported gains would not transfer to real scenes. The LANA evaluation already performs a smaller version of this for cloud and shadow; extending it to water, snow/ice, and terrain shadow is the missing check.
Extended reading notes
Core claim
The central claim, on the paper's own terms, is that the five masks required to isolate 'good quality' water pixels can be produced by one end-to-end segmentation model from just the six reflective HLS bands (blue, green, red, NIR, SWIR-1, SWIR-2), without DEMs, sun-position calculations, or separate Fmask runs. The model uses a shared backbone that outputs a common feature representation and five lightweight heads, one per mask, trained jointly with summed binary cross-entropy loss. Multi-tasking is not just a cost-saving trick: compared with single-task versions of the same architectures, the multi-task model is better or comparable on every mask and improves snow/ice F1 by more than 10 points on CNN backbones. The paper also claims that transfer learning from large remote-sensing datasets is decisive, with Satlas pre-training lifting Swin-T's water F1 from 80.73% to 91.10% and its cloud-shadow F1 from 35.81% to 63.10%. When the multi-task masker replaces the standard Fmask-plus-DEM pipeline inside an SSC estimation system, runtime for 400k samples falls from 86.84 days to as little as 2.78 days and SSC RMSE drops by 2.64 mg/L.
Load-bearing premise
The load-bearing premise is that DSWx's algorithmically generated masks are trustworthy enough to serve as ground truth, which the authors themselves qualify in Section V-D by saying the model is 'making a model of a model' — so any systematic DSWx error such as false snow/ice, dilated clouds, or missed clouds over water becomes the model's ceiling.
Editorial extensions
If this is right
- Replacing the standard pipeline with the multi-task model makes global-scale SSC processing feasible: 400,000 HLS samples per day would take 2.78 days with MobileNetv3 instead of 86.84 days on four CPU cores.
- Because all five masks come from one forward pass on six HLS bands, the same masking step can be rerun on historical HLS imagery, which DSWx and other auxiliary products do not cover.
- The best mask accuracy is obtained with Swin-T pre-trained on Satlas, but MobileNetv3 is within a few F1 points on most masks and is roughly six times faster, so the collection of models covers a practical speed-accuracy spectrum.
- Multi-task training improves over single-task training by more than 10 F1 points for snow/ice on CNN backbones, consistent with rare classes benefiting from shared context.
- Downstream SSC estimation improves when masking improves: RMSE drops by 2.64 mg/L and the 95th percentile error by 13.55 mg/L.
Reading between the lines
- Editorial inference: because the training target is DSWx, the model's ceiling is the quality of DSWx; the next step the paper implicitly points to is to use the same multi-task architecture with a larger set of manually reviewed labels, which should clarify whether the remaining errors come from the network or from the labels.
- Editorial inference: the same single-model masking step could replace separate Fmask and DEM computations in other reflectance-based analyses, such as water-color retrieval, flood mapping, or lake-ice studies, where the same five obstruction classes are removed.
- Editorial inference: the reported speed and memory numbers imply that near-daily global masking is feasible on modest CPU resources; one could test this by running the MobileNetv3 pipeline on a full year of global HLS tiles and checking whether error accumulates in high-latitude or mountainous regions.
- Editorial inference: since terrain shadow reaches about 97% F1 without a DEM input, the model may be learning shadow geometry implicitly from image context; a targeted evaluation on steep topography and extreme solar angles would show whether this generalizes or is an artifact of the DSWx training distribution.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a multi-task deep learning system that predicts five masks (water, cloud, cloud shadow, snow/ice, terrain shadow) directly from six HLS spectral bands, replacing the multi-source Fmask/DEM pipeline used in standard SSC estimation. The authors train and evaluate several architectures (DeepLabv3+, MobileNetv3, SegNet, ResNet50, Swin-T, ViT-B/16) with different pre-training strategies (ImageNet, Satlas, Prithvi) on a global dataset of DSWx labels, and compare against DeepWaterMap, MNDWI, Fmask, LANA, and U-Net Wieland. They report a water-mask F1 of 91.10% for Swin-T pretrained on Satlas versus 82.21% for DeepWaterMap, improved cloud and cloud shadow F1 on the manually labeled LANA benchmark, a roughly 30x runtime speedup for a 400k-sample SSC pipeline, and a 2.64 mg/L RMSE reduction in downstream SSC estimation.
Significance. If the accuracy claims hold, the paper would be a useful contribution to operational satellite masking: a single lightweight model that predicts all required masks from one input source, with favorable speed and memory, is practically valuable for global-scale hydrology and other remote sensing workflows. The paper has several concrete strengths: the multi-task formulation is clearly described; the global DSWx dataset is split spatially at the scene level to reduce leakage; DeepWaterMap is retrained on the same DSWx training split for a controlled comparison; runtime and storage overheads are measured in detail; and the LANA benchmark provides independent manually labeled evidence for cloud and cloud shadow performance. The downstream SSC experiment, while confounded, is a genuinely useful end-to-end evaluation and the efficiency gains are credible.
major comments (3)
- [Section IV.A.3, Table V, Section V-D] The headline water-mask F1 comparison is measured entirely against DSWx labels, the same algorithm-generated product used for training. The authors explicitly state in Section V-D that they are “making a model of a model” and that any DSWx errors become their errors. Under these conditions, the 9-point F1 gain over DeepWaterMap supports fidelity to DSWx, not necessarily true masking accuracy, and the known DSWx artifacts listed in Section V-D (dilated clouds, missed clouds over water, equatorial snow/ice mislabels) could be learned and reproduced by the model. Because water, snow/ice, and terrain shadow have no manual validation, the central accuracy claim needs either an independent water-mask evaluation (for example, manual labels or a reference water product) or a careful reframing as DSWx-fidelity only.
- [Section IV.A.4, Table VI] The LANA benchmark provides independent manual labels only for cloud and cloud shadow, so it does not validate the headline water, snow/ice, or terrain shadow results. Moreover, the abstract's “at least 6% improvement” is not consistently supported by Table VI: Swin-T improves cloud F1 by only 0.54 points over LANA (92.96% vs 92.42%), while the 6-point or larger gains occur for cloud shadow and clear. The adaptation of the six-band multi-task model to the LANA Landsat 8 benchmark is also underspecified (band selection, preprocessing, tile size, and training protocol are not stated), which is needed for reproducibility and for interpreting the comparison against LANA's eight-band model.
- [Section IV.B.3, Table X] The downstream SSC comparison does not isolate mask quality. The standard and proposed pipelines differ in mask source, reprojection and alignment overhead, cloud-cover filtering, and the feature distributions fed to the SSC model, so the 2.64 mg/L RMSE reduction could stem from any of these differences rather than from more accurate masks. Additionally, the multi-task model is trained on DSWx labels from April 2023 to March 2024 while the SSC test data span April 2013 to October 2021; the temporal domain shift is not discussed. The efficiency claims in Tables VII–IX are solid, but the accuracy transfer to historical imagery needs at least an acknowledgement or a small analysis.
minor comments (6)
- [Section IV.A.3, Table V] The text says MNDWI evaluated out of the box gives a 16.62% F1 score, while Table V reports 58.43% after cloud and shadow filtering; please clarify this discrepancy and state explicitly which number corresponds to the “out of the box” evaluation.
- [Section IV.A.3, Table V] The comparison between unfiltered deep learning outputs and a cloud-and-shadow-filtered MNDWI is not apples-to-apples; please state whether the filter is a DSWx-based mask and discuss how this affects the headline gap.
- [Table VIII] In the proposed-pipeline column, the row labeled “Water” with time 1.96 s appears to aggregate all mask predictions; please restructure the table so that readers can see which components correspond to the single multi-task inference call.
- [Tables III and IV] All F1, precision, recall, and IoU numbers are reported without confidence intervals or multiple-seed variation; given the modest differences among some architectures (for example, DeepLabv3+ versus MobileNetv3), a measure of variability would help assess whether the rankings are stable.
- [Section V-A] The statement that “cloud labels from DSWx have an F1 score of 89.81%” should specify the reference set (presumably LANA manual labels) and the model or comparison used to compute that score.
- [General] The manuscript does not include a data or code availability statement; since the authors claim a new dataset and trained models, a statement about releasing code, trained weights, and the exact data splits would significantly aid reproducibility.
Circularity Check
No circularity: held-out DSWx evaluation and independent LANA validation support the empirical claims; the 'model of a model' caveat is a validity limitation, not a circular derivation.
full rationale
The paper's central derivation is empirical rather than tautological. Models are trained on a spatial held-out split of DSWx labels and evaluated on a separate held-out test split (Section II-B, Section IV-A); DeepWaterMap is retrained on the same DSWx train set for comparison, so the reported water-mask F1 gain is a fair relative measure within the DSWx label distribution. The LANA benchmark provides independent, manually labeled evidence for cloud and cloud shadow improvements (Table VI). The downstream SSC comparison uses a described two-stage MLP and compares standard versus multi-task pipelines on the same SSC test set, so the RMSE reduction is not forced by construction. The paper's explicit limitation in Section V-D—'we are, in essence, making a model of a model by predicting DSWx output'—is an honest caveat about label fidelity and external validity, not evidence that any reported metric reduces to its training input. Self-citations such as [54] and [77] provide context or component models but are not load-bearing uniqueness claims or ansatz-smuggling devices. Thus no circular step can be exhibited from the paper's own equations or definitions.
Assumptions & free parameters
free parameters (1)
- SSC model combination thresholds (two values) =
Not disclosed
assumptions (3)
- domain assumption DSWx labels approximate ground truth for water, cloud, cloud shadow, snow/ice, and terrain shadow masks.
- domain assumption Six HLS reflective bands (Blue, Green, Red, NIR, SWIR1, SWIR2) contain sufficient information to predict all five masks, including terrain shadow and snow/ice, without thermal bands or DEM.
- domain assumption The LANA manually labeled dataset is a valid independent benchmark for cloud and cloud shadow generalization.
Cite this review
Pith. "Pith review of Improving Satellite Imagery Masking using Multi-task and Transfer Learning." pith.science (2026). https://pith.science/paper/SFKLA62X
@misc{pith2026241208545,
author = {Pith},
title = {Pith review of: Improving Satellite Imagery Masking using Multi-task and Transfer Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/SFKLA62X}},
note = {Machine review of arXiv:2412.08545}
}
read the original abstract
Many remote sensing applications employ masking of pixels in satellite imagery for subsequent measurements. For example, estimating water quality variables, such as Suspended Sediment Concentration (SSC) requires isolating pixels depicting water bodies unaffected by clouds, their shadows, terrain shadows, and snow and ice formation. A significant bottleneck is the reliance on a variety of data products (e.g., satellite imagery, elevation maps), and a lack of precision in individual steps affecting estimation accuracy. We propose to improve both the accuracy and computational efficiency of masking by developing a system that predicts all required masks from Harmonized Landsat and Sentinel (HLS) imagery. Our model employs multi-tasking to share computation and enable higher accuracy across tasks. We experiment with recent advances in deep network architectures and show that masking models can benefit from these, especially when combined with pre-training on large satellite imagery datasets. We present a collection of models offering different speed/accuracy trade-offs for masking. MobileNet variants are the fastest, and perform competitively with larger architectures. Transformer-based architectures are the slowest, but benefit the most from pre-training on large satellite imagery datasets. Our models provide a 9% F1 score improvement compared to previous work on water pixel identification. When integrated with an SSC estimation system, our models result in a 30x speedup while reducing estimation error by 2.64 mg/L, allowing for global-scale analysis. We also evaluate our model on a recently proposed cloud and cloud shadow estimation benchmark, where we outperform the current state-of-the-art model by at least 6% in F1 score.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
Remote sensing of suspended sediments and shallow coastal waters,
R.-R. Li and B.-C. Gao, “Remote sensing of suspended sediments and shallow coastal waters,” IEEE Transactions on Geoscience and Remote Sensing, vol. 41, no. 3, pp. 559–566, 2003. 1
work page 2003
-
[2]
J. C. Tilton, W. T. Lawrence, and A. J. Plaza, “Utilizing hierarchical segmentation to generate water and snow masks to facilitate monitoring change with remotely sensed image data,” GIScience & Remote Sensing, vol. 43, no. 1, pp. 39–66, 2006. 1
work page 2006
-
[3]
Automated masking of cloud and cloud shadow for forest change analysis using landsat images,
C. Huang, N. Thomas, S. N. Goward, J. G. Masek, Z. Zhu, J. R. Town- shend, and J. E. V ogelmann, “Automated masking of cloud and cloud shadow for forest change analysis using landsat images,” International Journal of Remote Sensing , vol. 31, no. 20, pp. 5449–5464, 2010. 1
work page 2010
-
[4]
R. A. Aravena, M. B. Lyons, and D. A. Keith, “High resolution forest masking for seasonal monitoring with a regionalized and colourimetri- cally assisted chorologic typology,” Remote Sensing, vol. 15, no. 14, p. 3457, 2023. 1
work page 2023
-
[5]
C. Atzberger, “Advances in remote sensing of agriculture: Context de- scription, existing operational monitoring systems and major information needs,” Remote sensing, vol. 5, no. 2, pp. 949–981, 2013. 1
work page 2013
-
[6]
S. Valero, D. Morin, J. Inglada, G. Sepulcre, M. Arias, O. Hagolle, G. Dedieu, S. Bontemps, P. Defourny, and B. Koetz, “Production of a dynamic cropland mask by processing remote sensing image series at high temporal and spatial resolutions,” Remote Sensing , vol. 8, no. 1, p. 55, 2016. 1
work page 2016
-
[7]
M. Huang, N. Chen, W. Du, Z. Chen, and J. Gong, “Dmblc: An indirect urban impervious surface area extraction approach by detecting and masking background land cover on google earth image,” Remote Sensing, vol. 10, no. 5, p. 766, 2018. 1
work page 2018
-
[8]
Using aviris data and multiple- masking techniques to map urban forest tree species,
Q. Xiao, S. Ustin, and E. McPherson, “Using aviris data and multiple- masking techniques to map urban forest tree species,” International Journal of Remote Sensing , vol. 25, no. 24, pp. 5637–5654, 2004. 1
work page 2004
Show all 79 references
-
[9]
Improved landsat operational land imager (oli) cloud and shadow detection with the learning attention network algorithm (lana),
H. K. Zhang, D. Luo, and D. P. Roy, “Improved landsat operational land imager (oli) cloud and shadow detection with the learning attention network algorithm (lana),” Remote Sensing, vol. 16, no. 8, p. 1321, 2024. 1, 2, 6, 10, 12, 17
2024
-
[10]
Fmask 4.0: Improved cloud and cloud shadow detection in landsats 4–8 and sentinel-2 imagery,
S. Qiu, Z. Zhu, and B. He, “Fmask 4.0: Improved cloud and cloud shadow detection in landsats 4–8 and sentinel-2 imagery,” Remote Sensing of Environment, vol. 231, p. 111205, 2019. 1, 6, 12, 17
2019
-
[11]
Improvement and expansion of the fmask algorithm: Cloud, cloud shadow, and snow detection for landsats 4–7, 8, and sentinel 2 images,
Z. Zhu, S. Wang, and C. E. Woodcock, “Improvement and expansion of the fmask algorithm: Cloud, cloud shadow, and snow detection for landsats 4–7, 8, and sentinel 2 images,” Remote sensing of Environment, vol. 159, pp. 269–277, 2015. 1
2015
-
[12]
A novel water change tracking algorithm for dynamic mapping of inland water using time- series remote sensing imagery,
X. Chen, L. Liu, X. Zhang, S. Xie, and L. Lei, “A novel water change tracking algorithm for dynamic mapping of inland water using time- series remote sensing imagery,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing , vol. 13, pp. 1661– 1674, 2020. 1
2020
-
[13]
Simple method to extract lake ice condition from landsat images,
X. Yang, T. M. Pavelsky, L. P. Bendezu, and S. Zhang, “Simple method to extract lake ice condition from landsat images,” IEEE Transactions on Geoscience and Remote Sensing , vol. 60, pp. 1–10, 2021. 1 19
2021
-
[14]
The color of rivers,
J. R. Gardner, X. Yang, S. N. Topp, M. R. Ross, E. H. Altenau, and T. M. Pavelsky, “The color of rivers,” Geophysical Research Letters , vol. 48, no. 1, p. e2020GL088946, 2021. 1
2021
-
[15]
Toward improved accuracy of remote sensing approaches for quantifying suspended sediment: Im- plications for suspended-sediment monitoring,
E. Dethier, C. Renshaw, and F. Magilligan, “Toward improved accuracy of remote sensing approaches for quantifying suspended sediment: Im- plications for suspended-sediment monitoring,” Journal of Geophysical Research: Earth Surface, vol. 125, no. 7, p. e2019JF005033, 2020. 1, 3
2020
-
[16]
Human activities change suspended sediment concentration along rivers,
J. Gardner, T. Pavelsky, S. Topp, X. Yang, M. R. Ross, and S. Co- hen, “Human activities change suspended sediment concentration along rivers,” Environmental Research Letters, vol. 18, no. 6, p. 064032, 2023. 1
2023
-
[17]
Suspended sediment monitoring and assessment for yellow river estuary from landsat tm and etm+ imagery,
M. Zhang, Q. Dong, T. Cui, C. Xue, and S. Zhang, “Suspended sediment monitoring and assessment for yellow river estuary from landsat tm and etm+ imagery,” Remote Sensing of Environment, vol. 146, pp. 136–147,
-
[18]
Simultaneous remote sensing of river discharge and suspended sedi- ment on the sagavanirktok river, alaska
T. Langhorst, T. Pavelsky, M. Harlan, E. Friedmann, and C. J. Gleason, “Simultaneous remote sensing of river discharge and suspended sedi- ment on the sagavanirktok river, alaska.” inAGU Fall Meeting Abstracts, vol. 2022, 2022, pp. H32G–06. 1
2022
-
[19]
Riverine sediment response to deforestation in the amazon basin,
A. Narayanan, S. Cohen, and J. R. Gardner, “Riverine sediment response to deforestation in the amazon basin,” Earth Surface Dynamics, vol. 12, no. 2, pp. 581–599, 2024. 1
2024
-
[20]
Modification of normalised difference water index (ndwi) to enhance open water features in remotely sensed imagery,
H. Xu, “Modification of normalised difference water index (ndwi) to enhance open water features in remotely sensed imagery,” International journal of remote sensing , vol. 27, no. 14, pp. 3025–3033, 2006. 1, 6, 17
2006
-
[21]
Water bodies’ mapping from sentinel-2 imagery with modified normalized difference water index at 10-m spatial resolution produced by sharpening the swir band,
Y . Du, Y . Zhang, F. Ling, Q. Wang, W. Li, and X. Li, “Water bodies’ mapping from sentinel-2 imagery with modified normalized difference water index at 10-m spatial resolution produced by sharpening the swir band,” Remote Sensing, vol. 8, no. 4, p. 354, 2016. 1
2016
-
[22]
Donchyts, J
G. Donchyts, J. Schellekens, H. Winsemius, E. Eisemann, and N. Van de Giesen, “A 30 m resolution surface water mask including estimation of positional and thematic differences using landsat 8, srtm and open- streetmap: a case study in the murray-darling basin, australia,” Remo...
2016
-
[23]
Estimating fractional snow cover from modis using the normalized difference snow index,
V . V . Salomonson and I. Appel, “Estimating fractional snow cover from modis using the normalized difference snow index,” Remote sensing of environment, vol. 89, no. 3, pp. 351–360, 2004. 1
2004
-
[24]
An improved scheme for correcting remote spectral surface reflectance simultaneously for terrestrial brdf and water-surface sunglint in coastal environments,
E. Greenberg, D. R. Thompson, D. Jensen, P. A. Townsend, N. Queally, A. Chlus, C. G. Fichot, J. P. Harringmeyer, and M. Simard, “An improved scheme for correcting remote spectral surface reflectance simultaneously for terrestrial brdf and water-surface sunglint in coastal envi...
2022
-
[25]
Correction of sunglint effects in high spatial resolution hyperspectral imagery using swir or nir bands and taking account of spectral variation of refractive index of water,
B.-C. Gao and R.-R. Li, “Correction of sunglint effects in high spatial resolution hyperspectral imagery using swir or nir bands and taking account of spectral variation of refractive index of water,” Advances in Environmental and Engineering Research, vol. 2, no. 3, pp. 1–15, 2021. 2
2021
-
[26]
A high-accuracy map of global terrain elevations,
D. Yamazaki, D. Ikeshima, R. Tawatari, T. Yamaguchi, F. O’Loughlin, J. C. Neal, C. C. Sampson, S. Kanae, and P. D. Bates, “A high-accuracy map of global terrain elevations,” Geophysical Research Letters, vol. 44, no. 11, pp. 5844–5853, 2017. 2, 3
2017
-
[27]
Delivering arbitrary-modal semantic segmentation,
J. Zhang, R. Liu, H. Shi, K. Yang, S. Reiß, K. Peng, H. Fu, K. Wang, and R. Stiefelhagen, “Delivering arbitrary-modal semantic segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 1136–1147. 2
2023
-
[28]
Side adapter network for open-vocabulary semantic segmentation,
M. Xu, Z. Zhang, F. Wei, H. Hu, and X. Bai, “Side adapter network for open-vocabulary semantic segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 2945–2954. 2
2023
-
[29]
Surface water mapping by deep learning,
F. Isikdogan, A. C. Bovik, and P. Passalacqua, “Surface water mapping by deep learning,” IEEE journal of selected topics in applied earth observations and remote sensing , vol. 10, no. 11, pp. 4909–4918, 2017. 2, 6
2017
-
[30]
Cloud and cloud shadow detection in landsat imagery based on deep convolutional neural networks,
D. Chai, S. Newsam, H. K. Zhang, Y . Qiu, and J. Huang, “Cloud and cloud shadow detection in landsat imagery based on deep convolutional neural networks,” Remote sensing of environment, vol. 225, pp. 307–316,
-
[31]
Attention is all you need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017. 2, 5
2017
-
[32]
Encoder- decoder with atrous separable convolution for semantic image segmen- tation,
L.-C. Chen, Y . Zhu, G. Papandreou, F. Schroff, and H. Adam, “Encoder- decoder with atrous separable convolution for semantic image segmen- tation,” in Proceedings of the European conference on computer vision (ECCV), 2018, pp. 801–818. 2, 5
2018
-
[33]
Searching for mobilenetv3,
A. Howard, M. Sandler, G. Chu, L.-C. Chen, B. Chen, M. Tan, W. Wang, Y . Zhu, R. Pang, V . Vasudevan et al. , “Searching for mobilenetv3,” in Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 1314–1324. 2, 5
2019
-
[34]
Segnet: A deep con- volutional encoder-decoder architecture for image segmentation,
V . Badrinarayanan, A. Kendall, and R. Cipolla, “Segnet: A deep con- volutional encoder-decoder architecture for image segmentation,” IEEE transactions on pattern analysis and machine intelligence , vol. 39, no. 12, pp. 2481–2495, 2017. 2, 5
2017
-
[35]
An image is worth 16x16 words: Transformers for image recognition at scale,
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” in International Conference on Learning Representations , 2021. 2, 5
2021
-
[36]
Swin transformer: Hierarchical vision transformer using shifted windows,
Z. Liu, Y . Lin, Y . Cao, H. Hu, Y . Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 10 012–10 022. 2, 5
2021
-
[37]
Revisiting unreasonable effectiveness of data in deep learning era,
C. Sun, A. Shrivastava, S. Singh, and A. Gupta, “Revisiting unreasonable effectiveness of data in deep learning era,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 843–852. 2
2017
-
[38]
Laion- 5b: An open large-scale dataset for training next generation image-text models,
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsmanet al., “Laion- 5b: An open large-scale dataset for training next generation image-text models,” Advances in Neural Information Processing Systems , vol. 35, pp....
2022
-
[39]
Imagenet: A large-scale hierarchical image database,
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in 2009 IEEE conference on computer vision and pattern recognition . Ieee, 2009, pp. 248–255. 2, 5
2009
-
[40]
Fast r-cnn,
R. Girshick, “Fast r-cnn,” in Proceedings of the IEEE international conference on computer vision , 2015, pp. 1440–1448. 2, 5
2015
-
[41]
Fully convolutional networks for semantic segmentation,
J. Long, E. Shelhamer, and T. Darrell, “Fully convolutional networks for semantic segmentation,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 3431–3440. 2, 5
2015
-
[42]
Mask r-cnn,
K. He, G. Gkioxari, P. Doll ´ar, and R. Girshick, “Mask r-cnn,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 2961–2969. 2, 5
2017
-
[43]
Foundation Models for Generalist Geospatial Artificial Intelligence,
J. Jakubik, S. Roy, C. E. Phillips, P. Fraccaro, D. Godwin, B. Zadrozny, D. Szwarcman, C. Gomes, G. Nyirjesy, B. Edwards, D. Kimura, N. Si- mumba, L. Chu, S. K. Mukkavilli, D. Lambhate, K. Das, R. Bangalore, D. Oliveira, M. Muszynski, K. Ankur, M. Ramasubramanian, I. Gurung, S...
-
[44]
Satlaspretrain: A large-scale dataset for remote sensing image under- standing,
F. Bastani, P. Wolters, R. Gupta, J. Ferdinando, and A. Kembhavi, “Satlaspretrain: A large-scale dataset for remote sensing image under- standing,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 16 772–16 782. 2, 5
2023
-
[45]
Improved automated detection of subpixel-scale inunda- tion—revised dynamic surface water extent (dswe) partial surface water tests,
J. W. Jones, “Improved automated detection of subpixel-scale inunda- tion—revised dynamic surface water extent (dswe) partial surface water tests,” Remote Sensing, vol. 11, no. 4, p. 374, 2019. 2, 16
2019
-
[46]
The harmonized landsat and sentinel-2 surface reflectance data set,
M. Claverie, J. Ju, J. G. Masek, J. L. Dungan, E. F. Vermote, J.-C. Roger, S. V . Skakun, and C. Justice, “The harmonized landsat and sentinel-2 surface reflectance data set,” Remote sensing of environment , vol. 219, pp. 145–161, 2018. 2
2018
-
[47]
A first look at the opera surface water extent and land surface disturbance products and their applications,
M. G. Bato, K. Devlin, R. Dhillon, M. Bonnema, S. Sangha, S. Niemoeller, A. Pickens, G. H. Shiroma, A. L. Handwerger, J. W. Jones et al. , “A first look at the opera surface water extent and land surface disturbance products and their applications,” in EGU General Assembly Con...
2023
-
[48]
Multitask learning,
R. Caruana, “Multitask learning,” Machine learning, vol. 28, pp. 41–75,
-
[49]
Learning to multitask,
Y . Zhang, Y . Wei, and Q. Yang, “Learning to multitask,” Advances in Neural Information Processing Systems , vol. 31, 2018. 2
2018
-
[50]
Seeing through the clouds with deepwatermap,
L. F. Isikdogan, A. Bovik, and P. Passalacqua, “Seeing through the clouds with deepwatermap,” IEEE Geoscience and Remote Sensing Letters, vol. 17, no. 10, pp. 1662–1666, 2019. 2, 6, 8, 17
2019
-
[51]
Suspended sediment concentration database, hydroshare,
E. Dethier, “Suspended sediment concentration database, hydroshare,” http://www.hydroshare.org/resource/ 2ee7d421618a4873b9906540d047ced4, 2019. 3
2019
-
[52]
Water quality at the global scale: Gemstat database and information system,
C. F ¨arber, D. Lisniak, P. Saile, S.-H. Kleber, M. Ehl, S. Dietrich, M. Fader, and S. Demuth, “Water quality at the global scale: Gemstat database and information system,” in EGU General Assembly Confer- ence Abstracts, 2018, p. 15984. 3 20
2018
-
[53]
A brief overview of the global river chemistry database, glorich,
J. Hartmann, R. Lauerwald, and N. Moosdorf, “A brief overview of the global river chemistry database, glorich,” Procedia Earth and Planetary Science, vol. 10, pp. 23–27, 2014. 3
2014
-
[54]
Modeling suspended sediment concentra- tion using artificial neural networks, an effort towards global sediment flux observations in rivers from space,
L. V . Lucchese, R. Daroya, T. Simmons, P. Prum, S. Maji, T. Pavelsky, C. Gleason, and J. Gardner, “Modeling suspended sediment concentra- tion using artificial neural networks, an effort towards global sediment flux observations in rivers from space,” Copernicus Meetings, Tec...
2024
-
[55]
Gradient-based learning applied to document recognition,
Y . LeCun, L. Bottou, Y . Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE , vol. 86, no. 11, pp. 2278–2324, 1998. 5
1998
-
[56]
Semantic scene segmen- tation in unstructured environment with modified deeplabv3+,
B. Baheti, S. Innani, S. Gajre, and S. Talbar, “Semantic scene segmen- tation in unstructured environment with modified deeplabv3+,” Pattern Recognition Letters, vol. 138, pp. 223–229, 2020. 5
2020
-
[57]
Simple convolutional neural network on image classification,
T. Guo, J. Dong, H. Li, and Y . Gao, “Simple convolutional neural network on image classification,” in 2017 IEEE 2nd International Conference on Big Data Analysis (ICBDA) . IEEE, 2017, pp. 721–724. 5
2017
-
[58]
Deformable detr: Deformable transformers for end-to-end object detection,
X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai, “Deformable detr: Deformable transformers for end-to-end object detection,”arXiv preprint arXiv:2010.04159, 2020. 5
2010 arXiv
-
[59]
Rectified linear units improve restricted boltz- mann machines,
V . Nair and G. E. Hinton, “Rectified linear units improve restricted boltz- mann machines,” in Proceedings of the 27th international conference on machine learning (ICML-10) , 2010, pp. 807–814. 5
2010
-
[60]
Rectifier nonlinearities improve neural network acoustic models,
A. L. Maas, A. Y . Hannun, A. Y . Ng et al. , “Rectifier nonlinearities improve neural network acoustic models,” in Proc. icml, vol. 30, no. 1. Atlanta, GA, 2013, p. 3. 5
2013
-
[61]
U-net: Convolutional networks for biomedical image segmentation,
O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in Medical image computing and computer-assisted intervention–MICCAI 2015: 18th international con- ference, Munich, Germany, October 5-9, 2015, proceedings, part III 18 ...
2015
-
[62]
Deep residual learning for image recognition,
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778. 5
2016
-
[63]
Neural machine translation by jointly learning to align and translate,
D. Bahdanau, K. Cho, and Y . Bengio, “Neural machine translation by jointly learning to align and translate,” arXiv preprint arXiv:1409.0473,
-
[64]
Decaf: A deep convolutional activation feature for generic visual recognition,
J. Donahue, Y . Jia, O. Vinyals, J. Hoffman, N. Zhang, E. Tzeng, and T. Darrell, “Decaf: A deep convolutional activation feature for generic visual recognition,” in International conference on machine learning . PMLR, 2014, pp. 647–655. 6
2014
-
[65]
Instance-aware semantic segmentation via multi-task network cascades,
J. Dai, K. He, and J. Sun, “Instance-aware semantic segmentation via multi-task network cascades,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 3150–3158. 6
2016
-
[66]
Rich feature hierarchies for accurate object detection and semantic segmentation,
R. Girshick, J. Donahue, T. Darrell, and J. Malik, “Rich feature hierarchies for accurate object detection and semantic segmentation,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2014, pp. 580–587. 6
2014
-
[67]
Multi-sensor cloud and cloud shadow segmentation with a convolutional neural network,
M. Wieland, Y . Li, and S. Martinis, “Multi-sensor cloud and cloud shadow segmentation with a convolutional neural network,” Remote Sensing of Environment, vol. 230, p. 111203, 2019. 6, 12
2019
-
[68]
A threshold selection method from gray-level his- tograms,
N. Otsu et al. , “A threshold selection method from gray-level his- tograms,” Automatica, vol. 11, no. 285-296, pp. 23–27, 1975. 6
1975
-
[69]
Sunmask algorithm,
J. Soimasuo, H. Cho, and M. Neteler, “Sunmask algorithm,” https:// grass.osgeo.org/grass83/manuals/r.sunmask.html, 2001. 7
2001
-
[70]
The astronomical almanac’s algorithm for approximate solar position (1950–2050),
J. J. Michalsky, “The astronomical almanac’s algorithm for approximate solar position (1950–2050),” Solar energy, vol. 40, no. 3, pp. 227–235,
1950
-
[71]
Fourier series representation of the position of the sun
J. Spencer, “Fourier series representation of the position of the sun.” Search, vol. 2, no. 5, p. 172, 1971. 7
1971
-
[72]
Sun-pointing programs and their accuracy,
J. C. Zimmerman, “Sun-pointing programs and their accuracy,” Sandia National Lab.(SNL-NM), Albuquerque, NM (United States), Tech. Rep.,
-
[73]
An introduction to computational geometry,
M. Minsky and S. Papert, “An introduction to computational geometry,” Cambridge tiass., HIT , vol. 479, no. 480, p. 104, 1969. 8
1969
-
[74]
An all-scale feature fusion network with boundary point prediction for cloud detection,
W. Wang and Z. Shi, “An all-scale feature fusion network with boundary point prediction for cloud detection,” IEEE Geoscience and Remote Sensing Letters, vol. 19, pp. 1–5, 2021. 11
2021
-
[75]
Cdunet: Cloud detection unet for remote sensing imagery,
K. Hu, D. Zhang, and M. Xia, “Cdunet: Cloud detection unet for remote sensing imagery,” Remote Sensing, vol. 13, no. 22, p. 4533, 2021. 11
2021
-
[76]
Cloud detection in optical remote sensing images with deep semi-supervised and active learning,
X. Yao, Q. Guo, and A. Li, “Cloud detection in optical remote sensing images with deep semi-supervised and active learning,” IEEE Geoscience and Remote Sensing Letters , vol. 20, pp. 1–5, 2023. 11
2023
-
[77]
Confluence: An open-source framework for swot discharge estimation and value- added products,
N. Tebaldi, C. Gleason, M. Durand, S. Coss, K. Andreadis, R. Dudley, D. Bjerklie, K. Larnier, H. Oubanas, P.-O. Malaterre et al., “Confluence: An open-source framework for swot discharge estimation and value- added products,” in AGU Fall Meeting Abstracts , vol. 2021, 2021, pp...
2021
-
[78]
The surface water and ocean topography (swot) mission river database (sword): A global river network for satellite data products,
E. H. Altenau, T. M. Pavelsky, M. T. Durand, X. Yang, R. P. d. M. Frasson, and L. Bendezu, “The surface water and ocean topography (swot) mission river database (sword): A global river network for satellite data products,” Water Resources Research , vol. 57, no. 7, p. e2021WR0...
2018
-
[2018]
He serves on the editorial board of the International Journal of Computer Vision (IJCV)
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.