REVIEW 5 major objections 8 minor 1 cited by
SeaMo: A Season-Aware Multimodal Foundation Model for Remote Sensing
T0 review · 5 major / 8 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read SeaMo fuses four seasons and SAR-optical data to top benchmarks
desk verdict SeaMo's season-aware pretraining is a plausible, well-benchmarked recipe, but the core ablation conflates data volume with temporal modeling, so the headline claim isn't yet supported. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is the Temporal-Multimodal (TM) fusion block, a cascade of cross-attention layers that runs after the encoder: for each season t, the optical tokens act as queries over SAR tokens of the same season plus tokens from season t−1, and vice versa, so temporal and modal information are fused in one pass. Around it sits the partial-overlap geospatial region selection, which crops each seasonal image from a slightly different offset so that regions overlap only partially, and the progressive pretraining strategy that first trains the encoder on a single season and then unfreezes temporal fusion across all four seasons. The masked autoencoder objective remains the same throughout: reconstruct masked patches in each modality, with the loss summed over seasons and modalities. What carries the argument is that these three choices are varied in ablations, and each is shown to improve downstream performance over its obvious alternative.
What would settle it
Pretrain SeaMo exactly as described but replace each location's four seasonal images with four images sampled from four different geographic locations, so no two crops share a ground footprint; if downstream performance stays roughly equal, the seasonal-consistency assumption is not what drives the gains.
Extended reading notes
Core claim
The central claim is that SeaMo's pretraining design—jointly encoding optical and SAR tokens, cropping partially overlapping regions across the four seasonal snapshots, and passing the visible tokens through temporal-multimodal fusion blocks before reconstruction—significantly improves the quality of the learned encoder compared with unimodal, single-season, or non-fusing alternatives. On fine-tuning, SeaMo is reported as state of the art or near it on EuroSAT (99.37% accuracy), fMoW-S2 (58.25% with 10% data), BigEarthNet optical (88.54 mAP), DFC2020 segmentation (49.79 mIoU), SegMunich (51.3 mIoU), OSCD change detection (54.54 F1), EuroSAT-SAR (89.69%), BigEarthNet-SAR (82.23 mAP), and DFC2020-SAR (49.54 mIoU). The paper attributes the gains to two mechanisms: the partial-overlap cropping forces the model to find correlations between regions that do not overlap across seasons, and the TM block lets each modality at each time point attend to the other modality and to the previous season, so the reconstruction task cannot be solved by copying a single aligned image. The progressive schedule—first uni-season multimodal learning, then multi-season temporal fusion—is presented as the stabilizer that makes the harder temporal task trainable.
Load-bearing premise
The paper assumes that images of one location taken in different seasons show enough stable geological structure that forcing the model to relate partially overlapping crops teaches time-invariant features rather than injecting seasonal noise; if transient surface changes dominate, the pretraining signal weakens.
Editorial extensions
If this is right
- Adding a temporal dimension to masked autoencoder pretraining on remote sensing data improves downstream performance; the ablation shows longer temporal sequences (up to four seasons) monotonically help.
- Multimodal pretraining transfers to unimodal downstream tasks: SeaMo outperforms SAR-only baselines on EuroSAT-SAR and BigEarthNet-SAR, and optical-only baselines on optical tasks.
- Partially overlapping crops across seasons beat both identical-region crops, which make the reconstruction too easy, and fully random non-overlapping crops.
- The TM block's fused design is preferred over a decoupled design because it matches or exceeds it while using fewer cross-attention layers and parameters.
- The model generalizes to unseen image sizes via positional embedding interpolation, but a reduction in spectral bands (from 12 to 9 or 4) hurts accuracy, signalling a remaining tokenizer limitation.
Reading between the lines
- If the seasonal-consistency mechanism is real, SeaMo's gains should grow when pretraining data cover regions with strong but regular seasonal cycles such as agriculture, and shrink when the four snapshots capture transient surface change such as snow or flood; this could be tested by splitting the pretraining set by biome.
- A reader could reinterpret the partial-overlap cropping as a cheap form of multi-view consistency that does not need true geology: any set of partially overlapping crops from the same location, even same-season repeats, might give similar gains; the paper does not run that control.
- The TM block is a general spatiotemporal fusion primitive; it could be lifted into video masked modeling or multi-sensor time-series models beyond optical/SAR pairs, though the paper only evaluates it in the remote sensing setting.
- The reported gains at 10% fine-tuning data suggest the encoder learns transferable features; an obvious next stress test is zero-shot or linear-probe evaluation on unseen sensors and resolutions, which the paper only partially covers.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. SeaMo is a masked-autoencoder visual foundation model pretrained on SSL4EO-S12, jointly modeling four-season Sentinel-2 optical and Sentinel-1 SAR inputs. The paper's contributions are (i) partially overlapping spatial region selection across seasons, (ii) a Temporal-Multimodal (TM) fusion block that propagates cross-attention information across seasons and modalities, and (iii) a progressive two-phase pretraining strategy (single-time multimodal learning followed by multi-time fusion). The model is evaluated by fine-tuning and linear probing on EuroSAT, fMoW-S2, BigEarthNet(-SAR), DFC2020(-SAR), SegMunich, OSCD, GLH-Water, and S1S2-Water, where it reports the best or second-best result in most tables.
Significance. If the claimed gains are reproducible, SeaMo offers a practical pretraining recipe for season-aware multimodal remote sensing foundation models. The paper's strengths include broad benchmark coverage across optical and SAR tasks, systematic ablations of the cropping strategy, TM-block design, temporal length, pretraining strategy, decoder depth, mask ratio, reconstruction target, and positional embedding; explicit reporting of computational cost; and several honest limitation statements, such as the OSCD result being acknowledged as not state-of-the-art and the pretraining-strategy ablation being acknowledged as difficult to interpret. The central architectural idea, the TM block, is clearly described and directly ablated. The main weakness is that the key causal claim—that explicit seasonal fusion "significantly enhances" performance—rests on single-run margins that are often smaller than typical fine-tuning noise, so the statistical significance and protocol-matched validity of the gains are not yet established.
major comments (5)
- [§4.5.3, Table 14] The temporal-length ablation in Table 14 conflates the amount of pretraining data with the presence of temporal structure. The T=1 row uses only the first season, which is 25% of SSL4EO-S12, and excludes the TM block, while the T=4 row uses all four seasons, the TM block, and the progressive two-phase schedule. The comparison therefore changes the number of distinct training images, the total number of training epochs, the architecture, and the training schedule simultaneously. A masked autoencoder trained on four times as many distinct images would be expected to improve even without any season-aware mechanism, so the inference from this table to "explicit seasonal modeling helps" is not supported. The paper should provide a data-matched and schedule-matched control, such as the Multimodal-Temporal row in Table 12 run on the same total epochs, or an epoch-scaled T=1 condition.
- [§4.2, §4.5.2, Tables 1-14] No error bars, confidence intervals, or multiple-seed results appear anywhere in the paper. Several of the margins that carry the central claim are very small: in Table 11 the default Fuse design is 99.37 vs 99.41 for Decouple on EuroSAT and 82.23 vs 82.30 on BEN-SAR; in Table 14 the T=3-to-T=4 improvement is 0.23 mAP on BEN and 0.07 mAP on BEN-SAR; in Table 1 the EuroSAT advantage over DOFA is 0.07 points. The manuscript itself states in Section 4.5.4 that "It is challenging to draw a definitive conclusion from these experiments," which is in tension with the abstract's and Section 3.3.2's claim that seasonal fusion "significantly enhance[s]" performance. At minimum, the central ablations and the comparisons against the closest baselines should be repeated with at least three seeds and reported as mean ± std.
- [§4.5.2, §4.5.4, Tables 11-12] The chosen default configuration is not the best on all tasks. In Table 11, Decouple outperforms the adopted Fuse design on EuroSAT (99.41 vs 99.37) and BEN-SAR (82.30 vs 82.23), and Table 12 shows that Siamese-Temporal beats the final Multimodal-Temporal-TM on BEN-SAR (+0.45 mAP). The paper acknowledges the latter but chooses Fuse on computational grounds. That is a legitimate efficiency trade-off, but it should be presented as such, not as evidence that the proposed Fuse design is the accuracy-maximizing instantiation of the TM block. Reporting both variants with variance would clarify whether the differences are meaningful.
- [§4.2, Tables 1-8] The benchmark comparison protocol is underspecified. The tables appear to mix numbers reproduced by the authors with numbers quoted from prior publications (e.g., SatMAE is described in the text as pretrained on fMoW-Sentinel, while other baselines come from different pretraining corpora), and the table notes only partially clarify the 10% fine-tuning rule. Since several reported advantages are below 0.1 point (Table 1 EuroSAT: 99.37 vs 99.30; Table 5 OSCD: 54.54 vs 54.29), small differences in downstream protocol could change the ranking. For each entry, the paper should state whether the result was obtained under the authors' own fine-tuning pipeline or quoted from the original paper, and provide the corresponding protocol details.
- [§3.2, Table 10] The design rationale for partial-overlap cropping depends on the Section 3.2 assumption that "the geological attributes of a given area tend to remain consistent over time." When seasonal imagery is dominated by transient surface changes, the cross-season reconstruction objective could inject noise rather than stable structure. Table 10 provides empirical support for partial overlap on the tested tasks, but it does not distinguish between learning stable temporal correlations and simply benefiting from a harder multi-view reconstruction task. A concrete control would be to shuffle the seasonal order or to pair crops from different locations during temporal fusion; if the benefit persists, the "season-aware" interpretation would need to be revised.
minor comments (8)
- [Table 10] The column header "OCSD" should be "OSCD."
- [§4.2.1] The sentence about SatMAE being "pretrained and then fine-tuned on the fMoW-Sentinel" is confusing; please clarify which pretraining corpus and which fine-tuning split are used for each baseline in Tables 1 and 2.
- [§3.2] The phrase "geological attributes" is too narrow; land-cover or surface attributes would better describe the properties that are expected to remain consistent across seasons.
- [Algorithm 1] The fusion operator f[·] is used in the pseudocode before it is defined; please define it in the text and use a consistent notation for the concatenation inside the brackets.
- [References] References [53] and [55] are the same paper and should be consolidated.
- [Figure 15] Please add axis labels and specify whether the x-axis counts epochs of the second pretraining phase or total pretraining epochs, and state the exact epoch values at each plotted point.
- [Table 9] The baselines in Table 9 (UNet, ResNet, ViT) are not foundation models; a sentence in the caption or text should make this explicit so the comparison is not overinterpreted.
- [Data availability] The Data Availability section lists only the datasets; the authors should consider releasing code and pretrained weights to support reproducibility.
Circularity Check
No significant circularity: SeaMo's season-aware fusion claim rests on benchmark comparisons and ablations, not on self-referential derivation or fitted predictions.
full rationale
The paper's contribution is an architecture and pretraining recipe evaluated on downstream benchmarks, and no load-bearing step reduces to its own inputs by construction. The TM block is tested by ablations on held-out tasks (Tables 10-12), the cropping strategy is ablated in Table 10, and the pretraining objective is a standard masked reconstruction MSE over visible and masked patches, with no fitted parameter relabeled as a prediction. Self-citations to SpectralGPT and DOFA appear only as published baselines or evaluation protocols, not as premises that force the central claim. The temporal-length ablation in Table 14 is confounded with data quantity and training schedule, but that is a validity concern rather than a circular reduction: the conclusion does not hold by definition or by self-citation. The structural assumption that geological attributes remain consistent over time motivates the partial-overlap cropping strategy but is not equivalent to the downstream performance claims. Because the paper is self-contained against external benchmarks and its ablations are the actual evidence for the season-aware mechanism, no circular step with a quotable reduction can be exhibited.
Assumptions & free parameters
free parameters (5)
- Mask ratio =
75%
- Partial overlap crop rate =
stated default [75%,100%], empirically best [50%,100%]
- Decoder depth =
4 blocks (default), 6 blocks best
- Temporal length =
4 seasons
- Progressive schedule =
20 epochs single-time + 200 epochs multi-time
assumptions (4)
- domain assumption Geological attributes of a given area remain consistent over time, aside from human-induced changes.
- domain assumption SSL4EO-S12's four seasonal snapshots and Sentinel-1/Sentinel-2 pairs are adequately aligned in time and geography.
- standard math MAE-style reconstruction of normalized masked patches is a valid proxy for learning transferable features in multimodal, multi-seasonal RS data.
- domain assumption A single tokenizer per modality can adequately encode 12-band optical and 2-band SAR channels for the fusion encoder.
invented entities (1)
-
Temporal-Multimodal (TM) fusion block
Cite this review
Pith. "Pith review of SeaMo: A Season-Aware Multimodal Foundation Model for Remote Sensing." pith.science (2026). https://pith.science/paper/SQTQHFCZ
@misc{pith2026241219237,
author = {Pith},
title = {Pith review of: SeaMo: A Season-Aware Multimodal Foundation Model for Remote Sensing},
year = {2026},
howpublished = {\url{https://pith.science/paper/SQTQHFCZ}},
note = {Machine review of arXiv:2412.19237}
}
read the original abstract
Remote Sensing (RS) data encapsulates rich multi-dimensional information essential for Earth observation. Its vast volume, diverse sources, and temporal continuity make it particularly well-suited for developing large Visual Foundation Models (VFMs). These models serve as powerful feature extractors, leveraging extensive RS data for pretraining and subsequent fine-tuning in various geoscientific applications. However, existing VFMs in the RS domain often concentrate on specific image characteristics, neglecting the full season-aware potential of RS data. To bridge this gap, we introduce SeaMo, a novel VFM that effectively integrates multimodal and multi-seasonal RS information. SeaMo leverages a masked image modeling framework to fully exploit the spatial, spectral, and seasonal dimensions of RS data. Specifically, we employ unaligned spatial region selection to capture spatial heterogeneity, incorporate multi-source inputs for enhanced multimodal integration, and introduce temporal-multimodal fusion blocks to assimilate seasonal variations effectively. By explicitly modeling the complex, season-dependent attributes of RS data, SeaMo enhances generalization, robustness, and adaptability across geoscientific tasks. Extensive experiments and ablation studies demonstrate its superior performance, underscoring its potential as a foundational model for Earth observation.
Figures
Figures from the paper (12 more)
Forward citations
Cited by 1 Pith paper
-
Deploying Geospatial Foundation Models in the Real World: Lessons from WorldCereal
A structured protocol for deploying geospatial foundation models is introduced and validated in WorldCereal, where fine-tuned Presto outperforms a fully-supervised CatBoost baseline in crop mapping.
Reference graph
Works this paper leans on
-
[1]
G. Vivone, L.-J. Deng, S. Deng, D. Hong, M. Jiang, C. Li, W. Li, H. Shen, X. Wu, J.-L. Xiao, J. Yao, M. Zhang, J. Chanussot, S. Garc ´ ıa, A. Plaza, Deep learning in remote sensing image fusion: Methods, protocols, data, and future perspectives, IEEE Geoscience and Remote Sensing Magazine 13 (1) (2025) 269–
work page 2025
-
[2]
D. Hong, B. Zhang, X. Li, Y. Li, C. Li, J. Yao, N. Yokoya, H. Li, P. Ghamisi, X. Jia, A. Plaza, P. Gamba, J. A. Benediktsson, J. Chanussot, SpectralGPT: Spectral remote sensing foundation model, IEEE Transactions on Pattern Analysis and Machine Intelligence 46 (8) (2024) 5227–5244. doi:10.1109/TPAMI.2024. 3362475. 24
-
[3]
C. Li, B. Zhang, D. Hong, J. Zhou, G. Vivone, S. Li, J. Chanus- sot, Casformer: Cascaded transformers for fusion-aware compu- tational hyperspectral imaging, Inf. Fusion 108 (2024) 102408. doi:10.1016/J.INFFUS.2024.102408
arXiv 2024
-
[4]
D. Hong, C. Li, B. Zhang, N. Yokoya, J. A. Benediktsson, J. Chanussot, Multimodal artificial intelligence foundation mod- els: Unleashing the power of remote sensing big data in earth observation, The Innovation Geoscience 2 (1) (2024) 100055. doi:10.59717/j.xinn-geo.2024.100055
arXiv 2024
- [5]
-
[6]
R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brun- skill, et al., On the opportunities and risks of foundation models (2022). arXiv:2108.07258
arXiv 2022
-
[7]
X. Wanyan, S. Seneviratne, S. Shen, M. Kirley, Extending global-local view alignment for self-supervised learning with re- mote sensing imagery, in: 2024 IEEE/CVF Conference on Com- puter Vision and Pattern Recognition Workshops (CVPR W), 2024, pp. 2443–2453. doi:10.1109/CVPRW63382.2024.00251
arXiv 2024
- [8]
Show all 57 references
-
[12]
X. Sun, P. Wang, W. Lu, Z. Zhu, X. Lu, Q. He, J. Li, X. Rong, Z. Yang, H. Chang, Q. He, G. Yang, R. Wang, J. Lu, K. Fu, Ringmo: A remote sensing foundation model with masked im- age modeling, IEEE Transactions on Geoscience and Remote Sensing 61 (2023) 1–22. doi:10.1109/TGRS.2...
2023
-
[13]
Y. Cong, S. Khanna, C. Meng, P. Liu, E. Rozi, Y. He, M. Burke, D. Lobell, S. Ermon, SatMAE: Pre-training transformers for temporal and multi-spectral satellite imagery, in: Advances in Neural Information Processing Systems, Vol. 35, Curran Asso- ciates, Inc., 2022, pp. 197–211
2022
-
[14]
K. He, X. Chen, S. Xie, Y. Li, P. Doll´ ar, R. Girshick, Masked autoencoders are scalable vision learners, in: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 15979–15988. doi:10.1109/CVPR52688. 2022.01553
2022
-
[17]
Fuller, K
A. Fuller, K. Millard, J. Green, CROMA: Remote sensing representations with contrastive radar-optical masked autoen- coders, in: Advances in Neural Information Processing Systems, Vol. 36, Curran Associates, Inc., 2023, pp. 5506–5538
2023
-
[18]
X. Guo, J. Lao, B. Dang, Y. Zhang, L. Yu, L. Ru, L. Zhong, Z. Huang, K. Wu, D. Hu, H. He, J. Wang, J. Chen, M. Yang, Y. Zhang, Y. Li, SkySense: A multi-modal remote sensing foun- dation model towards universal interpretation for earth observa- tion imagery, in: 2024 IEEE/CVF C...
2024
-
[19]
Bastani, P
F. Bastani, P. Wolters, R. Gupta, J. Ferdinando, A. Kembhavi, Satlaspretrain: A large-scale dataset for remote sensing image understanding, in: 2023 IEEE/CVF International Conference on Computer Vision (ICCV), 2023, pp. 16726–16736. doi:10. 1109/ICCV51070.2023.01538
2023
-
[20]
Srivastava, E
N. Srivastava, E. Mansimov, R. Salakhutdinov, Unsupervised learning of video representations using lstms, in: Proceedings of the 32nd International Conference on International Conference on Machine Learning - Volume 37, ICML’15, JMLR.org, 2015, p. 843–852
2015
-
[21]
Vondrick, H
C. Vondrick, H. Pirsiavash, A. Torralba, Anticipating visual representations from unlabeled video, in: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 98–106. doi:10.1109/CVPR.2016.18
2016 doi
-
[22]
T. Chen, S. Kornblith, M. Norouzi, G. Hinton, A simple framework for contrastive learning of visual representations, in: Proceedings of the 37th International Conference on Ma- chine Learning, Vol. 119 of Proceedings of Machine Learning 25 Research, PMLR, 2020, pp. 1597–1607
2020
-
[23]
K. He, H. Fan, Y. Wu, S. Xie, R. Girshick, Momentum con- trast for unsupervised visual representation learning, in: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recog- nition (CVPR), 2020, pp. 9726–9735. doi:10.1109/CVPR42600. 2020.00975
2020
-
[24]
Z. Xie, Z. Zhang, Y. Cao, Y. Lin, J. Bao, Z. Yao, Q. Dai, H. Hu, SimMIM: a simple framework for masked image mod- eling, in: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 9643–9653. doi: 10.1109/CVPR52688.2022.00943
2022
-
[26]
Bachmann, D
R. Bachmann, D. Mizrahi, A. Atanov, A. Zamir, MultiMAE: Multi-modal multi-task masked autoencoders, in: Computer Vision – ECCV 2022, Springer Nature Switzerland, Cham, 2022, pp. 348–367. doi:10.1007/978-3-031-19836-6_20
2022 doi
-
[27]
Feichtenhofer, H
C. Feichtenhofer, H. Fan, Y. Li, K. He, Masked autoencoders as spatiotemporal learners, in: Proceedings of the 36th Inter- national Conference on Neural Information Processing Systems (NeurIPS 2022), NIPS ’22, Curran Associates Inc., Red Hook, NY, USA, 2022, pp. 35946–35958
2022
-
[28]
L. Zhou, H. Liu, J. Bae, J. He, D. Samaras, P. Prasanna, Self pre-training with masked autoencoders for medical image classification and segmentation, in: 2023 IEEE 20th Interna- tional Symposium on Biomedical Imaging (ISBI), 2023, pp. 1–6. doi:10.1109/ISBI53787.2023.10230477
2023
-
[29]
Nguyen, J
T. Nguyen, J. Brandstetter, A. Kapoor, J. K. Gupta, A. Grover, ClimaX: a foundation model for weather and climate, in: Proceedings of the 40th International Conference on Machine Learning, ICML’23, JMLR.org, 2023, pp. 25904–25938
2023
-
[30]
Mendieta, B
M. Mendieta, B. Han, X. Shi, Y. Zhu, C. Chen, Towards geospatial foundation models via continual pretraining, in: 2023 IEEE/CVF International Conference on Computer Vi- sion (ICCV), 2023, pp. 16760–16770. doi:10.1109/ICCV51070. 2023.01541
2023
-
[31]
M. Tang, A. Cozma, K. Georgiou, H. Qi, Cross-scale mae: a tale of multi-scale exploitation in remote sensing, in: Proceedings of the 37th International Conference on Neural Information Pro- cessing Systems, NIPS ’23, Curran Associates Inc., Red Hook, NY, USA, 2023, pp. 20054–20066
2023
-
[32]
Ayush, B
K. Ayush, B. Uzkent, C. Meng, K. Tanmay, M. Burke, D. Lo- bell, S. Ermon, Geography-aware self-supervised learning, in: 2021 IEEE/CVF International Conference on Computer Vi- sion (ICCV), 2021, pp. 10161–10170. doi:10.1109/ICCV48922. 2021.01002
2021
-
[33]
Y. Wang, C. M. Albrecht, X. X. Zhu, Multilabel-guided soft contrastive learning for efficient earth observation pretrain- ing, IEEE Transactions on Geoscience and Remote Sensing 62 (2024) 1–16. doi:10.1109/TGRS.2024.3466896
2024
-
[34]
F. Liu, D. Chen, Z. Guan, X. Zhou, J. Zhu, Q. Ye, L. Fu, J. Zhou, Remoteclip: A vision language foundation model for remote sensing, IEEE Transactions on Geoscience and Remote Sensing 62 (2024) 1–16. doi:10.1109/TGRS.2024.3390838
2024
-
[35]
V. V. Cepeda, G. K. Nayak, M. Shah, GeoCLIP: clip-inspired alignment between locations and images for effective worldwide geo-localization, in: Proceedings of the 37th International Con- ference on Neural Information Processing Systems, NIPS ’23, Curran Associates Inc., Red Ho...
2023
-
[36]
X. Li, C. Wen, Y. Hu, Z. Yuan, X. X. Zhu, Vision-language models in remote sensing: Current progress and future trends, IEEE Geoscience and Remote Sensing Magazine 12 (2) (2024) 32–66. doi:10.1109/MGRS.2024.3383473
2024
-
[37]
Y. Wang, N. A. A. Braham, Z. Xiong, C. Liu, C. M. Albrecht, X. X. Zhu, Ssl4eo-s12: A large-scale multimodal, multitempo- ral dataset for self-supervised learning in earth observation [soft- ware and data sets], IEEE Geoscience and Remote Sensing Mag- azine 11 (3) (2023) 98–106...
2023
-
[38]
Vaswani, N
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, I. Polosukhin, Attention is all you need, in: Proceedings of the 31st International Conference on Neural Information Processing Systems, NIPS’17, Curran Associates Inc., Red Hook, NY, USA, 2017,...
2017
-
[39]
Dosovitskiy, L
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, N. Houlsby, An image is worth 16x16 words: Transformers for image recognition at scale, ICLR
-
[40]
Gupta, J
A. Gupta, J. Wu, J. Deng, L. Fei-Fei, Siamese masked autoen- coders, in: Proceedings of the 37th International Conference on Neural Information Processing Systems, NIPS ’23, Curran Associates Inc., Red Hook, NY, USA, 2023, pp. 40676–40693
2023
-
[41]
Eyma¨ el, R
A. Eyma¨ el, R. Vandeghen, A. Cioppa, S. Giancola, B. Ghanem, M. Van Droogenbroeck, Efficient image pre- training with siamese cropped masked autoencoders, in: Com- puter Vision – ECCV 2024, Springer Nature Switzerland, Cham, 2025, pp. 348–366. doi:10.1007/978-3-031-73337-6_ 20
2024 doi
-
[42]
Xiong, Y
Z. Xiong, Y. Wang, F. Zhang, A. J. Stewart, J. Hanna, D. Borth, I. Papoutsis, B. L. Saux, G. Camps-Valls, X. X. Zhu, Neural plasticity-inspired multimodal foundation model for earth observation (2024). arXiv:2403.15356
2024
-
[43]
Caron, H
M. Caron, H. Touvron, I. Misra, H. Jegou, J. Mairal, P. Bo- 26 janowski, A. Joulin, Emerging properties in self-supervised vi- sion transformers, in: 2021 IEEE/CVF International Confer- ence on Computer Vision (ICCV), 2021, pp. 9630–9640. doi: 10.1109/ICCV48922.2021.00951
2021
-
[44]
Assran, Q
M. Assran, Q. Duval, I. Misra, P. Bojanowski, P. Vincent, M. Rabbat, Y. LeCun, N. Ballas, Self-supervised learning from images with a joint-embedding predictive architecture, in: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 15619–15629....
2023
-
[45]
Helber, B
P. Helber, B. Bischke, A. Dengel, D. Borth, Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification, IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing 12 (7) (2019) 2217–
2019
-
[46]
Sumbul, M
G. Sumbul, M. Charfuelan, B. Demir, V. Markl, Bigearthnet: A large-scale benchmark archive for remote sensing image un- derstanding, in: IGARSS 2019 - 2019 IEEE International Geo- science and Remote Sensing Symposium, 2019, pp. 5901–5904. doi:10.1109/IGARSS.2019.8900532
2019
-
[47]
Sumbul, A
G. Sumbul, A. de Wall, T. Kreuziger, F. Marcelino, H. Costa, P. Benevides, M. Caetano, B. Demir, V. Markl, Bigearthnet- mm: A large-scale, multimodal, multilabel benchmark archive for remote sensing image classification and retrieval [software and data sets], IEEE Geoscience a...
2021
-
[48]
Neumann, A
M. Neumann, A. S. Pinto, X. Zhai, N. Houlsby, In-domain rep- resentation learning for remote sensing, in: AI for Earth Sci- ences Workshop at International Conference on Learning Rep- resentations (ICLR), 2020, pp. 1–20
2020
-
[49]
Robinson, K
C. Robinson, K. Malkin, N. Jojic, H. Chen, R. Qin, C. Xiao, M. Schmitt, P. Ghamisi, R. H¨ ansch, N. Yokoya, Global land- cover mapping with weak supervision: Outcome of the 2020 ieee grss data fusion contest, IEEE Journal of Selected Topics in Applied Earth Observations and Re...
2021
- [50]
-
[51]
R. C. Daudt, B. Le Saux, A. Boulch, Y. Gousseau, Urban change detection for multispectral earth observation using con- volutional neural networks, in: IGARSS 2018 - 2018 IEEE In- ternational Geoscience and Remote Sensing Symposium, 2018, pp. 2115–2118. doi:10.1109/IGARSS.2018.8518015
2018
-
[52]
T. Xiao, Y. Liu, B. Zhou, Y. Jiang, J. Sun, Unified perceptual parsing for scene understanding, in: Computer Vision – ECCV 2018: 15th European Conference, Munich, Germany, September 8–14, 2018, Proceedings, Part V, Springer-Verlag, Berlin, Hei- delberg, 2018, p. 432–448. doi:1...
2018 doi
-
[54]
Y. Wang, H. H. Hern´ andez, C. M. Albrecht, X. X. Zhu, Fea- ture guided masked autoencoder for self-supervised learning in remote sensing, IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing 18 (2025) 321–336. doi:10.1109/JSTARS.2024.3493237
2025
-
[55]
Y. Wang, C. M. Albrecht, X. X. Zhu, Self-supervised vision transformers for joint sar-optical representation learning, in: IGARSS 2022 - 2022 IEEE International Geoscience and Re- mote Sensing Symposium, 2022, pp. 139–142. doi:10.1109/ IGARSS46834.2022.9883983
2022
-
[56]
Fuller, K
A. Fuller, K. Millard, J. R. Green, Satvit: Pretraining transformers for earth observation, IEEE Geoscience and Re- mote Sensing Letters 19 (2022) 1–5. doi:10.1109/LGRS.2022. 3201489
2022 doi
-
[57]
Chan-To-Hing, B
H. Chan-To-Hing, B. Veeravalli, Fus-mae: A cross-attention- based data fusion approach for masked autoencoders in remote sensing, in: IGARSS 2024 - 2024 IEEE International Geoscience and Remote Sensing Symposium, 2024, pp. 6953–6958. doi: 10.1109/IGARSS53475.2024.10642424
2024
-
[58]
Ronneberger, P
O. Ronneberger, P. Fischer, T. Brox, U-Net: Convolutional networks for biomedical image segmentation, in: N. Navab, J. Hornegger, W. M. Wells, A. F. Frangi (Eds.), Medical Im- age Computing and Computer-Assisted Intervention – MICCAI 2015, Springer International Publishing, Ch...
2015
-
[59]
K. He, X. Zhang, S. Ren, J. Sun, Deep residual learning for image recognition, in: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 770–778. doi:10.1109/CVPR.2016.90
2016 doi
-
[60]
Y. Li, B. Dang, W. Li, Y. Zhang, Glh-water: a large-scale dataset for global surface water detection in large-size very- high-resolution satellite imagery, in: Proceedings of the Thirty- Eighth AAAI Conference on Artificial Intelligence and Thirty- Sixth Conference on Innovati...
2024 doi
-
[61]
Wieland, F
M. Wieland, F. Fichtner, S. Martinis, S. Groth, C. Krullikowski, S. Plank, M. Motagh, S1s2-water: A global dataset for semantic segmentation of water bodies from sentinel- 1 and sentinel-2 27 satellite images, IEEE Journal of Selected Topics in Applied Earth Observations and R...
2024
-
[241]
doi:10.1007/978-3-319-24574-4\_28
-
[310]
doi:10.1109/MGRS.2024.3495516
2024
-
[2226]
doi:10.1109/JSTARS.2019.2918242
2019
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.