REVIEW 4 major objections 4 minor 66 references
A multispectral Earth-observation model trained by dual-teacher distillation—a contrastive multispectral teacher plus a frozen optical vision foundation model—beats much larger foundation models on optical and multispectral benchmarks using
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 21:29 UTC pith:RS6WKPYP
load-bearing objection DEO reports strong, broad multispectral EO benchmark results for a dual-teacher contrastive distillation scheme, but the paper's headline mechanism—objective-level consistency with a VFM teacher—is asserted rather than tested. the 4 major comments →
Brewing Stronger Features: Dual-Teacher Distillation for Multispectral Earth Observation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On the paper's own terms, the central discovery is that objective-level consistency—matching the student's pretraining objective to the teacher's—determines how well optical vision-foundation-model knowledge transfers to a multispectral student. Prior EO pretraining typically couples masked image modeling with VFM distillation, but because MIM optimizes local reconstruction rather than global semantic structure, the student's latent space stays misaligned with the VFM's contrastively structured feature space. The proposed model instead pairs a contrastive self-distillation multispectral teacher (global/local view alignment with a coding-rate regularizer) with a frozen optical VFM teacher tha
What carries the argument
The central mechanism is 'dual-teacher contrastive distillation': a single compact transformer student is supervised by two teachers. Teacher one is a contrastive self-distillation multispectral teacher, updated by exponential moving average of the student, which receives global views of 10-channel multispectral imagery and enforces cosine similarity with the student's global-plus-local views while a coding-rate regularizer discourages feature collapse. Teacher two is a frozen optical vision foundation model, which receives global views of the 3-channel optical bands and transfers its class token and patch tokens from final and intermediate layers into separate student projection heads on a
Load-bearing premise
The claim hinges on the assumption that objective-level consistency with the optical teacher—not the extra teacher, the extra optical data, or the larger supervision—is what drives the benchmark gains.
What would settle it
Pretrain the same student backbone on the same 0.5M-image dataset with the same two teachers, but change only the student's multispectral objective from contrastive self-distillation to masked image modeling; if downstream segmentation, change-detection, and classification scores stay at the same level, the paper's objective-alignment explanation is falsified even though the benchmark numbers may stand.
If this is right
- If the central claim is right, contrastive distillation is a data-efficient pretraining recipe for multispectral EO: state-of-the-art results come from an 87M-parameter model and 0.5M pretraining images, far fewer than several competing foundation models.
- A single dual-teacher model serves both optical-only and multispectral downstream tasks without degrading either, so one checkpoint can replace separate RGB and multispectral backbones.
- Objective compatibility with modern optical VFMs means new, stronger optical teachers can be dropped into the pipeline as they appear, letting EO models inherit general-computer-vision progress cheaply.
- Low-data fine-tuning results indicate the pretrained representations transfer well when labels are scarce, which matters for rapid disaster-response mapping.
- Distilling a patch-based optical teacher into a hierarchical transformer backbone with a small patch size produces fine-grained features for dense prediction, so pixel-level tasks benefit from global semantic priors.
Where Pith is reading between the lines
- Editorial inference: The benchmark gains might come from the additional optical teacher, the high-resolution aerial optical data substituted into pretraining, or the larger total supervision rather than from objective alignment per se; a controlled experiment that keeps backbone, data, and teachers fixed while toggling only the student's objective (contrastive vs masked image modeling) would isola
- Editorial inference: The practice of replacing low-resolution multispectral optical bands with high-resolution aerial imagery suggests that 'privileged' high-resolution data can be injected at pretraining time; a natural extension is to test the same substitution with other modalities or resolutions.
- Editorial inference: If objective compatibility is the true driver, then future VFMs trained under a different paradigm (for example, generative or JEPA-style objectives) would each need a purpose-built distillation recipe, implying there is no single universal distillation scheme for EO.
- Editorial inference: The method's reliance on spatially aligned optical and multispectral inputs, and the absence of a strong teacher for radar (SAR) data, sets a boundary on its reach; a testable extension is to add a third teacher for a non-visual modality and check whether the same compatibility argument still holds.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes DEO, a dual-teacher self-supervised pretraining method for multispectral Earth observation. A Swin-Tiny student is trained jointly with (i) a contrastive self-distillation multispectral teacher on 10-band Sentinel-2 imagery, using an EMA teacher and a coding-rate regularizer, and (ii) a frozen optical VFM teacher (DINOv3) whose class and patch tokens are distilled from RGB channels. Pretraining uses 500k images from fMoW-Sentinel with 150k RGB bands replaced by high-resolution aerial counterparts. The paper reports state-of-the-art results across semantic segmentation (GEO-Bench plus SpaceNet, Sen1Floods11, PASTIS), bi-temporal change detection (LEVIR, OSCD), linear-probe classification, and a 10% low-data regime, with an average improvement of 3.64 points in segmentation, 1.2 in change detection, and 1.31 in classification. The central causal claim is that aligning the student's contrastive objective with the contrastive self-distillation objective of the optical VFM is what drives the gains.
Significance. If the empirical claims hold, DEO provides a compelling alternative to the dominant MIM-based pretraining recipe for multispectral EOFMs. Its strengths are the breadth of evaluation (13 datasets, multiple tasks, low-data regime), the use of standard held-out splits, and the clear component-wise ablation that builds from a contrastive baseline. The result is also practically significant: the method reaches strong performance with only 0.5M pretraining images and an 87M-parameter backbone, much less than many competitors. However, the paper's most distinctive intellectual claim, that objective-level consistency with the VFM teacher is the mechanism, is not directly tested, and several evaluation details (single runs, unspecified input-channel adaptation) undermine confidence in the exact margins. These issues are fixable with additional experiments and reporting, so the paper is a worthwhile candidate after major revision.
major comments (4)
- [Section 3.4, Contribution 2, Figure 3, Table 5] The headline mechanism claim, that matching the student's contrastive objective with the VFM's contrastive self-distillation objective drives the gains, is not supported by the experiments as designed. The only cross-paradigm comparison is against Copernicus-FM, which differs simultaneously in pretraining objective (MIM vs. contrastive), backbone (ViT vs. Swin), pretraining data, and distillation losses. The internal ablation in Table 5 starts from a contrastive multispectral baseline and adds the DINOv3 teacher; it never replaces the student's objective with an MIM objective while holding backbone, data, augmentations, and distillation targets fixed. Consequently, the results are equally compatible with the alternative explanation that the gains come from the extra optical teacher's features, the high-resolution optical replacement (+0.38 overall in Table 5), or the larger effective sup
- [Tables 1-7] All results are reported from single runs with no error bars, standard deviations, or number of seeds. This is a serious limitation because many of the claimed improvements are small relative to typical fine-tuning variance. For example, in Table 5 the overall increment for adding high-resolution optical data is only 0.38 points, and several component additions are 0.15-0.27 points. In Table 1, DEO is worse than SatDiFuser on GB-cattle by 2.77 points, and in Table 3 DEO is worse than TerraFM on GB-ben by 3.72 points, even though the averages are favorable. Without multiple seeds or confidence intervals, the precise headline margins (3.64, 1.2, 1.31) cannot be distinguished from noise. Please provide at least 3-5 seeds for the main tables and for the key ablation rows, or report error bars from a bootstrap over evaluation runs.
- [Section 3.3, Tables 1 and 3, Tables 12-13] The paper does not specify how DEO, pretrained with a 10-channel patch embedding, is adapted to downstream datasets with different band counts. The downstream evaluation includes 4-band (GB-chesapeake), 12-band (m-bigearthnet), 13-band (GB-cashew, GB-SA-crop-type, Sen1Floods11, m-eurosat), and 18-band (m-so2sat) inputs. The only description is employing a 10-channel patch embedding layer in Section 3.3, and the supplementary dataset tables list band counts without any adaptation protocol. This is essential for reproducibility and for interpreting the multispectral results: if the first 10 bands are used, some Sentinel-2 bands are discarded; if weights are averaged, duplicated, or randomly initialized, the comparison changes. Please state the exact channel-matching procedure used for each dataset, and ideally ablate it.
- [Section 4.2, Supplement B.2] The change-detection evaluation is not a frozen-backbone evaluation. Section 4.2 says the protocol follows [37,45,54], and Supplement B.2 states, following related work [45,54], we also train the backbone for change detection. In addition, for ViT-based methods a UNet decoder is used, while Swin-based methods use UPerNet; this means the decoder architecture is not held fixed across methods. The paper should make this explicit in the main text and clarify that the 1.2-point improvement in Table 2 is measured under a full-fine-tuning protocol, not a representation-transfer protocol. If the authors intend Table 2 as evidence of better pretrained features, they should also report the frozen-backbone version.
minor comments (4)
- [Abstract / Section 5] The stated average improvement of 3.64 percentage points in semantic segmentation does not directly match Table 1. From Table 1, the overall-average gap to the best prior method is 3.73, the per-dataset average over the previous best per dataset is about 2.77, and the per-modality average is about 2.21. Please state exactly how the 3.64 number is computed.
- [Figure 3] The PCA feature visualization is used to support the objective-alignment mechanism. Please add a quantitative feature-space alignment metric (e.g., CKA or k-nearest-neighbor agreement) between DEO, Copernicus-FM, and DINOv3 features, or explicitly label Figure 3 as illustrative only.
- [Supplement B.2] The sentence For all experiments, we use a batch size of 64 and fine-tune for 50 epochs using a learning rate is incomplete. Please provide the learning-rate values in the text or refer clearly to Table 11.
- [Table 11] Learning rates vary substantially across datasets and methods (e.g., 10^-4 to 10^-1 for different datasets, and SatDiFuser uses its own values). Please describe the hyperparameter-selection procedure to rule out per-dataset test-set overfitting.
Circularity Check
No significant circularity; the benchmark claims are held-out measurements and the method's losses are not fitted-then-relabeled predictions.
full rationale
I examined the derivation chain for circular reductions. The method's losses (Eqs. 3-5) combine a standard contrastive self-distillation objective (DINO-style with a coding-rate regularizer) and distillation from a frozen optical VFM (DINOv3). None of these equations is defined in terms of a downstream benchmark metric; the downstream numbers are obtained by fine-tuning on standard, held-out splits (GEO-Bench, SpaceNet, Sen1Floods11, PASTIS, LEVIR, OSCD), so there is no fitted input being relabeled as a prediction. The paper's central causal claim—that objective-level consistency with the contrastive VFM teacher drives the gains—is supported only by a confounded comparison against Copernicus-FM, which differs in backbone, pretraining data, and losses; this is a correctness risk in the attribution, but it is not circularity, since the comparison is not used to construct the method or to fit the reported numbers. The only self-citation of note is the change-detection evaluation protocol taken from the authors' prior TGRS paper [45]; that citation supplies an evaluation procedure (feature subtraction, UNet decoder choices) and is not a load-bearing premise of the method's derivation or of the benchmark results. No uniqueness theorem, ansatz, or known-result renaming is invoked in a way that reduces the paper's contributions to its own inputs. The limitations section also honestly notes dependence on the optical VFM's strength, which is a boundary condition rather than a circular step. Overall, the derivation is self-contained with respect to its empirical claims; no circularity is present.
Axiom & Free-Parameter Ledger
free parameters (5)
- Composite loss weights (α1, α2, α3, γ) =
α1=1, α2=0.5, α3=0.5, γ=1
- Number of high-resolution optical replacements in pretraining mix =
150,000
- EMA rate for multispectral teacher =
0.996
- Projection-head bottleneck dimensions =
256 (MS), 1024 (optical)
- Global/local view counts and crop ranges =
n=2 globals, m=10 locals; crops (0.4,1) and (0.05,0.4)
axioms (6)
- domain assumption Contrastive self-distillation with cosine similarity plus a coding-rate regularizer produces non-collapsed, semantically structured representations.
- domain assumption A frozen optical VFM (DINOv3) trained on natural/satellite imagery provides transferable semantic priors for EO tasks across the domain shift.
- domain assumption Sentinel-2 10-band input (atmospheric bands removed) plus the RGB subset is a sufficient representation for the target segmentation/change/classification tasks.
- domain assumption The fMoW-RGB aerial images and fMoW-Sentinel tiles are sufficiently spatially aligned to be treated as co-located views of the same scene for the 'high-res privileged knowledge' transfer.
- domain assumption The benchmark test splits used for component/teacher selection are acceptable as held-out evidence for the final numbers.
- domain assumption Per-method learning rates (Table 11), per-architecture decoders (UNet for ViT, UPerNet for Swin), and 50-epoch fine-tuning budgets give every method an equally fair evaluation.
read the original abstract
Foundation models are transforming Earth Observation (EO), yet the diversity of EO sensors and modalities makes a single universal model unrealistic. Multiple specialized EO foundation models (EOFMs) will likely coexist, making efficient knowledge transfer across modalities essential. Most existing EO pretraining relies on masked image modeling, which emphasizes local reconstruction but provides limited control over global semantic structure. To address this, we propose a dual-teacher contrastive distillation framework for multispectral imagery that aligns the student's pretraining objective with the contrastive self-distillation paradigm of modern optical vision foundation models (VFMs). Our approach combines a multispectral teacher with an optical VFM teacher, enabling coherent cross-modal representation learning. Experiments across diverse optical and multispectral benchmarks show that our model adapts to multispectral data without compromising performance on optical-only inputs, achieving state-of-the-art results in both settings, with an average improvement of 3.64 percentage points in semantic segmentation, 1.2 in change detection, and 1.31 in classification tasks. This demonstrates that contrastive distillation provides a principled and efficient approach to scalable representation learning across heterogeneous EO data sources. Project page: \textcolor{magenta}{https://wolfilip.github.io/DEO/}.
Figures
Reference graph
Works this paper leans on
-
[1]
Terrafm: A Scalable Foundation Model for Unified Multisensor Earth Observation
Muhammad Sohail anish, Muhammad Akhtar Munir, Syed Roshaan Ali Shah, Muhammad Haris Khan, Rao Muham- mad Anwer, Jorma Laaksonen, Fahad Shahbaz Khan, and Salman Khan. Terrafm: A Scalable Foundation Model for Unified Multisensor Earth Observation. In9th International Conference on Learning Representations, ICLR, 2026. 1, 2, 5, 6, 8, 14
2026
-
[2]
Omnisat: Self-Supervised Modality Fusion for Earth Observation
Guillaume Astruc, Nicolas Gonthier, Clement Mallet, and Loic Landrieu. Omnisat: Self-Supervised Modality Fusion for Earth Observation. InEuropean Conference on Computer Vision, pages 409–427. Springer, 2024. 2
2024
-
[3]
Anysat: One Earth Observation Model for Many Resolutions, Scales, and Modalities
Guillaume Astruc, Nicolas Gonthier, Clement Mallet, and Loic Landrieu. Anysat: One Earth Observation Model for Many Resolutions, Scales, and Modalities. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 19530–19540, 2025. 1, 2
2025
-
[4]
Mohammed Baharoon, Waseem Qureshi, Jiahong Ouyang, Yanwu Xu, Abdulrhman Aljouie, and Wei Peng. Evaluat- ing General Purpose Vision Foundation Models for Medical Image Analysis: An Experimental Study of Dinov2 on Radi- ology Benchmarks.arXiv preprint arXiv:2312.02366, 2023. 2
Pith/arXiv arXiv 2023
-
[5]
How learning by re- construction produces uninformative features for perception
Randall Balestriero and Yann Lecun. How learning by re- construction produces uninformative features for perception. InProceedings of the 41st International Conference on Ma- chine Learning, pages 2566–2585. PMLR, 2024. 2, 8
2024
-
[6]
Vi- creg: Variance-invariance-covariance regularization for self- supervised learning
Adrien Bardes, Jean Ponce, and Yann LeCun. Vi- creg: Variance-invariance-covariance regularization for self- supervised learning. InThe Tenth International Conference on Learning Representations, ICLR, 2022. 2
2022
-
[7]
Less is More? Data Spe- cialization for Self-Supervised Remote Sensing Models
Alvard Barseghyan, Ani Vanyan, Hakob Tamazyan, Evan Shelhamer, and Hrant Khachatrian. Less is More? Data Spe- cialization for Self-Supervised Remote Sensing Models. In TerraBytes-ICML 2025 Workshop, 2025. 2
2025
-
[8]
A Foundation Model for the Earth System.Nature, pages 1– 8, 2025
Cristian Bodnar, Wessel P Bruinsma, Ana Lucic, Megan Stanley, Anna Allen, Johannes Brandstetter, Patrick Garvan, Maik Riechert, Jonathan A Weyn, Haiyu Dong, et al. A Foundation Model for the Earth System.Nature, pages 1– 8, 2025. 1
2025
-
[9]
Sen1floods11: A Georeferenced Dataset to Train and Test Deep Learning Flood Algorithms for Sentinel-1
Derrick Bonafilia, Beth Tellman, Tyler Anderson, and Erica Issenberg. Sen1floods11: A Georeferenced Dataset to Train and Test Deep Learning Flood Algorithms for Sentinel-1. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pages 210–211,
-
[10]
Christopher F Brown, Michal R Kazmierski, Valerie J Pasquarella, William J Rucklidge, Masha Samsikova, Chen- hui Zhang, Evan Shelhamer, Estefania Lahera, Olivia Wiles, Simon Ilyushchenko, et al. Alphaearth Founda- tions: An Embedding Field Model for Accurate and Efficient Global Mapping From Sparse Label Data.arXiv preprint arXiv:2507.22291, 2025. 1
Pith/arXiv arXiv 2025
-
[11]
Emerg- ing Properties in Self-Supervised Vision Transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Herv ´e J´egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerg- ing Properties in Self-Supervised Vision Transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9650–9660, 2021. 2, 3, 4, 5
2021
-
[12]
A Spatial-Temporal Attention- Based Method and A New Dataset for Remote Sensing Im- age Change Detection.Remote Sensing, 12:1662, 2020
Hao Chen and Zhenwei Shi. A Spatial-Temporal Attention- Based Method and A New Dataset for Remote Sensing Im- age Change Detection.Remote Sensing, 12:1662, 2020. 6, 8, 13, 16
2020
-
[13]
A Simple Framework for Contrastive Learn- ing of Visual Representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Ge- offrey Hinton. A Simple Framework for Contrastive Learn- ing of Visual Representations. InInternational Conference on Machine Learning, pages 1597–1607. PmLR, 2020. 2, 4
2020
-
[14]
Exploring Simple Siamese Representation Learning
Xinlei Chen and Kaiming He. Exploring Simple Siamese Representation Learning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15750–15758, 2021. 3
2021
-
[15]
Functional Map of the World
Gordon Christie, Neil Fendley, James Wilson, and Ryan Mukherjee. Functional Map of the World. InProceed- ings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 6172–6180, 2018. 5
2018
-
[16]
Satmae: Pre-Training Transformers for Temporal and Multi-Spectral Satellite Imagery.Advances in Neural Information Processing Systems, 35:197–211, 2022
Yezhen Cong, Samar Khanna, Chenlin Meng, Patrick Liu, Erik Rozi, Yutong He, Marshall Burke, David Lobell, and Stefano Ermon. Satmae: Pre-Training Transformers for Temporal and Multi-Spectral Satellite Imagery.Advances in Neural Information Processing Systems, 35:197–211, 2022. 2, 5
2022
-
[17]
Urban Change Detection for Multispectral Earth Observation Using Convolutional Neural Networks
Rodrigo Caye Daudt, Bertr Le Saux, Alexandre Boulch, and Yann Gousseau. Urban Change Detection for Multispectral Earth Observation Using Convolutional Neural Networks. In IEEE International Geoscience and Remote Sensing Sympo- sium, pages 2115–2118. IEEE, 2018. 6, 8, 13, 16
2018
-
[18]
Robsense: A Robust Multi-Modal Foundation Model for Remote Sensing With Static, Temporal, and Incomplete Data Adaptability
Minh Kha Do, Kang Han, Phu Lai, Khoa T Phan, and Wei Xiang. Robsense: A Robust Multi-Modal Foundation Model for Remote Sensing With Static, Temporal, and Incomplete Data Adaptability. InProceedings of the Computer Vi- sion and Pattern Recognition Conference, pages 7427–7436,
-
[19]
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl- vain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In9th International Conference on Learning Repre- sentations, ICLR, 2021. 5
2021
-
[20]
DUNIA: pixel-sized embeddings via cross-modal alignment for earth observation applications
Ibrahim Fayad, Max Zimmer, Martin Schwartz, Fabian Gieseke, Philippe Ciais, Gabriel Belouze, Sarah Brood, Aur´elien de Truchis, and Alexandre d’Aspremont. DUNIA: pixel-sized embeddings via cross-modal alignment for earth observation applications. InForty-second International Con- ference on Machine Learning, ICML, 2025. 1, 2
2025
-
[21]
Croma: Remote Sensing Representations With Contrastive Radar- 9 Optical Masked Autoencoders.Advances in Neural Infor- mation Processing Systems, 36:5506–5538, 2023
Anthony Fuller, Koreen Millard, and James Green. Croma: Remote Sensing Representations With Contrastive Radar- 9 Optical Masked Autoencoders.Advances in Neural Infor- mation Processing Systems, 36:5506–5538, 2023. 2, 5, 6, 8, 14
2023
-
[22]
Panoptic Segmentation of Satellite Image Time Series With Convo- lutional Temporal Attention Networks
Vivien Sainte Fare Garnot and Loic Landrieu. Panoptic Segmentation of Satellite Image Time Series With Convo- lutional Temporal Attention Networks. InProceedings of the IEEE/CVF International Conference on Computer Vi- sion, pages 4872–4881, 2021. 5, 15
2021
-
[23]
Crossearth: Geospatial vision foundation model for domain generalizable remote sensing semantic segmentation.IEEE Transactions on Pat- tern Analysis and Machine Intelligence, 2025
Ziyang Gong, Zhixiang Wei, Di Wang, Xiaoxing Hu, Xi- anzheng Ma, Hongruixuan Chen, Yuru Jia, Yupeng Deng, Zhenming Ji, Xiangwei Zhu, et al. Crossearth: Geospatial vision foundation model for domain generalizable remote sensing semantic segmentation.IEEE Transactions on Pat- tern Analysis and Machine Intelligence, 2025. 2
2025
-
[24]
Bootstrap Your Own Latent-a New Ap- proach to Self-Supervised Learning.Advances in neural in- formation processing systems, 33:21271–21284, 2020
Jean-Bastien Grill, Florian Strub, Florent Altch ´e, Corentin Tallec, Pierre Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Guo, Mohammad Ghesh- laghi Azar, et al. Bootstrap Your Own Latent-a New Ap- proach to Self-Supervised Learning.Advances in neural in- formation processing systems, 33:21271–21284, 2020. 2, 3
2020
-
[25]
Bridging Remote Sensors With Multisensor Geospa- tial Foundation Models
Boran Han, Shuai Zhang, Xingjian Shi, and Markus Reich- stein. Bridging Remote Sensors With Multisensor Geospa- tial Foundation Models. InProceedings of the Ieee/cvf Con- ference on Computer Vision and Pattern Recognition, pages 27852–27862, 2024. 2
2024
-
[26]
Momentum Contrast for Unsupervised Visual Rep- resentation Learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum Contrast for Unsupervised Visual Rep- resentation Learning. InProceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 9729–9738, 2020. 2, 3
2020
-
[27]
Masked Autoencoders Are Scal- able Vision Learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Doll´ar, and Ross Girshick. Masked Autoencoders Are Scal- able Vision Learners. InProceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 16000–16009, 2022. 2
2022
-
[28]
Greg Heinrich, Mike Ranzinger, Hongxu Yin, Yao Lu, Jan Kautz, Andrew Tao, Bryan Catanzaro, and Pavlo Molchanov. Radiov2. 5: Improved Baselines for Agglomerative Vision Foundation Models. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 22487–22497,
-
[29]
Yuru Jia, Valerio Marsocci, Ziyang Gong, Xue Yang, Maarten Vergauwen, and Andrea Nascetti. Can generative geospatial diffusion models excel as discriminative geospa- tial foundation models? InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 8429– 8440, 2025. 1, 2, 5, 6, 8, 14
2025
-
[30]
Lobell, and Ste- fano Ermon
Samar Khanna, Patrick Liu, Linqi Zhou, Chenlin Meng, Robin Rombach, Marshall Burke, David B. Lobell, and Ste- fano Ermon. Diffusionsat: A generative foundation model for satellite imagery. InThe Twelfth International Confer- ence on Learning Representations, 2024. 2
2024
-
[31]
Kingma and Jimmy Ba
Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In3rd International Conference on Learning Representations, ICLR, 2015. 5
2015
-
[32]
Segment Any- thing
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer White- head, Alexander C Berg, Wan-Yen Lo, et al. Segment Any- thing. InProceedings of the IEEE/CVF International Con- ference on Computer Vision, pages 4015–4026, 2023. 2
2023
-
[33]
Geo- bench: Toward Foundation Models for Earth Monitoring
Alexandre Lacoste, Nils Lehmann, Pau Rodriguez, Evan Sherwin, Hannah Kerner, Bj¨orn L¨utjens, Jeremy Irvin, David Dao, Hamed Alemohammad, Alexandre Drouin, et al. Geo- bench: Toward Foundation Models for Earth Monitoring. Advances in Neural Information Processing Systems, 36: 51080–51093, 2023. 1, 5, 6, 8, 13, 15, 16
2023
-
[34]
Yansheng Li, Xinwei Li, Yongjun Zhang, Daifeng Peng, and Lorenzo Bruzzone. Cost-efficient Information Extraction From Massive Remote Sensing Data: When Weakly Super- vised Deep Learning Meets Remote Sensing Big Data.Inter- national Journal of Applied Earth Observation and Geoin- formation, 120:103345, 2023. 1
2023
-
[35]
Masked Angle-Aware Autoen- coder for Remote Sensing Images
Zhihao Li, Biao Hou, Siteng Ma, Zitong Wu, Xianpeng Guo, Bo Ren, and Licheng Jiao. Masked Angle-Aware Autoen- coder for Remote Sensing Images. InEuropean Conference on Computer Vision, pages 260–278. Springer, 2024. 2
2024
-
[36]
Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10012–10022, 2021. 5
2021
-
[37]
Towards Geospatial Foundation Models via Continual Pretraining
Mat ´ıas Mendieta, Boran Han, Xingjian Shi, Yi Zhu, and Chen Chen. Towards Geospatial Foundation Models via Continual Pretraining. InProceedings of the IEEE/CVF In- ternational Conference on Computer Vision, pages 16806– 16816, 2023. 2, 4, 5, 6, 8, 14
2023
-
[38]
SF Meneses III and AC Blanco. Rapid Mapping and As- sessment of Damages Due to Typhoon Rai Using Sentinel-1 Synthetic Aperture Radar Data.The International Archives of the Photogrammetry, Remote Sensing and Spatial Infor- mation Sciences, 43:1139–1146, 2022. 8
2022
-
[39]
Represen- tation Learning With Contrastive Predictive Coding.arXiv preprint arXiv:1807.03748, 2018
Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Represen- tation Learning With Contrastive Predictive Coding.arXiv preprint arXiv:1807.03748, 2018. 2
Pith/arXiv arXiv 2018
-
[40]
Maxime Oquab, Timoth ´ee Darcet, Th´eo Moutakanni, Huy V . V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel HAZIZA, Francisco Massa, Alaaeldin El-Nouby, Mido Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael Rabbat, Vasu Sharma, Gabriel Synnaeve, Hu Xu, Herve Je- gou, Julien Mairal, Patr...
2024
-
[41]
Learning Transferable Visual Models From Natural Language Super- vision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning Transferable Visual Models From Natural Language Super- vision. InInternational Conference on Machine Learning, pages 8748–8763. PmLR, 2021. 2
2021
-
[42]
Am-radio: Agglomerative Vision Foundation Model Reduce All Domains Into One
Mike Ranzinger, Greg Heinrich, Jan Kautz, and Pavlo Molchanov. Am-radio: Agglomerative Vision Foundation Model Reduce All Domains Into One. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12490–12500, 2024. 2 10
2024
-
[43]
Scale-mae: A Scale-Aware Masked Autoencoder for Multiscale Geospatial Representation Learning
Colorado J Reed, Ritwik Gupta, Shufan Li, Sarah Brockman, Christopher Funk, Brian Clipp, Kurt Keutzer, Salvatore Can- dido, Matt Uyttendaele, and Trevor Darrell. Scale-mae: A Scale-Aware Masked Autoencoder for Multiscale Geospatial Representation Learning. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 4088– 4099, 2023. 1...
2023
-
[44]
Position: Mission Critical–Satellite Data is a Distinct Modality in Machine Learning
Esther Rolf, Konstantin Klemmer, Caleb Robinson, and Hannah Kerner. Position: Mission Critical–Satellite Data is a Distinct Modality in Machine Learning. InForty-first International Conference on Machine Learning, 2024. 1
2024
-
[45]
Be the Change You Want to See: Revisiting Remote Sens- ing Change Detection Practices.IEEE Transactions on Geo- science and Remote Sensing, 63:1–11, 2025
Bla ˇz Rolih, Matic Fuˇcka, Filip Wolf, and LukaˇCehovin Zajc. Be the Change You Want to See: Revisiting Remote Sens- ing Change Detection Practices.IEEE Transactions on Geo- science and Remote Sensing, 63:1–11, 2025. 6, 14
2025
-
[46]
U- net: Convolutional networks for biomedical image segmen- tation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U- net: Convolutional networks for biomedical image segmen- tation. InInternational Conference on Medical image com- puting and computer-assisted intervention, pages 234–241. Springer, 2015. 14
2015
-
[47]
Dune: Distilling a Universal Encoder From Heterogeneous 2d and 3d Teachers
Mert B ¨ulent Sarıyıldız, Philippe Weinzaepfel, Thomas Lu- cas, Pau de Jorge, Diane Larlus, and Yannis Kalantidis. Dune: Distilling a Universal Encoder From Heterogeneous 2d and 3d Teachers. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 30084–30094,
-
[48]
Understanding Contrastive Versus Recon- structive Self-Supervised Learning of Vision Transform- ers
Shashank Shekhar, Florian Bordes, Pascal Vincent, and A Morcos. Understanding Contrastive Versus Recon- structive Self-Supervised Learning of Vision Transform- ers. InNeurIPS 2022 Workshop: Self-Supervised Learn- ing—Theory and Practice2022, 2022. 8
2022
-
[49]
Dinov3.arXiv preprint arXiv:2508.10104, 2025
Oriane Sim ´eoni, Huy V V o, Maximilian Seitzer, Federico Baldassarre, Maxime Oquab, Cijo Jose, Vasil Khalidov, Marc Szafraniec, Seungeun Yi, Micha ¨el Ramamonjisoa, et al. Dinov3.arXiv preprint arXiv:2508.10104, 2025. 2, 3, 4, 5, 6, 7, 8, 14
Pith/arXiv arXiv 2025
-
[50]
Galileo: Learning global & local features of many remote sensing modalities
Gabriel Tseng, Anthony Fuller, Marlena Reil, Henry Her- zog, Patrick Beukema, Favyen Bastani, James R Green, Evan Shelhamer, Hannah Kerner, and David Rolnick. Galileo: Learning global & local features of many remote sensing modalities. InForty-second International Conference on Machine Learning, 2025. 1, 2, 5
2025
-
[51]
To- wards Advanced Wildfire Analysis: A Siamese Network- Based Change Detection Approach Through Self-Supervised Learning
Dimitris Valsamis, Alexandros Oikonomidis, Chrysoula Chatzichristaki, Anastasia Moumtzidou, Ilias Gialam- poukidis, Stefanos Vrochidis, and Ioannis Kompatsiaris. To- wards Advanced Wildfire Analysis: A Siamese Network- Based Change Detection Approach Through Self-Supervised Learning. In2024 International Conference on Content- Based Multimedia Indexing (C...
2024
-
[52]
Spacenet: A Remote Sensing Dataset and Challenge Series
Adam Van Etten, Dave Lindenbaum, and Todd M Bacastow. Spacenet: A Remote Sensing Dataset and Challenge Series. arXiv preprint arXiv:1807.01232, 2018. 1, 5, 8, 15
Pith/arXiv arXiv 2018
-
[53]
Panopticon: Advancing Any- Sensor Foundation Models for Earth Observation
Leonard Waldmann, Ando Shah, Yi Wang, Nils Lehmann, Adam Stewart, Zhitong Xiong, Xiao Xiang Zhu, Stefan Bauer, and John Chuang. Panopticon: Advancing Any- Sensor Foundation Models for Earth Observation. InPro- ceedings of the Computer Vision and Pattern Recognition Conference, pages 2204–2214, 2025. 1, 2, 4
2025
-
[54]
MTP: Advancing Remote Sensing Foun- dation Model via Multi-Task Pretraining.IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2024
Di Wang, Jing Zhang, Minqiang Xu, Lin Liu, Dongsheng Wang, Erzhong Gao, Chengxi Han, Haonan Guo, Bo Du, Dacheng Tao, et al. MTP: Advancing Remote Sensing Foun- dation Model via Multi-Task Pretraining.IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2024. 6, 14
2024
-
[55]
Multi- label Guided Soft Contrastive Learning for Efficient Earth Observation Pretraining.IEEE Transactions on Geoscience and Remote Sensing, 2024
Yi Wang, Conrad M Albrecht, and Xiao Xiang Zhu. Multi- label Guided Soft Contrastive Learning for Efficient Earth Observation Pretraining.IEEE Transactions on Geoscience and Remote Sensing, 2024. 2
2024
-
[56]
Towards a Unified Copernicus Foundation Model for Earth Vision
Yi Wang, Zhitong Xiong, Chenying Liu, Adam J Stewart, Thomas Dujardin, Nikolaos Ioannis Bountos, Angelos Za- vras, Franziska Gerken, Ioannis Papoutsis, Laura Leal-Taix´e, et al. Towards a Unified Copernicus Foundation Model for Earth Vision. InProceedings of the IEEE/CVF International Conference on Computer Vision, 2025. 1, 2, 4, 5, 6, 8, 14
2025
-
[57]
Extending Global-Local View Alignment for Self-Supervised Learning With Remote Sensing Imagery
Xinye Wanyan, Sachith Seneviratne, Shuchang Shen, and Michael Kirley. Extending Global-Local View Alignment for Self-Supervised Learning With Remote Sensing Imagery. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2443–2453, 2024. 2
2024
-
[58]
Simplifying DINO via Coding Rate Regularization
Ziyang Wu, Jingyuan Zhang, Druv Pai, XuDong Wang, Chandan Singh, Jianwei Yang, Jianfeng Gao, and Yi Ma. Simplifying DINO via Coding Rate Regularization. In Forty-second International Conference on Machine Learn- ing, 2025. 2, 3, 13
2025
-
[59]
Foundation Models for Remote Sensing and Earth Observation: A Sur- vey.IEEE Geoscience and Remote Sensing Magazine, 2025
Aoran Xiao, Weihao Xuan, Junjue Wang, Jiaxing Huang, Dacheng Tao, Shijian Lu, and Naoto Yokoya. Foundation Models for Remote Sensing and Earth Observation: A Sur- vey.IEEE Geoscience and Remote Sensing Magazine, 2025. 1
2025
-
[60]
Unified Perceptual Parsing for Scene Understand- ing
Tete Xiao, Yingcheng Liu, Bolei Zhou, Yuning Jiang, and Jian Sun. Unified Perceptual Parsing for Scene Understand- ing. InProceedings of the European Conference on Com- puter Vision (ECCV), pages 418–434, 2018. 5, 6, 14
2018
-
[61]
Simmim: A Simple Framework for Masked Image Modeling
Zhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin, Jianmin Bao, Zhuliang Yao, Qi Dai, and Han Hu. Simmim: A Simple Framework for Masked Image Modeling. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9653–9663, 2022. 2
2022
-
[62]
Neu- ral Plasticity-Inspired Foundation Model for Observing the Earth Crossing Modalities.CoRR, 2024
Zhitong Xiong, Yi Wang, Fahong Zhang, Adam J Stewart, Jo¨elle Hanna, Damian Borth, Ioannis Papoutsis, Bertrand Le Saux, Gustau Camps-Valls, and Xiao Xiang Zhu. Neu- ral Plasticity-Inspired Foundation Model for Observing the Earth Crossing Modalities.CoRR, 2024. 2
2024
-
[63]
One for All: Toward Unified Foundation Models for Earth Vision
Zhitong Xiong, Yi Wang, Fahong Zhang, and Xiao Xiang Zhu. One for All: Toward Unified Foundation Models for Earth Vision. InIGARSS 2024-2024 IEEE International Geoscience and Remote Sensing Symposium, pages 2734–
2024
-
[64]
Barlow Twins: Self-Supervised Learning via Redundancy Reduction
Jure Zbontar, Li Jing, Ishan Misra, Yann LeCun, and St´ephane Deny. Barlow Twins: Self-Supervised Learning via Redundancy Reduction. InInternational Conference on Machine Learning, pages 12310–12320. PMLR, 2021. 2, 4
2021
-
[65]
SkySense V2: A Uni- 11 fied Foundation Model for Multi-Modal Remote Sensing
Yingying Zhang, Lixiang Ru, Kang Wu, Lei Yu, Lei Liang, Yansheng Li, and Jingdong Chen. SkySense V2: A Uni- 11 fied Foundation Model for Multi-Modal Remote Sensing. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025. 1, 2
2025
-
[66]
ibot: Image bert pre-training with online tokenizer.International Conference on Learning Representations (ICLR), 2022
Jinghao Zhou, Chen Wei, Huiyu Wang, Wei Shen, Cihang Xie, Alan Yuille, and Tao Kong. ibot: Image bert pre-training with online tokenizer.International Conference on Learning Representations (ICLR), 2022. 2 12 Brewing Stronger Features: Dual-Teacher Distillation for Multispectral Earth Observation Supplementary Material A. Pretraining In this section, we w...
2022
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.