Pith. sign in

REVIEW 5 major objections 7 minor 83 references

CGEarthEye:A High-Resolution Remote Sensing Vision Foundation Model Based on the Jilin-1 Satellite Constellation

T0 review · 5 major / 7 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read A 2.1-billion-parameter vision model family pretrained on 15 million sub-meter Jilin-1 satellite images reports state-of-the-art accuracy on 10 remote sensing benchmarks across four Earth-observation tasks.

desk verdict The new dataset is the real news; the claim of 'consistently SOTA' is contradicted by the paper's own tables, and with softened claims and released assets this is a solid empirical study. read the letter →

arxiv 2507.00356 v1 pith:IS3RJFTW submitted 2025-07-01 cs.CV cs.AI

classification cs.CVcs.AI
keywords remotesensingfoundationmodelself-supervisedlearningVisionTransformerJilin-1satelliteconstellationseasonalcontrastivemaskedimagemodelingsub-meterimageryEarthobservation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to show that a vision foundation model trained on a large, global, sub-meter-resolution corpus from the Jilin-1 satellite constellation learns features that transfer better to Earth-observation tasks than models pretrained on medium-resolution or natural imagery. To do this it builds JLSSD, a 15-million-image self-supervised dataset with quarterly temporal sampling, and pretrains five Vision Transformer backbones with a hybrid objective that mixes augmentation-based contrast, seasonal contrast, and masked-patch token contrast. Across 10 benchmarks covering scene classification, object detection, semantic segmentation, and change detection, the largest model (CGEarthEye-Giant, 1.1 billion parameters) is reported to reach state-of-the-art or near-state-of-the-art accuracy even with a frozen backbone. If correct, this means high-resolution commercial satellite constellations can serve as a general-purpose pretraining source that lowers the cost of downstream adaptation while matching or beating fully fine-tuned task-specific models.

What carries the argument

The load-bearing mechanism is the joint teacher-student pretraining objective. The teacher branch encodes two global views and supplies stable anchors; the student branch encodes eight local crops, three seasonal views, and two masked variants, and is trained to match the teacher's class tokens and masked-patch tokens through three cross-entropy losses: augmentation-aware contrastive loss, seasonal contrastive loss, and masked-patch token contrastive loss. The student is updated by backpropagation while the teacher is updated by exponential moving average, preventing collapse. The other half of the machinery is JLSSD's sampling design: 1-kilometer grid cells stratified by ESA WorldCover land cover, elevation bins, and administrative region, with quarterly mosaics in China and annual mosaics outside, which is what supplies the seasonal signal and global diversity.

What would settle it

Run a near-duplicate or embedding-similarity search between JLSSD images and the test sets of RESISC-45, AID, DIOR, DIOR-R, LoveDA, iSAID, Potsdam, LEVIR-CD, SYSU-CD, and CDD; if a substantial fraction of benchmark test images appear in pretraining, the SOTA claim collapses. A cleaner experiment is to retrain CGEarthEye on JLSSD with all Chinese-sampled images removed and check whether the frozen-backbone gains on Chinese-heavy benchmarks like LoveDA and DIOR survive.

Watch

Extended reading notes

Core claim

The paper's central claim is that CGEarthEye consistently reaches state-of-the-art performance on 10 benchmarks across four remote sensing tasks when the pretrained backbone is frozen during fine-tuning. The evidence is a set of comparisons: frozen CGEarthEye-Giant surpasses all compared remote sensing foundation models, including fully fine-tuned SkySense and MTP, on RESISC-45, AID, DIOR, DIOR-R, LoveDA, and SYSU-CD, and lands within a fraction of a point on iSAID, Potsdam, LEVIR-CD, and CDD where fully fine-tuned SkySense or MTP lead. The paper attributes this to pretraining on JLSSD, which provides 15 million 0.75-meter images from 10 million global grid cells with quarterly temporal views for Chinese regions and annual mosaics elsewhere, combined with a three-part self-supervised objective. A supporting analysis shows that scaling from 22 million to 1.1 billion parameters raises classification accuracy by up to 8 percentage points on the harder fMoW benchmark, and that the pretrained model converges faster and outperforms the natural-image model DINOv2 on all four tasks.

Load-bearing premise

The load-bearing premise is that JLSSD's sampled global grid cells are representative of the world and do not overlap with the benchmark imagery; if the pretraining set contains test images or concentrates on the same Chinese regions as the benchmarks, the reported frozen-backbone improvements would be inflated.

Editorial extensions

If this is right

  • Frozen-backbone fine-tuning becomes a competitive deployment mode for sub-meter Earth-observation models, so users with limited GPU budgets can still reach near-top accuracy.
  • Larger pretrained backbones pay off most on hard, diverse benchmarks like fMoW; on easier benchmarks, 307-million-parameter models saturate, so model size can be matched to task difficulty.
  • The comparison with DINOv2 implies that remote-sensing-specific pretraining data and objectives, not just model scale, drive downstream gains in Earth-observation tasks.
  • The Jilin-1 application cases indicate the same frozen features transfer to operational mapping tasks such as crane detection, building extraction, and building change detection.
  • Full fine-tuning still beats frozen backbones across classification benchmarks, so releasing the full model family lets users trade accuracy against compute.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper does not report a leakage or overlap analysis between JLSSD and the 10 public benchmarks; testing for near-duplicate imagery would determine how much of the frozen-backbone gain is due to transferable features rather than coincidental geographic overlap.
  • If the seasonal contrastive loss is a key driver, then change detection should benefit disproportionately; the reported first-place SYSU-CD result is consistent with that, and an ablation that removes only the seasonal term would isolate it.
  • The global sampling strategy is built around Jilin-1's coverage and China-centric quarterly mosaics; other constellations with different revisit patterns could test whether the method transfers or whether the gains are tied to this particular data distribution.
  • The 150-day, 16-GPU pretraining budget and released weights make CGEarthEye a practical reference point for commercial sub-meter constellations, but reproducibility depends on access to the JLSSD imagery itself; publishing a sample or coordinate list would let others measure distribution shift.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 7 minor

Summary. The paper introduces CGEarthEye, a family of five ViT-based remote sensing foundation models with parameter counts from 22M to 1.1B (2.1B total), pretrained on JLSSD, a newly constructed dataset of roughly 15 million sub-meter-resolution images from the Jilin-1 satellite constellation. The pretraining objective combines augmentation-aware contrastive learning, seasonal alignment contrastive learning, and masked-patch token contrastive learning. The authors evaluate the frozen-backbone variants on 10 benchmarks covering scene classification (RESISC-45, AID), object detection (DIOR, DIOR-R), semantic segmentation (LoveDA, iSAID, Potsdam), and change detection (LEVIR-CD, SYSU-CD, CDD), reporting state-of-the-art results and additional studies on parameter scaling, convergence, comparison to DINOv2, and practical deployment for crane detection, building extraction, and change detection in Longhua District, Shenzhen.

Significance. If the central claims are supported, this is a valuable industrial-scale contribution: it demonstrates that proprietary sub-meter commercial satellite imagery can support large-scale self-supervised pretraining for remote sensing, and the frozen-backbone transfer results on some benchmarks (e.g., DIOR/DIOR-R detection, SYSU-CD) are genuinely competitive with fully fine-tuned prior models. The promised release of code and weights (mentioned in the abstract) is a strength. However, the headline claim of consistent state-of-the-art performance is contradicted by the paper's own tables, and the absence of any overlap analysis between the pretraining dataset and the evaluation benchmarks leaves a load-bearing concern unresolved. The paper would be significantly strengthened by correcting the overstated claims, providing leakage analysis, and reconciling the numerical inconsistencies.

major comments (5)
  1. [Abstract; §3.2; §5] The claim that CGEarthEye 'consistently achieves state-of-the-art (SOTA) performance' is contradicted by the paper's own results. In Table 3, frozen CGEarthEye (ViT-G) scores 0.9584 on RESISC-45 and 0.9760 on AID, below fully fine-tuned SkySense (0.9632 and 0.9768) and below CGEarthEye's own fully fine-tuned variant (0.9675 and 0.9769). In Table 5, frozen CGEarthEye (0.6951 mIoU on iSAID, 0.9353 on Potsdam) trails fully fine-tuned SkySense (0.7091 and 0.9399). In Table 6, frozen CGEarthEye (0.9246 on LEVIR-CD, 0.9804 on CDD) trails MTP (0.9267 and 0.9837) and SkySense on LEVIR-CD (0.9258). Counting all ten benchmark rows, frozen CGEarthEye is top-ranked on only four (DIOR, DIOR-R, LoveDA, SYSU-CD). The headline should be revised to a benchmark-specific claim such as 'competitive with or better than fully fine-tuned state-of-the-art models on several benchmarks,' and the conclusion in §5 must be corrected accordingly.
  2. [§2.1] No overlap or leakage analysis is provided between the JLSSD pretraining dataset and the nine evaluation benchmarks (RESISC-45, AID, DIOR, LoveDA, iSAID, Potsdam, LEVIR-CD, SYSU-CD, CDD). The sampling procedure partitions the global into 1 km × 1 km grid cells and samples from all land-cover classes, which could plausibly include imagery from these benchmark datasets, especially those derived from Google Earth or publicly available aerial/satellite sources. If JLSSD inadvertently contains training images from the test distributions, the reported frozen-backbone gains would be inflated. Please provide an overlap analysis (e.g., geolocation-based exclusion, image hashing, or a statement of why overlap is impossible) for all benchmark datasets.
  3. [§4.6, Table 9] The numerical results for crane detection are internally inconsistent. The text states 'CGEarthEye identified 221 cranes compared to YOLOv8's 234 detections,' but Table 9 lists YOLOv8 detections as 213. Furthermore, the recall values in Table 9 (0.7472 for YOLOv8 and 0.8261 for CGEarthEye) do not equal the detection counts divided by ground truth (213/253 = 0.842 and 221/253 = 0.874, respectively). If recall is computed as true positives divided by ground truths, then the detection counts and recalls imply different numbers of true positives; the metric definitions (e.g., whether 'detections' includes false positives) must be clarified and the values reconciled.
  4. [§2.1] The dataset composition is numerically inconsistent with the stated scale. The paper reports 8.06 million quarterly image samples from 2.015 million Chinese locations and 7.985 million annual mosaic samples from 7.985 million non-Chinese locations, which sum to 16.045 million images, not '15 million.' Also, JLSSD is described as a 'large-scale supervised dataset' in §2.1, though it is used for self-supervised learning; this should be corrected to 'self-supervised dataset.' These inconsistencies undermine the precision of a key contribution claim.
  5. [§3.2.1, Table 3] The comparison protocol is not apples-to-apples, and the text's characterization is misleading. The section states that CGEarthEye 'significantly outperforms existing remote sensing foundation models' and 'even with a frozen backbone ... consistently surpasses other vision foundation models.' However, Table 3 shows frozen CGEarthEye (0.9584) below fully fine-tuned SkySense (0.9632) on RESISC-45 and below fully fine-tuned SkySense and CGEarthEye on AID. The statement should explicitly distinguish frozen vs. fine-tuned settings and avoid claiming consistent superiority over all compared models when the table shows otherwise.
minor comments (7)
  1. [Abstract and throughout] There are numerous typos and grammatical errors, e.g., 'rinterpretation', 'fundation', 'strategie', 'Potsdom', 'containsg', 'Trainging', and 'RFVFNs'. The manuscript would benefit from careful proofreading.
  2. [§2.2.3] The loss equations are not properly typeset (e.g., 'classtoken logt s T S L p p=∑∑' and 'patch s i logti iL p p=∑'). Please use standard LaTeX formatting so the mathematical definitions are legible and unambiguous.
  3. [§2.1] The description of the grid-stratified sampling formula is garbled. The notation 'c_i_rS' and the definition of '_ _ _c i r sampleM' are not rendered correctly. Please rewrite the formula and its definitions clearly.
  4. [§2.1, §2.2.1] It is unclear how the three 'seasonal contrastive views' are generated for non-Chinese regions, which are sampled from annual mosaics rather than quarterly data. Please specify the temporal composition of the seasonal views for both Chinese and non-Chinese samples.
  5. [§3.2.3] The sentence 'Data processing follows SkySense and MTP' is vague. Please specify the exact preprocessing steps (e.g., band selection, normalization, tiling) used for LoveDA, iSAID, and Potsdam so that the results are reproducible.
  6. [§4.6] The 'quadrat sampling methodology' used for accuracy assessment in the three deployment studies is not described. Please explain how ground truth counts (e.g., 253 cranes, 38,857 buildings) were obtained and how the quadrats were selected, as this directly affects the validity of the reported recall, precision, and F1 values.
  7. [Abstract] The statement 'The code and pre-trained model weights will be released at: https://github.com/1921134176/CGEarthEye' currently points to a non-existent repository. Please provide a working link or state the intended release date.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the central results are external benchmark measurements, and the pre-training objective does not use downstream labels or fitted constants.

full rationale

The paper's derivation chain is: construct JLSSD by stratified grid sampling from Jilin-1 imagery; pre-train ViT backbones with three SSL losses (augmentation contrast, seasonal contrast, masked-patch contrast); transfer to external benchmarks. None of these stages defines a benchmark quantity in terms of itself. The contrastive and reconstruction losses use only image views and teacher-student features, not downstream labels or test-set statistics, so the reported benchmark accuracies are genuine post-hoc measurements. The comparisons in Tables 3-6 are against external published models; the authors' self-citations in references [7] and [8] are contextual Jilin-1 application studies and are not load-bearing for the SSL or SOTA claims. The paper does not invoke any uniqueness theorem or ansatz by self-citation. Possible weaknesses, such as a lack of overlap analysis between JLSSD and the evaluation benchmarks, or the tension between the abstract's 'consistently achieves SOTA' and Tables 5-6 where frozen CGEarthEye trails SkySense/MTP on iSAID, Potsdam, LEVIR-CD and CDD, are empirical and correctness concerns, not circularity. No fitted parameter is renamed as a prediction, and no result is equivalent to its input by construction. Therefore the circularity score is 0.

Assumptions & free parameters 5 free parameters · 4 assumptions · 0 invented entities

The central empirical claims rest on the JLSSD pretraining corpus and on standard SSL components. There are no fitted physical parameters, but the dataset stratification, sampling formula, loss weighting, and key hyperparameters are hand choices. The main unverified premises are that the auxiliary maps used for stratification are reliable, that quarterly mosaics are co-registered, and that JLSSD does not overlap the benchmark test sets. These assumptions are not validated externally in the paper.

free parameters (5)
  • EMA momentum m = 0.992
    Set by hand in Section 2.2.3; controls teacher parameter updates and prevents collapse during pretraining.
  • Masking ratio range = 10%-50%
    Chosen in Section 2.2.1; determines the difficulty of the masked patch token contrastive objective.
  • Global and local crop resolutions = 224/98, then 518 global
    Training schedule in Section 3.1; crop scales are design choices that affect representation granularity.
  • Sampling strata definitions = 7 land cover classes; 24 elevation bins; county/country admin
    Section 2.1; the stratification and sampling formula determine JLSSD composition and global balance.
  • Pretraining loss weights = 1, 1, 1 (implicit)
    The final loss in Section 2.2.3 sums three terms without reported weighting; equal weighting is a hand choice.
assumptions (4)
  • domain assumption ESA WorldCover land cover labels and global DEM are accurate enough for stratifying the 1 km grid cells used to build JLSSD.
    Sampling in Section 2.1 partitions grids by land cover, elevation, and administrative region; errors in these auxiliary maps would bias the sampled distribution.
  • domain assumption Quarterly Jilin-1 mosaics within 2023 are co-registered and radiometrically consistent, so seasonal views of the same grid represent the same ground area.
    The seasonal contrast loss in Section 2.2.3 aligns features across seasons; misregistration would turn seasonal alignment into a noise-fitting task.
  • ad hoc to paper The 1 km grid stratification and cluster filtering produce a globally balanced and non-redundant sample; removing redundant scenes does not bias toward easy scenes.
    No external validation of JLSSD representativeness is provided; the sampling formula and cluster filtering are introduced specifically for this dataset.
  • standard math Standard ViT, EMA, contrastive, and masked-image-modeling components from the cited literature behave as described and do not require re-derivation.
    The framework relies on ViT [60], EMA [61], contrastive objectives, and masked patch modeling from prior work without re-deriving them.

how reviews work

0 comments
Cite this review

Pith. "Pith review of CGEarthEye:A High-Resolution Remote Sensing Vision Foundation Model Based on the Jilin-1 Satellite Constellation." pith.science (2026). https://pith.science/paper/IS3RJFTW

@misc{pith2026250700356,
  author       = {Pith},
  title        = {Pith review of: CGEarthEye:A High-Resolution Remote Sensing Vision Foundation Model Based on the Jilin-1 Satellite Constellation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IS3RJFTW}},
  note         = {Machine review of arXiv:2507.00356}
}
read the original abstract

Deep learning methods have significantly advanced the development of intelligent rinterpretation in remote sensing (RS), with foundational model research based on large-scale pre-training paradigms rapidly reshaping various domains of Earth Observation (EO). However, compared to the open accessibility and high spatiotemporal coverage of medium-resolution data, the limited acquisition channels for ultra-high-resolution optical RS imagery have constrained the progress of high-resolution remote sensing vision foundation models (RSVFM). As the world's largest sub-meter-level commercial RS satellite constellation, the Jilin-1 constellation possesses abundant sub-meter-level image resources. This study proposes CGEarthEye, a RSVFM framework specifically designed for Jilin-1 satellite characteristics, comprising five backbones with different parameter scales with totaling 2.1 billion parameters. To enhance the representational capacity of the foundation model, we developed JLSSD, the first 15-million-scale multi-temporal self-supervised learning (SSL) dataset featuring global coverage with quarterly temporal sampling within a single year, constructed through multi-level representation clustering and sampling strategies. The framework integrates seasonal contrast, augmentation-based contrast, and masked patch token contrastive strategies for pre-training. Comprehensive evaluations across 10 benchmark datasets covering four typical RS tasks demonstrate that the CGEarthEye consistently achieves state-of-the-art (SOTA) performance. Further analysis reveals CGEarthEye's superior characteristics in feature visualization, model convergence, parameter efficiency, and practical mapping applications. This study anticipates that the exceptional representation capabilities of CGEarthEye will facilitate broader and more efficient applications of Jilin-1 data in traditional EO application.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

83 extracted references · 79 canonical work pages

  1. [1]

    Classifying vegetation communities karst wetland synergistic use of image fusion and object-based machine learning algorithm with Jilin-1 and UAV multispectral images,

    B. Fu, P. Zuo, M. Liu, G. Lan, H. He, Z. Lao, Y. Zhang, D. Fan, and E. Gao, “Classifying vegetation communities karst wetland synergistic use of image fusion and object-based machine learning algorithm with Jilin-1 and UAV multispectral images,” Ecol. Indic., vol. 140, 2022 JUL. 2022

  2. [2]

    Multi-Object Tracking in Satellite Videos With Graph-Based Multitask Modeling,

    Q. He, X. Sun, Z. Yan, B. Li, and K. Fu, “Multi-Object Tracking in Satellite Videos With Graph-Based Multitask Modeling,” IEEE Trans. Geosci. Remote Sens., vol. 60, 2022. 2022

  3. [3]

    Satellite Video Super-Resolution via Multiscale Deformable Convolution Alignment and Temporal Grouping Projection,

    Y. Xiao, X. Su, Q. Yuan, D. Liu, H. Shen, and L. Zhang, “Satellite Video Super-Resolution via Multiscale Deformable Convolution Alignment and Temporal Grouping Projection,” IEEE Trans. Geosci. Remote Sens., vol. 60, 2022. 2022. 31 31

  4. [4]

    Detecting and Tracking Small and Dense Moving Objects in Satellite Videos: A Benchmark,

    Q. Yin, Q. Hu, H. Liu, F. Zhang, Y. Wang, Z. Lin, W. An, and Y. Guo, “Detecting and Tracking Small and Dense Moving Objects in Satellite Videos: A Benchmark,” IEEE Trans. Geosci. Remote Sens., vol. 60,

  5. [5]

    Analyzing spatial variability in night-time lights using a high spatial resolution color Jilin-1 image - Jerusalem as a case study,

    E. Guk, and N. Levin, “Analyzing spatial variability in night-time lights using a high spatial resolution color Jilin-1 image - Jerusalem as a case study,” ISPRS J. Photogramm. Remote Sens., vol. 163, pp. 121-136, 2020 MAY. 2020

  6. [6]

    Regional rare-earth element supply and demand balanced with circular economy strategies,

    P. Wang, Y. Y. Yang, O. Heidrich, L. Y. Chen, L. H. Chen, T. Fishman, and W. Q. Chen, “Regional rare-earth element supply and demand balanced with circular economy strategies,” Nat. Geosci., vol. 17, no. 1, JAN. 2024

  7. [7]

    AFWS: Angle-Free Weakly Supervised Rotating Object Detection for Remote Sensing Images,

    J. Lu, Q. Hu, R. Zhu, Y. Wei, and T. Li, “AFWS: Angle-Free Weakly Supervised Rotating Object Detection for Remote Sensing Images,” IEEE Trans. Geosci. Remote Sens., vol. 62, 2024. 2024

  8. [8]

    Full Convolution Neural Network Combined with Contextual Feature Representation for Cropland Extraction from High-Resolution Remote Sensing Images,

    Z. Li, S. Chen, X. Meng, R. Zhu, J. Lu, L. Cao, and P. Lu, “Full Convolution Neural Network Combined with Contextual Feature Representation for Cropland Extraction from High-Resolution Remote Sensing Images,” REMOTE SENSING, vol. 14, no. 9, 2022 MAY. 2022

Show all 83 references
  1. [9]

    Deep Residual Learning for Image Recognition,

    K. He, X. Zhang, S. Ren, J. Sun, and Ieee, “Deep Residual Learning for Image Recognition,” in 2016 IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2016, pp. 770-778

  2. [10]

    Rethinking Atrous Convolution for Semantic Image Segmentation,

    L.-C. Chen, G. Papandreou, F. Schroff, and H. J. a. e.-p. Adam, "Rethinking Atrous Convolution for Semantic Image Segmentation," https://ui.adsabs.harvard.edu/abs/2017arXiv170605587C, [June 01, 2017, 2017]

  3. [11]

    High-Resolution Representations for Labeling Pixels and Regions,

    K. Sun, Y. Zhao, B. Jiang, T. Cheng, B. Xiao, D. Liu, Y. Mu, X. Wang, W. Liu, and J. Wang, “High-Resolution Representations for Labeling Pixels and Regions,” Arxiv, 2019. 2019

  4. [12]

    A ConvNet for the 2020s,

    Z. Liu, H. Mao, C.-Y. Wu, C. Feichtenhofer, T. Darrell, S. Xie, and S. O. C. Ieee Comp, “A ConvNet for the 2020s,” in 2022 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2022, pp. 11966-11976

  5. [13]

    CLNet: Cross-layer convolutional neural network for change detection in optical remote sensing imagery,

    Z. Zheng, Y. Wan, Y. Zhang, S. Xiang, D. Peng, and B. Zhang, “CLNet: Cross-layer convolutional neural network for change detection in optical remote sensing imagery,” ISPRS J. Photogramm. Remote Sens., vol. 175, pp. 247-267, 2021 MAY. 2021

  6. [14]

    SAMPolyBuild: Adapting the Segment Anything Model for polygonal building extraction,

    C. Wang, J. Chen, Y. Meng, Y. Deng, K. Li, and Y. Kong, “SAMPolyBuild: Adapting the Segment Anything Model for polygonal building extraction,” ISPRS J. Photogramm. Remote Sens., vol. 218, pp. 707-720. 2024

  7. [15]

    Remote Sensing Image Scene Classification Based on an Enhanced Attention Module,

    Z. Zhao, J. Li, Z. Luo, J. Li, and C. Chen, “Remote Sensing Image Scene Classification Based on an Enhanced Attention Module,” IEEE Geosci. Remote Sens. Lett., vol. 18, no. 11, pp. 1926-1930, 2021 NOV. 2021

  8. [16]

    Remote Sensing Scene Classification via Multi-Branch Local Attention Network,

    S.-B. Chen, Q.-S. Wei, W.-Z. Wang, J. Tang, B. Luo, and Z.-Y. Wang, “Remote Sensing Scene Classification via Multi-Branch Local Attention Network,” IEEE Trans. Image Process., vol. 31, pp. 99-109, 2022. 2022

  9. [17]

    SNUNet-CD: A Densely Connected Siamese Network for Change Detection of VHR Images,

    S. Fang, K. Li, J. Shao, and Z. Li, “SNUNet-CD: A Densely Connected Siamese Network for Change Detection of VHR Images,” IEEE Geosci. Remote Sens. Lett., vol. 19, 2022. 2022

  10. [18]

    Advancing Plain Vision Transformer Toward Remote Sensing Foundation Model,

    D. Wang, Q. Zhang, Y. Xu, J. Zhang, B. Du, D. Tao, and L. Zhang, “Advancing Plain Vision Transformer Toward Remote Sensing Foundation Model,” IEEE Trans. Geosci. Remote Sens., vol. 61, 2023. 2023

  11. [19]

    ChangeCLIP: Remote sensing change detection with multimodal vision-language representation learning,

    S. J. Dong, L. B. Wang, B. Du, and X. L. Meng, “ChangeCLIP: Remote sensing change detection with multimodal vision-language representation learning,” ISPRS J. Photogramm. Remote Sens., vol. 208, pp. 53-69, FEB. 2024. 32 32

  12. [20]

    An Empirical Study of Remote Sensing Pretraining,

    D. Wang, J. Zhang, B. Du, G.-S. Xia, and D. Tao, “An Empirical Study of Remote Sensing Pretraining,” IEEE Trans. Geosci. Remote Sens., vol. 61, 2023. 2023

  13. [21]

    An Empirical Study of Training Self-Supervised Vision Transformers,

    X. Chen, S. Xie, K. He, and Ieee, “An Empirical Study of Training Self-Supervised Vision Transformers,” in 2021 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2021), 2021, pp. 9620-9629

  14. [22]

    AID: A Benchmark Data Set for Performance Evaluation of Aerial Scene Classification,

    G.-S. Xia, J. Hu, F. Hu, B. Shi, X. Bai, Y. Zhong, L. Zhang, and X. Lu, “AID: A Benchmark Data Set for Performance Evaluation of Aerial Scene Classification,” IEEE Trans. Geosci. Remote Sens., vol. 55, no. 7, pp. 3965-3981, 2017 JUL. 2017

  15. [23]

    OpenSARShip: A Dataset Dedicated to Sentinel-1 Ship Interpretation,

    L. Huang, B. Liu, B. Li, W. Guo, W. Yu, Z. Zhang, and W. Yu, “OpenSARShip: A Dataset Dedicated to Sentinel-1 Ship Interpretation,” IEEE J. Sel. Topics Appl. Earth Observ. Remote Sens., vol. 11, no. 1, pp. 195-208, 2018 JAN. 2018

  16. [24]

    SSL4EO-S12: A large-scale multimodal, multitemporal dataset for self-supervised learning in Earth observation [Software and Data Sets],

    Y. Wang, N. A. A. Braham, Z. Xiong, C. Liu, C. M. Albrecht, and X. X. Zhu, “SSL4EO-S12: A large-scale multimodal, multitemporal dataset for self-supervised learning in Earth observation [Software and Data Sets],” IEEE Geosci. Remote Sens. Mag., vol. 11, no. 3, pp. 98-106, 2023...

  17. [25]

    Functional Map of the World,

    G. Christie, N. Fendley, J. Wilson, R. Mukherjee, and Ieee, “Functional Map of the World,” in 2018 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2018, pp. 6172-6180

  18. [26]

    BIGEARTHNET: A LARGE-SCALE BENCHMARK ARCHIVE FOR REMOTE SENSING IMAGE UNDERSTANDING,

    G. Sumbul, M. Charfuelan, B. Demir, V. Markl, and Ieee, “BIGEARTHNET: A LARGE-SCALE BENCHMARK ARCHIVE FOR REMOTE SENSING IMAGE UNDERSTANDING,” in 2019 IEEE INTERNATIONAL GEOSCIENCE AND REMOTE SENSING SYMPOSIUM (IGARSS 2019), 2019, pp. 5901-5904

  19. [27]

    BigEarthNet-MM A large-scale, multimodal, multilabel benchmark archive for remote sensing image classification and retrieval,

    G. Sumbul, A. de Wall, T. Kreuziger, F. Marcelino, H. Costa, P. Benevides, M. Caetano, B. Demir, and V. Markl, “BigEarthNet-MM A large-scale, multimodal, multilabel benchmark archive for remote sensing image classification and retrieval,” IEEE Geosci. Remote Sens. Mag., vol. 9...

  20. [28]

    ImageNet: A Large-Scale Hierarchical Image Database,

    J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, F.-F. Li, and Ieee, “ImageNet: A Large-Scale Hierarchical Image Database,” in CVPR: 2009 IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION, VOLS 1-4, 2009, pp. 248-255

  21. [29]

    A Global Context-aware and Batch-independent Network for road extraction from VHR satellite imagery,

    Q. Zhu, Y. Zhang, L. Wang, Y. Zhong, Q. Guan, X. Lu, L. Zhang, and D. Li, “A Global Context-aware and Batch-independent Network for road extraction from VHR satellite imagery,” ISPRS J. Photogramm. Remote Sens., vol. 175, pp. 353-365. 2021

  22. [30]

    RingMo-SAM: A Foundation Model for Segment Anything in Multimodal Remote-Sensing Images,

    Z. Yan, J. Li, X. Li, R. Zhou, W. Zhang, Y. Feng, W. Diao, K. Fu, and X. Sun, “RingMo-SAM: A Foundation Model for Segment Anything in Multimodal Remote-Sensing Images,” IEEE Trans. Geosci. Remote Sens., vol. 61, pp. 1-16. 2023

  23. [31]

    SDCluster: A clustering based self-supervised pre-training method for semantic segmentation of remote sensing images,

    H. Xu, C. Zhang, P. Yue, and K. Wang, “SDCluster: A clustering based self-supervised pre-training method for semantic segmentation of remote sensing images,” ISPRS J. Photogramm. Remote Sens., vol. 223, pp. 1-14. 2025

  24. [32]

    Self-Supervised Pretraining and Controlled Augmentation Improve Rare Wildlife Recognition in UAV Images,

    X. Zheng, B. Kellenberger, R. Gong, I. Hajnsek, D. Tuia, and I. C. Soc, “Self-Supervised Pretraining and Controlled Augmentation Improve Rare Wildlife Recognition in UAV Images,” in 2021 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION WORKSHOPS (ICCVW 2021), 2021, pp. 732-741

  25. [33]

    Towards Geospatial Foundation Models via Continual Pretraining,

    M. Mendieta, B. Han, X. Shi, Y. Zhu, C. Chen, and Ieee, “Towards Geospatial Foundation Models via Continual Pretraining,” in 2023 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2023), 2023, pp. 16760-16770. 33 33

  26. [34]

    Self-supervised Learning in Remote Sensing: A Review,

    Y. Wang, C. M. Albrecht, N. A. A. Braham, L. Mou, and X. Zhu, “Self-supervised Learning in Remote Sensing: A Review,” Arxiv, 2022. 2022

  27. [35]

    Consecutive Pretraining: A Knowledge Transfer Learning Strategy with Relevant Unlabeled Data for Remote Sensing Domain,

    T. Zhang, P. Gao, H. Dong, Y. Zhuang, G. Wang, W. Zhang, and H. Chen, “Consecutive Pretraining: A Knowledge Transfer Learning Strategy with Relevant Unlabeled Data for Remote Sensing Domain,” Arxiv, 2022. 2022

  28. [36]

    DINOv2: Learning Robust Visual Features without Supervision,

    M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, M. Assran, N. Ballas, W. Galuba, R. Howes, P.-Y. Huang, S.-W. Li, I. Misra, M. Rabbat, V. Sharma, G. Synnaeve, H. Xu, H. Jegou, J. Mairal, P. Labatut, A. Joul...

  29. [37]

    Generative Pretraining from Pixels,

    M. Chen, A. Radford, R. Child, J. Wu, H. Jun, D. Luan, and I. Sutskever, “Generative Pretraining from Pixels,” in INTERNATIONAL CONFERENCE ON MACHINE LEARNING, VOL 119, 2020

  30. [38]

    Momentum Contrast for Unsupervised Visual Representation Learning,

    K. He, H. Fan, Y. Wu, S. Xie, R. Girshick, and Ieee, “Momentum Contrast for Unsupervised Visual Representation Learning,” in 2020 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2020), 2020, pp. 9726-9735

  31. [39]

    iBOT: Image BERT Pre-Training with Online Tokenizer,

    J. Zhou, C. Wei, H. Wang, W. Shen, C. Xie, A. Yuille, and T. J. a. e.-p. Kong, "iBOT: Image BERT Pre-Training with Online Tokenizer," https://ui.adsabs.harvard.edu/abs/2021arXiv211107832Z, [November 01, 2021, 2021]

  32. [40]

    OmniMAE: Single Model Masked Pretraining on Images and Videos,

    R. Girdhar, A. El-Nouby, M. Singh, K. Vasudev Alwala, A. Joulin, and I. J. a. e.-p. Misra, "OmniMAE: Single Model Masked Pretraining on Images and Videos," https://ui.adsabs.harvard.edu/abs/2022arXiv220608356G, [June 01, 2022, 2022]

  33. [41]

    Masked Autoencoders Are Scalable Vision Learners,

    K. He, X. Chen, S. Xie, Y. Li, P. Dollar, R. Girshick, and S. O. C. Ieee Comp, “Masked Autoencoders Are Scalable Vision Learners,” in 2022 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2022), 2022, pp. 15979-15988

  34. [42]

    Masked Autoencoders that Listen,

    P.-Y. Huang, H. Xu, J. Li, A. Baevski, M. Auli, W. Galuba, F. Metze, and C. J. a. e.-p. Feichtenhofer, "Masked Autoencoders that Listen," https://ui.adsabs.harvard.edu/abs/2022arXiv220706405H, [July 01, 2022, 2022]

  35. [43]

    VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training,

    Z. Tong, Y. Song, J. Wang, and L. J. a. e.-p. Wang, "VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training," https://ui.adsabs.harvard.edu/abs/2022arXiv220312602T, [March 01, 2022, 2022]

  36. [44]

    CSP: Self-Supervised Contrastive Spatial Pre-Training for Geospatial-Visual Representations,

    G. C. Mai, N. Lao, Y. T. He, J. M. Song, and S. Ermon, “CSP: Self-Supervised Contrastive Spatial Pre-Training for Geospatial-Visual Representations,” in INTERNATIONAL CONFERENCE ON MACHINE LEARNING, VOL 202, 2023

  37. [45]

    GeoCLIP: Clip-Inspired Alignment between Locations and Images for Effective Worldwide Geo-localization,

    V. V. Cepeda, G. K. Nayak, and M. Shah, “GeoCLIP: Clip-Inspired Alignment between Locations and Images for Effective Worldwide Geo-localization,” in ADVANCES IN NEURAL INFORMATION PROCESSING SYSTEMS 36 (NEURIPS 2023), 2023

  38. [46]

    Geography-Aware Self-Supervised Learning,

    K. Ayush, B. Uzkent, C. Meng, K. Tanmay, M. Burke, D. Lobell, S. Ermon, and Ieee, “Geography-Aware Self-Supervised Learning,” in 2021 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2021), 2021, pp. 10161-10170

  39. [47]

    Change-Aware Sampling and Contrastive Learning for Satellite Images,

    U. Mall, B. Hariharan, K. Bala, and Ieee, “Change-Aware Sampling and Contrastive Learning for Satellite Images,” in 2023 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION, CVPR, 2023, pp. 5261-5270. 34 34

  40. [48]

    Seasonal Contrast: Unsupervised Pre-Training from Uncurated Remote Sensing Data,

    O. Manas, A. Lacoste, X. Giro-i-Nieto, D. Vazquez, P. Rodriguez, and Ieee, “Seasonal Contrast: Unsupervised Pre-Training from Uncurated Remote Sensing Data,” in 2021 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2021), 2021, pp. 9394-9403

  41. [49]

    A Billion-scale Foundation Model for Remote Sensing Images,

    K. Cha, J. Seo, and T. Lee, “A Billion-scale Foundation Model for Remote Sensing Images,” arXiv e-prints, pp. arXiv:2304.05215. 2023

  42. [50]

    RingMo: A Remote Sensing Foundation Model With Masked Image Modeling,

    X. Sun, P. Wang, W. Lu, Z. Zhu, X. Lu, Q. He, J. Li, X. Rong, Z. Yang, H. Chang, Q. He, G. Yang, R. Wang, J. Lu, and K. Fu, “RingMo: A Remote Sensing Foundation Model With Masked Image Modeling,” IEEE Trans. Geosci. Remote Sens., vol. 61, pp. 1-22. 2023

  43. [51]

    SatMAE: Pre-training Transformers for Temporal and Multi-Spectral Satellite Imagery,

    Y. Cong, S. Khanna, C. Meng, P. Liu, E. Rozi, Y. He, M. Burke, D. B. Lobell, and S. Ermon, “SatMAE: Pre-training Transformers for Temporal and Multi-Spectral Satellite Imagery,” Arxiv, 2022. 2022

  44. [52]

    Scale-MAE: A Scale-Aware Masked Autoencoder for Multiscale Geospatial Representation Learning,

    C. J. Reed, R. Gupta, S. Li, S. Brockman, C. Funk, B. Clipp, K. Keutzer, S. Candido, M. Uyttendaele, T. Darrell, and Ieee, “Scale-MAE: A Scale-Aware Masked Autoencoder for Multiscale Geospatial Representation Learning,” in 2023 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VIS...

  45. [53]

    SpectralGPT: Spectral Remote Sensing Foundation Model,

    D. Hong, B. Zhang, X. Li, Y. Li, C. Li, J. Yao, N. Yokoya, H. Li, P. Ghamisi, X. Jia, A. Plaza, P. Gamba, J. A. Benediktsson, and J. Chanussot, “SpectralGPT: Spectral Remote Sensing Foundation Model,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 46, no. 8, pp. 5227-5244, 2024...

  46. [54]

    CMID: A Unified Self-Supervised Learning Framework for Remote Sensing Image Understanding,

    D. Muhtar, X. Zhang, P. Xiao, Z. Li, and F. Gu, “CMID: A Unified Self-Supervised Learning Framework for Remote Sensing Image Understanding,” IEEE Trans. Geosci. Remote Sens., vol. 61, 2023. 2023

  47. [55]

    Contrastive Masked Autoencoders are Stronger Vision Learners,

    Z. Huang, X. Jin, C. Lu, Q. Hou, M.-M. Cheng, D. Fu, X. Shen, and J. Feng, “Contrastive Masked Autoencoders are Stronger Vision Learners,” Arxiv, 2024. 2024

  48. [56]

    CROMA: Remote Sensing Representations with Contrastive Radar-Optical Masked Autoencoders,

    A. Fuller, K. Millard, and J. R. Green, “CROMA: Remote Sensing Representations with Contrastive Radar-Optical Masked Autoencoders,” in ADVANCES IN NEURAL INFORMATION PROCESSING SYSTEMS 36 (NEURIPS 2023), 2023

  49. [57]

    SkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation Imagery,

    X. Guo, J. Lao, B. Dang, Y. Zhang, L. Yu, L. Ru, L. Zhong, Z. Huang, K. Wu, D. Hu, H. He, J. Wang, J. Chen, M. Yang, Y. Zhang, Y. Li, and I. C. Soc, “SkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation Imagery,” in 202...

  50. [58]

    Major TOM: Expandable Datasets for Earth Observation,

    A. Francis, and M. J. a. e.-p. Czerkawski, "Major TOM: Expandable Datasets for Earth Observation," https://ui.adsabs.harvard.edu/abs/2024arXiv240212095F, [February 01, 2024, 2024]

  51. [59]

    On Creating Benchmark Dataset for Aerial Image Interpretation: Reviews, Guidances, and Million-AID,

    Y. Long, G.-S. Xia, S. Li, W. Yang, M. Y. Yang, X. X. Zhu, L. Zhang, and D. Li, “On Creating Benchmark Dataset for Aerial Image Interpretation: Reviews, Guidances, and Million-AID,” IEEE J. Sel. Topics Appl. Earth Observ. Remote Sens., vol. 14, pp. 4205-4230, 2021. 2021

  52. [60]

    An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,” Arxiv, 2021. 2021

  53. [61]

    Bootstrap your own latent: A new approach to self-supervised Learning,

    J.-B. Grill, F. Strub, F. Altché, C. Tallec, P. H. Richemond, E. Buchatskaya, C. Doersch, B. Avila Pires, Z. D. Guo, M. Gheshlaghi Azar, B. Piot, K. Kavukcuoglu, R. Munos, and M. J. a. e.-p. Valko, "Bootstrap your own latent: A new approach to self-supervised Learning," https:...

  54. [62]

    FLASHATTENTION: Fast and Memory-Efficient Exact Attention with IO-Awareness,

    T. Dao, D. Y. Fu, S. Ermon, A. Rudra, and C. Re, “FLASHATTENTION: Fast and Memory-Efficient Exact Attention with IO-Awareness,” in ADVANCES IN NEURAL INFORMATION PROCESSING SYSTEMS 35 (NEURIPS 2022), 2022. 35 35

  55. [63]

    PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel,

    Y. Zhao, A. Gu, R. Varma, L. Luo, C.-C. Huang, M. Xu, L. Wright, H. Shojanazeri, M. Ott, S. Shleifer, A. Desmaison, C. Balioglu, P. Damania, B. Nguyen, G. Chauhan, Y. Hao, A. Mathews, and S. J. a. e.-p. Li, "PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel," ht...

  56. [64]

    Remote Sensing Image Scene Classification: Benchmark and State of the Art,

    G. Cheng, J. Han, and X. J. a. e.-p. Lu, "Remote Sensing Image Scene Classification: Benchmark and State of the Art," https://ui.adsabs.harvard.edu/abs/2017arXiv170300121C, [February 01, 2017, 2017]

  57. [65]

    SatlasPretrain: A Large-Scale Dataset for Remote Sensing Image Understanding,

    F. Bastani, P. Wolters, R. Gupta, J. Ferdinando, A. Kembhavi, and Ieee, “SatlasPretrain: A Large-Scale Dataset for Remote Sensing Image Understanding,” in 2023 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2023), 2023, pp. 16726-+

  58. [66]

    MTP: Advancing Remote Sensing Foundation Model via Multitask Pretraining,

    D. Wang, J. Zhang, M. Xu, L. Liu, D. Wang, E. Gao, C. Han, H. Guo, B. Du, D. Tao, and L. Zhang, “MTP: Advancing Remote Sensing Foundation Model via Multitask Pretraining,” IEEE J. Sel. Topics Appl. Earth Observ. Remote Sens., vol. 17, pp. 11632-11654, January 01, 2024. 2024

  59. [67]

    Object detection in optical remote sensing images: A survey and a new benchmark,

    K. Li, G. Wan, G. Cheng, L. Meng, and J. Han, “Object detection in optical remote sensing images: A survey and a new benchmark,” ISPRS J. Photogramm. Remote Sens., vol. 159, pp. 296-307, 2020 JAN. 2020

  60. [68]

    Anchor-Free Oriented Proposal Generator for Object Detection,

    G. Cheng, J. Wang, K. Li, X. Xie, C. Lang, Y. Yao, and J. Han, “Anchor-Free Oriented Proposal Generator for Object Detection,” IEEE Trans. Geosci. Remote Sens., vol. 60, 2022. 2022

  61. [69]

    LoveDA: A Remote Sensing Land-Cover Dataset for Domain Adaptive Semantic Segmentation,

    J. Wang, Z. Zheng, A. Ma, X. Lu, and Y. J. a. e.-p. Zhong, "LoveDA: A Remote Sensing Land-Cover Dataset for Domain Adaptive Semantic Segmentation," https://ui.adsabs.harvard.edu/abs/2021arXiv211008733W, [October 01, 2021, 2021]

  62. [70]

    iSAID: A Large-scale Dataset for Instance Segmentation in Aerial Images,

    S. Waqas Zamir, A. Arora, A. Gupta, S. Khan, G. Sun, F. Shahbaz Khan, F. Zhu, L. Shao, G.-S. Xia, and X. J. a. e.-p. Bai, "iSAID: A Large-scale Dataset for Instance Segmentation in Aerial Images," https://ui.adsabs.harvard.edu/abs/2019arXiv190512886W, [May 01, 2019, 2019]

  63. [71]

    A Spatial-Temporal Attention-Based Method and a New Dataset for Remote Sensing Image Change Detection,

    H. Chen, and Z. Shi, “A Spatial-Temporal Attention-Based Method and a New Dataset for Remote Sensing Image Change Detection,” REMOTE SENSING, vol. 12, no. 10, 2020 MAY. 2020

  64. [72]

    A Deeply Supervised Attention Metric-Based Network and an Open Aerial Image Dataset for Remote Sensing Change Detection,

    Q. Shi, M. Liu, S. Li, X. Liu, F. Wang, and L. Zhang, “A Deeply Supervised Attention Metric-Based Network and an Open Aerial Image Dataset for Remote Sensing Change Detection,” IEEE Trans. Geosci. Remote Sens., vol. 60, 2022. 2022

  65. [73]

    Change Detection in Remote Sensing Images Using Conditional Adversarial Networks,

    M. A. Lebedev, Y. V. Vizilter, O. V. Vygolov, V. A. Knyaz, and A. Y. Rubis, “Change Detection in Remote Sensing Images Using Conditional Adversarial Networks,” The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences, vol. XLII-2, pp. 5...

  66. [74]

    A Transformer-Based Siamese Network for Change Detection,

    W. G. C. Bandara, and V. M. Patel, “A Transformer-Based Siamese Network for Change Detection,” Arxiv,

  67. [75]

    Remote Sensing Image Change Detection With Transformers,

    H. Chen, Z. Qi, and Z. Shi, “Remote Sensing Image Change Detection With Transformers,” IEEE Trans. Geosci. Remote Sens., vol. 60, 2022. 2022

  68. [76]

    HANet: A Hierarchical Attention Network for Change Detection With Bitemporal Very-High-Resolution Remote Sensing Images,

    C. X. Han, C. Wu, H. A. Guo, M. Q. Hu, and H. R. X. Chen, “HANet: A Hierarchical Attention Network for Change Detection With Bitemporal Very-High-Resolution Remote Sensing Images,” IEEE J. Sel. Topics Appl. Earth Observ. Remote Sens., vol. 16, pp. 3867-3878. 2023

  69. [77]

    Change Guiding Network: Incorporating Change Prior to Guide Change Detection in Remote Sensing Imagery,

    C. X. Han, C. Wu, H. N. Guo, M. Q. Hu, J. P. Li, and H. R. X. Chen, “Change Guiding Network: Incorporating Change Prior to Guide Change Detection in Remote Sensing Imagery,” IEEE J. Sel. Topics Appl. Earth Observ. Remote Sens., vol. 16, pp. 8395-8407. 2023

  70. [78]

    Exchanging Dual-Encoder-Decoder: A New Strategy for Change Detection With Semantic Guidance and Spatial Localization,

    S. J. Zhao, X. L. Zhang, P. F. Xiao, and G. J. He, “Exchanging Dual-Encoder-Decoder: A New Strategy for Change Detection With Semantic Guidance and Spatial Localization,” IEEE Trans. Geosci. Remote Sens., vol. 61. 2023. 36 36

  71. [79]

    C2F-SemiCD: A Coarse-to-Fine Semi-Supervised Change Detection Method Based on Consistency Regularization in High-Resolution Remote Sensing Images,

    C. X. Han, C. Wu, M. Q. Hu, J. P. Li, and H. R. X. Chen, “C2F-SemiCD: A Coarse-to-Fine Semi-Supervised Change Detection Method Based on Consistency Regularization in High-Resolution Remote Sensing Images,” IEEE Trans. Geosci. Remote Sens., vol. 62. 2024

  72. [80]

    MutSimNet: Mutually Reinforcing Similarity Learning for RS Image Change Detection,

    X. Liu, Y. Liu, L. C. Jiao, L. L. Li, F. Liu, S. Y. Yang, and B. Hou, “MutSimNet: Mutually Reinforcing Similarity Learning for RS Image Change Detection,” IEEE Trans. Geosci. Remote Sens., vol. 62. 2024

  73. [81]

    Candidate-Aware and Change-Guided Learning for Remote Sensing Change Detection,

    F. Liu, Y. G. Liu, J. Liu, X. Tang, and L. Xiao, “Candidate-Aware and Change-Guided Learning for Remote Sensing Change Detection,” IEEE Trans. Geosci. Remote Sens., vol. 62. 2024

  74. [82]

    UAV-YOLOv8: A Small-Object-Detection Model Based on Improved YOLOv8 for UAV Aerial Photography Scenarios,

    G. Wang, Y. Chen, P. An, H. Hong, J. Hu, and T. Huang, “UAV-YOLOv8: A Small-Object-Detection Model Based on Improved YOLOv8 for UAV Aerial Photography Scenarios,” SENSORS, vol. 23, no. 16, 2023 AUG. 2023

  75. [83]

    Swin Transformer: Hierarchical Vision Transformer using Shifted Windows,

    Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. J. a. e.-p. Guo, "Swin Transformer: Hierarchical Vision Transformer using Shifted Windows," https://ui.adsabs.harvard.edu/abs/2021arXiv210314030L, [March 01, 2021, 2021]

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.