Pith. sign in

REVIEW 3 major objections 6 minor 5 cited by

SegEarth-OV claims the first annotation-free open-vocabulary segmentation pipeline for remote sensing, using a trained upsampler and a global-bias subtraction to lift mIoU across 17 optical and 8 SAR benchmarks.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

SegEarth-OV performs annotation-free open-vocabulary segmentation of remote-sensing images by upsampling CLIP features, removing global bias, and distilling optical knowledge into a SAR encoder.

T0 review reviewed 2026-08-05 challenge →

load-bearing objection The optical part is a solid, incremental extension of the authors' own SegEarth-OV; the SAR part, AlignEarth, is the genuinely new contribution but its headline SOTA numbers are undermined by an apparent train/evaluation overlap that must be resolved. the 3 major comments →

arxiv 2508.18067 v1 pith:RFKTNNTU submitted 2025-08-25 cs.CV

Annotation-Free Open-Vocabulary Segmentation for Remote-Sensing Images

classification cs.CV
keywords Semantic segmentationOpen-vocabularyRemote sensing imageVision-language modelAnnotation-freeFeature upsamplingGlobal bias alleviationSAR
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Open-vocabulary image segmentation—labeling every pixel with classes named on the fly—has worked mainly on natural images, where CLIP's coarse features still carry enough spatial detail. This paper argues that remote sensing imagery breaks that assumption: targets vary over orders of magnitude in scale, fine boundaries matter, and a global image token leaks into every local patch token. To fix this, SegEarth-OV trains a universal upsampler (SimFeatUp) once on unlabeled remote sensing images to restore high-resolution spatial detail, and subtracts the global token from each patch token to sharpen local semantics. For SAR imagery, where no large text-paired model exists, AlignEarth distills optical CLIP semantics into a SAR encoder using paired optical-SAR images, avoiding SAR captions entirely. The paper reports large mIoU gains over prior annotation-free methods on 17 optical and 8 SAR benchmarks, bringing open-vocabulary labeling into reach for Earth observation without dense annotations.

Core claim

The paper's central claim is that annotation-free open-vocabulary segmentation of remote sensing images is achievable without task-specific post-training or pixel labels. Three mechanisms carry it. SimFeatUp takes features from an early CLIP block, upsamples them with a single shared parameterized joint bilateral upsampling module run repeatedly, and a content retention network forces the result to reconstruct the input image, preventing semantic drift; one training pass on unlabeled remote sensing images makes the weights universal. Global Bias Alleviation subtracts a scaled copy of the replicated [CLS] token from every patch token (Eq. 9), removing global context that falsely activates unr

What carries the argument

The load-bearing identity is O_hat = O[1:hw+1] − λ O[0]: subtract a fraction (λ=0.3) of the replicated global [CLS] token from every local patch token to remove the global bias that CLIP injects into local features. Around this sits SimFeatUp, a universal upsampler built from a single parameterized joint bilateral upsampling (JBU) module executed repeatedly, guided by spatial distance and feature similarity in the high-resolution RGB image, and regularized by a content retention network that reconstructs the input image from the upsampled features. The third mechanism, AlignEarth, transfers the CLIP semantic space to SAR: a trainable SAR encoder is aligned to a frozen optical encoder by glob

Load-bearing premise

The paper assumes that the global context contaminating each local patch feature is a constant fraction of the image's global token, so subtracting the same fixed amount works for every image, dataset, and sensor.

What would settle it

Measure, per image or per region, the cosine similarity between local patch tokens and the [CLS] token across a scene-diverse remote sensing benchmark, then vary the subtraction strength λ from 0 to 1; if the optimal λ scatters widely rather than clustering near 0.3, the constant-bias assumption is not the operative cause of the reported gains.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • With SimFeatUp and Global Bias Alleviation attached, even the strongest annotation-free OVSS methods gain more: on OpenEarthMap the improvement grows from +3.3 for MaskCLIP to +8.1 for ClearCLIP (Table 6).
  • The upsampler transfers across modalities without retraining: the same SimFeatUp weights process features from the AlignEarth SAR encoder and from optical CLIP (Table 4).
  • AlignEarth removes the need for SAR-specific image-text pairs; paired optical-SAR images alone bring every tested OVSS baseline to a much higher mIoU on SAR data.
  • The full framework reaches 60.2% IoU on flood detection (WBS-SI), suggesting annotation-free OVSS is viable for disaster-relevant water mapping.
  • SimFeatUp also improves MaskCLIP, SCLIP, and ClearCLIP on natural image benchmarks (Table 8), indicating the upsampler is not remote-sensing-specific.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Inference: The fixed subtraction strength of 0.3 implies a single global-correction constant works across 25 datasets; a per-image or per-region adaptively estimated strength is a natural next test and might push gains further.
  • Inference: Because AlignEarth uses only paired optical-SAR images, unpaired SAR corpora and synthetic SAR could be added to improve the distilled encoder without any text annotation.
  • Inference: If the framework generalizes as claimed, pixel-level labeling of novel categories in disaster response or land-cover monitoring could be done on demand from text prompts, with no training labels.
  • Inference: The natural-image gains suggest SimFeatUp may serve as a general feature refiner for any frozen vision-language model, not only remote sensing encoders.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. SegEarth-OV is presented as the first annotation-free open-vocabulary segmentation framework for remote sensing. For optical images, it proposes SimFeatUp, a FeatUp-derived upsampler trained once on 16k images from Million-AID and equipped with an image-reconstruction content-retention loss, earlier-stage feature input, a single shared JBU module, and an 11x11 kernel, plus a Global Bias Alleviation operation that subtracts a scaled [CLS] token from patch tokens. For SAR, AlignEarth distills an optical CLIP encoder into a SAR encoder using paired optical-SAR images and a combination of global contrastive, global distillation, and local region-level distillation losses. The paper reports state-of-the-art mIoU on 8 optical semantic-segmentation datasets, 9 single-class extraction datasets, and 8 SAR datasets, with component ablations and plug-and-play experiments on natural images.

Significance. If the empirical claims hold, this is a practically useful contribution: the optical results are broad (17 datasets), SimFeatUp is evaluated as a plug-and-play module under several baselines, and SimFeatUp's training is independent of the optical evaluation benchmarks. The natural-image results in Table 8 also suggest some generality beyond remote sensing, and the authors state that code and models will be released, which would support reproducibility. However, the SAR results—one of the two headline advances—are currently not trustworthy because of an apparent overlap between AlignEarth's training pairs and the SAR evaluation benchmarks. The fixed lambda in Eq. (9) and the lack of variance reporting are secondary but require attention.

major comments (3)
  1. [§5.1, Table 4] AlignEarth's training pairs are taken from SpaceNet6, MSAW, QXS-SAROPT, SAR2Opt, DFC2023 Track1, PIE-RGB-SAR, DFC2025 Track1, WHU-OPT-SAR, DDHR-SK, FUSAR-Map, and YESeg-OPT-SAR. The SAR test sets listed in §5.1 are PIE-SAR [95], YESeg-SAR [96], FUSAR-Map [97], DDHR-Korea/SD/XA [98], WHU-SAR [99], and OpenEarthMap-SAR [100]. By the cited references, PIE-RGB-SAR is [95], YESeg-OPT-SAR is [96], FUSAR-Map is [97], DDHR-SK is [98], WHU-OPT-SAR is [99], and DFC2025 Track1 is [100]. Thus the SAR encoder is distilled on the same benchmarks that are later used for evaluation in Table 4. No train/test split or exclusion of evaluation images is reported. This means the large gains in Table 4 (e.g., 51.1 vs 12.1 on PIE-SAR) may reflect transductive adaptation to the test distributions, and the comparison to baselines that were not given this adaptation is not apples-to-apples. The authors need to ei
  2. [§3.3, Eq. (9)] Global Bias Alleviation subtracts a fixed fraction lambda=0.3 of the [CLS] token from every patch token. The implementation sets lambda=0.3 for all datasets, but no sensitivity analysis is reported; Table 7 only toggles the operation on/off. Since the bias magnitude may depend on scene content, sensor, or category, the optimum of lambda is not guaranteed to be stable, and the claimed universality of GBA across 25 datasets is not established. Please report a lambda sweep on at least one optical dataset and one SAR dataset, and similarly justify gamma, K, and the JBU kernel size.
  3. [§5.4, Table 7] All quantitative claims are single-run numbers with no variance. Several component contributions in Table 7 are small (0.3–1.7 mIoU), and the 'JBU One' and 'Large Kernel' rows are within likely run-to-run noise for a training-based upsampler. Without multiple seeds or a significance test, the incremental ablation story is not fully supported. This also applies to the modest gains in Table 8. I request variance bars or seed information for the main tables, at least for Tables 4 and 7.
minor comments (6)
  1. [§5.2/Table 1] The phrase 'slight class rename tricks' understates the per-dataset prompt engineering, which includes multiple synonyms and unions. Although the same vocabularies are used for all baselines, the paper should state how these vocabularies were chosen and whether choices were fixed before seeing test labels, to avoid the impression of test-set-specific tuning.
  2. [Conclusion] Typo: 'a annotation-free OVSS method' should be 'an annotation-free OVSS method'.
  3. [§3.2.4] The sentence 'only < 0.3M parameters' is unclear; give the actual parameter counts for FeatUp and SimFeatUp.
  4. [§5.1] The relationship between DFC2025 Track1 and OpenEarthMap-SAR should be explicit; currently both names are used (also in Table 4), making the overlap analysis difficult for the reader.
  5. [Table 3] The 896×896 row is missing for WHUSat.II; explain why or provide the number.
  6. [§5.3.1] The 'Oracle' is described in one sentence as a fully supervised SegFormer with full training data; provide the training protocol and dataset splits to make the upper-bound comparison interpretable.

Circularity Check

1 steps flagged

SAR benchmark comparison is circular: AlignEarth is trained on the same datasets used for SAR evaluation; optical results are independent.

specific steps
  1. fitted input called prediction [Section 5.1 (SAR Remote Sensing Datasets; Optical-SAR Paired Datasets for AlignEarth) and Table 4]
    "SAR Remote Sensing Datasets. To validate AlignEarth and extend SegEarth-OV’s capabilities to SAR images, we conduct experiments on 8 SAR segmentation datasets, including PIE-SAR [95], YESeg-SAR [96], FUSAR-Map [97], DDHR-Korea [98], DDHR-SD [98], DDHR-XA [98], WHU-SAR [99] and OpenEarthMap-SAR [100]. ... To train AlignEarth’s SAR image encoder, we collect a large-scale paired optical-SAR dataset comprising 65k image pairs from various public sources. The datasets include: SpaceNet 6 [73], MSAW [74], QXS-SAROPT [75], SAR2Opt [76], DFC2023 Track1 [102], PIE-RGB-SAR [95], DFC2025 Track1 [100], WH"

    The AlignEarth SAR encoder is trained on paired optical-SAR images drawn from the same published dataset families used as the SAR evaluation benchmarks: PIE-RGB-SAR/PIE-SAR (both [95]), YESeg-OPT-SAR/YESeg-SAR (both [96]), FUSAR-Map (same [97]), DDHR-SK/DDHR-Korea (both [98]), WHU-OPT-SAR/WHU-SAR (both [99]), and DFC2025 Track1/OpenEarthMap-SAR (both [100]). The paper reports no split excluding evaluation tiles, so the large Table 4 gains (e.g., 48.6 vs 16.5 on DDHR-Korea, 51.1 vs 12.1 on PIE-SAR) can reflect transductive adaptation to the test distributions rather than general cross-modal transfer. This is fitting an encoder to the evaluation data and then presenting the outcome as a prediction/generalization, which is circular with respect to the SAR SOTA claim. Baselines are not given t

full rationale

The optical experiments are self-contained against external benchmarks: SimFeatUp is trained on Million-AID independently of the optical evaluation datasets, Global Bias Alleviation is a fixed lambda=0.3 subtraction with no training, and the prompt vocabulary in Table 1 is applied uniformly to all methods. Self-citations to the earlier SegEarth-OV paper [11] are contextual and do not carry the derivation. The central SAR generalization claim, however, is undermined by train/evaluation overlap: Section 5.1 lists the AlignEarth training pairs as coming from PIE-RGB-SAR, YESeg-OPT-SAR, FUSAR-Map, DDHR-SK, WHU-OPT-SAR, and DFC2025 Track1, while Table 4 evaluates on PIE-SAR, YESeg-SAR, FUSAR-Map, DDHR-Korea, WHU-SAR, and OpenEarthMap-SAR — the same dataset families with the same reference IDs. Since no split excluding evaluation tiles is reported, the 'dramatic improvements' on SAR can reflect transductive adaptation to the test distributions, making the SAR SOTA claim a fitted input presented as prediction. This is a partial circularity: the optical claims are independent, but the SAR headline result is not.

Axiom & Free-Parameter Ledger

6 free parameters · 6 axioms · 0 invented entities

No new physical or mathematical entities are introduced. The ledger captures the hand-set hyperparameters (lambda, gamma, K, kernel size, prompt vocabulary) and the domain assumptions that the distilled SAR features are interchangeable with CLIP features and that the global bias is a uniform additive component.

free parameters (6)
  • lambda (global bias intensity) = 0.3 (fixed for all datasets)
    Controls how much of the [CLS] token is subtracted in Eq. (9); chosen by hand, presumably on the evaluated benchmarks; performance depends on this value.
  • gamma (image reconstruction loss weight) = 0.1
    Balances L_img and L_rec in Eq. (5); hand-set.
  • K (local distillation region size) = 7
    Granularity of region-level alignment in Eq. (12); hand-set.
  • JBU kernel size = 11 x 11 (vs 7 x 7 in FeatUp)
    Manually expanded to capture multi-scale RS objects; ablation shows +1.0/+0.2/+0.5 IoU on OpenEarthMap/WHUSat.II/WBS-SI.
  • Per-dataset prompt class names and synonyms = Customized per dataset (Table 1), e.g., 'grass' instead of 'low vegetation' on Potsdam/Vaihingen
    Open-vocabulary scores depend on prompt vocabulary; authors state they use 'class rename tricks for all methods'. This is hand-tuned to the test domains.
  • tau initialization in contrastive loss = 0.07 (learnable)
    Initialization for AlignEarth contrastive loss; standard CLIP value.
axioms (6)
  • domain assumption CLIP's frozen image-text embedding space is semantically meaningful for remote-sensing categories.
    The whole pipeline computes similarity between CLIP patch features and text embeddings of class names; if the domain gap is too large, scores are meaningless. Section 3.1.1.
  • ad hoc to paper Global bias in CLIP patch tokens can be modeled as an additive constant multiple of the [CLS] token.
    Eq. (9) postulates subtracting lambda times the CLS token from all patch tokens; this is the core mechanism of Global Bias Alleviation and is an ad hoc assumption about feature structure.
  • domain assumption Earlier-layer features X[1:hw+1] of CLIP contain more localizable information than final features, and remain compatible with the projection layer.
    Section 3.2.2 chooses Eq. (6) as SimFeatUp input; no formal proof, only empirical abaltion.
  • domain assumption Paired optical-SAR images provide a sufficient bridge for semantic knowledge transfer, despite imperfect co-registration.
    AlignEarth (Section 4) uses region averaging to tolerate misalignment; assumes residual error is zero-mean.
  • domain assumption A SAR encoder trained by distillation into the CLIP space can be interchanged with the CLIP image encoder for SimFeatUp and GBA without retraining.
    The paper claims SimFeatUp directly processes Esar features (Section 4.2); this assumes the distilled feature distribution matches CLIP's.
  • domain assumption Million-AID raw images are a sufficient training distribution for a universal upsampler across all test RS images and modalities.
    SimFeatUp is trained on 16k random Million-AID images (Section 5.1); assumes coverage of objects and scales in all evaluation sets.

reviewed 2026-08-05 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Annotation-Free Open-Vocabulary Segmentation for Remote-Sensing Images." pith.science (2026). https://pith.science/paper/RFKTNNTU

@misc{pith2026250818067,
  author       = {Pith},
  title        = {Pith review of: Annotation-Free Open-Vocabulary Segmentation for Remote-Sensing Images},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/RFKTNNTU}},
  note         = {Machine review of arXiv:2508.18067}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Semantic segmentation of remote sensing (RS) images is pivotal for comprehensive Earth observation, but the demand for interpreting new object categories, coupled with the high expense of manual annotation, poses significant challenges. Although open-vocabulary semantic segmentation (OVSS) offers a promising solution, existing frameworks designed for natural images are insufficient for the unique complexities of RS data. They struggle with vast scale variations and fine-grained details, and their adaptation often relies on extensive, costly annotations. To address this critical gap, this paper introduces SegEarth-OV, the first framework for annotation-free open-vocabulary segmentation of RS images. Specifically, we propose SimFeatUp, a universal upsampler that robustly restores high-resolution spatial details from coarse features, correcting distorted target shapes without any task-specific post-training. We also present a simple yet effective Global Bias Alleviation operation to subtract the inherent global context from patch features, significantly enhancing local semantic fidelity. These components empower SegEarth-OV to effectively harness the rich semantics of pre-trained VLMs, making OVSS possible in optical RS contexts. Furthermore, to extend the framework's universality to other challenging RS modalities like SAR images, where large-scale VLMs are unavailable and expensive to create, we introduce AlignEarth, which is a distillation-based strategy and can efficiently transfer semantic knowledge from an optical VLM encoder to an SAR encoder, bypassing the need to build SAR foundation models from scratch and enabling universal OVSS across diverse sensor types. Extensive experiments on both optical and SAR datasets validate that SegEarth-OV can achieve dramatic improvements over the SOTA methods, establishing a robust foundation for annotation-free and open-world Earth observation.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Towards Realistic Open-Vocabulary Remote Sensing Segmentation: Benchmark and Baseline

    cs.CV 2026-04 unverdicted novelty 7.0

    OVRSISBenchV2 is a realistic benchmark expanding scene and category coverage for open-vocabulary remote sensing segmentation, with Pi-Seg baseline showing strong transfer via positive-incentive noise perturbations.

  2. Towards Realistic Open-Vocabulary Remote Sensing Segmentation: Benchmark and Baseline

    cs.CV 2026-04 unverdicted novelty 5.5

    OVRSISBenchV2 expands open-vocabulary remote-sensing segmentation evaluation to 170K images and 128 categories, and Pi-Seg uses positive-incentive noise to improve transfer on that harder benchmark.

  3. Prompt-Calibrated SAM 3 for Open-Vocabulary Remote Sensing Semantic Segmentation

    cs.CV 2026-06 unverdicted novelty 5.0

    ProC-SAM3 calibrates SAM 3 prompts for remote sensing OVSS via offline MLLM-generated pools with constraints, cached embeddings, and presence-guided residual fusion, reaching 56.1% average mIoU on eight benchmarks (3....

  4. SegEarth-OV3: Exploring SAM 3 for Open-Vocabulary Semantic Segmentation in Remote Sensing Images

    cs.CV 2025-12 unverdicted novelty 5.0

    SAM 3 can be applied training-free to remote sensing open-vocabulary segmentation and change detection by fusing its semantic and instance heads and filtering with presence scores.

  5. Prompt-Calibrated SAM 3 for Open-Vocabulary Remote Sensing Semantic Segmentation

    cs.CV 2026-06 conditional novelty 4.0

    ProC-SAM3 builds a curated per-dataset prompt pool with an MLLM, caches text embeddings, and fuses SAM 3 mask outputs through a presence gate, reaching 56.1% average mIoU on eight remote-sensing benchmarks.

Reference graph

Works this paper leans on

110 extracted references · 55 canonical work pages · cited by 3 Pith papers

  1. [1]

    Learning to holistically detect bridges from large-size vhr remote sensing im- agery,

    Y. Li, J. Luo, Y. Zhang, Y. Tan, J.-G. Yu, and S. Bai, “Learning to holistically detect bridges from large-size vhr remote sensing im- agery,” IEEE transactions on pattern analysis and machine intelligence , vol. 46, no. 12, pp. 11 507–11 523, 2024. 14

  2. [2]

    Hypersigma: Hyperspectral intelligence comprehension foundation model,

    D. Wang, M. Hu, Y. Jin, Y. Miao, J. Yang, Y. Xu, X. Qin, J. Ma, L. Sun, C. Li et al. , “Hypersigma: Hyperspectral intelligence comprehension foundation model,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025

  3. [3]

    Fully convolutional change detection framework with generative adversarial network for unsupervised, weakly supervised and regional supervised change detection,

    C. Wu, B. Du, and L. Zhang, “Fully convolutional change detection framework with generative adversarial network for unsupervised, weakly supervised and regional supervised change detection,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 45, no. 8, pp. 9774–9788, 2023

  4. [4]

    Farseg++: Foreground-aware relation network for geospatial object segmen- tation in high spatial resolution remote sensing imagery,

    Z. Zheng, Y. Zhong, J. Wang, A. Ma, and L. Zhang, “Farseg++: Foreground-aware relation network for geospatial object segmen- tation in high spatial resolution remote sensing imagery,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023

  5. [5]

    Segearth-r1: Geospatial pixel reasoning via large language model,

    K. Li, Z. Xin, L. Pang, C. Pang, Y. Deng, J. Yao, G. Xia, D. Meng, Z. Wang, and X. Cao, “Segearth-r1: Geospatial pixel reasoning via large language model,” arXiv preprint arXiv:2504.09644, 2025

  6. [6]

    Learning transferable visual models from natural language supervision,

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agar- wal, G. Sastry, A. Askell, P . Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International conference on machine learning . PMLR, 2021, pp. 8748–8763

  7. [7]

    Open vocabulary semantic segmentation with patch aligned contrastive learning,

    J. Mukhoti, T.-Y. Lin, O. Poursaeed, R. Wang, A. Shah, P . H. Torr, and S.-N. Lim, “Open vocabulary semantic segmentation with patch aligned contrastive learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 19 413–19 423

  8. [8]

    Perceptual grouping in contrastive vision-language models,

    K. Ranasinghe, B. McKinzie, S. Ravi, Y. Yang, A. Toshev, and J. Shlens, “Perceptual grouping in contrastive vision-language models,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 5571–5584

  9. [9]

    Sclip: Rethinking self-attention for dense vision-language inference,

    F. Wang, J. Mei, and A. Yuille, “Sclip: Rethinking self-attention for dense vision-language inference,” arXiv preprint arXiv:2312.01597, 2023

  10. [10]

    Sim´eoni, H

    O. Sim´eoni, H. V . Vo, M. Seitzer, F. Baldassarre, M. Oquab, C. Jose, V . Khalidov, M. Szafraniec, S. Yi, M. Ramamonjisoaet al., “Dinov3,” arXiv preprint arXiv:2508.10104, 2025

  11. [11]

    Segearth-ov: Towards training-free open-vocabulary segmenta- tion for remote sensing images,

    K. Li, R. Liu, X. Cao, X. Bai, F. Zhou, D. Meng, and Z. Wang, “Segearth-ov: Towards training-free open-vocabulary segmenta- tion for remote sensing images,” in Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 10 545–10 556

  12. [12]

    Dynamicearth: How far are we from open-vocabulary change detection?

    K. Li, X. Cao, Y. Deng, C. Pang, Z. Xin, D. Meng, and Z. Wang, “Dynamicearth: How far are we from open-vocabulary change detection?” arXiv preprint arXiv:2501.12931, 2025

  13. [13]

    Special: Zero-shot hyperspectral image classification with clip,

    L. Pang, J. Yao, K. Li, and X. Cao, “Special: Zero-shot hyperspectral image classification with clip,” arXiv preprint arXiv:2501.16222 , 2025

  14. [14]

    Earthmind: Towards multi-granular and multi-sensor earth observation with large multimodal models,

    Y. Shu, B. Ren, Z. Xiong, D. P . Paudel, L. Van Gool, B. Demir, N. Sebe, and P . Rota, “Earthmind: Towards multi-granular and multi-sensor earth observation with large multimodal models,” arXiv preprint arXiv:2506.01667, 2025

  15. [15]

    Towards open vocabulary learning: A survey,

    J. Wu, X. Li, S. Xu, H. Yuan, H. Ding, Y. Yang, X. Li, J. Zhang, Y. Tong, X. Jiang et al. , “Towards open vocabulary learning: A survey,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024

  16. [16]

    A survey on open-vocabulary detection and segmentation: Past, present, and future,

    C. Zhu and L. Chen, “A survey on open-vocabulary detection and segmentation: Past, present, and future,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024

  17. [17]

    Reproducible scaling laws for contrastive language-image learning,

    M. Cherti, R. Beaumont, R. Wightman, M. Wortsman, G. Ilharco, C. Gordon, C. Schuhmann, L. Schmidt, and J. Jitsev, “Reproducible scaling laws for contrastive language-image learning,” in Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 2818–2829

  18. [18]

    Alip: Adaptive language-image pre-training with synthetic cap- tion,

    K. Yang, J. Deng, X. An, J. Li, Z. Feng, J. Guo, J. Yang, and T. Liu, “Alip: Adaptive language-image pre-training with synthetic cap- tion,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 2922–2931

  19. [19]

    Demystifying clip data,

    H. Xu, S. Xie, X. E. Tan, P .-Y. Huang, R. Howes, V . Sharma, S.-W. Li, G. Ghosh, L. Zettlemoyer, and C. Feichtenhofer, “Demystifying clip data,” in The Twelfth International Conference on Learning Representations, 2024

  20. [20]

    Is synthetic data all we need? benchmarking the robustness of models trained with synthetic images,

    K. Singh, T. Navaratnam, J. Holmer, S. Schaub-Meyer, and S. Roth, “Is synthetic data all we need? benchmarking the robustness of models trained with synthetic images,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 2505–2515

  21. [21]

    Scaling language-image pre-training via masking,

    Y. Li, H. Fan, R. Hu, C. Feichtenhofer, and K. He, “Scaling language-image pre-training via masking,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 23 390–23 400

  22. [22]

    Improving clip training with language rewrites,

    L. Fan, D. Krishnan, P . Isola, D. Katabi, and Y. Tian, “Improving clip training with language rewrites,” Advances in Neural Informa- tion Processing Systems, vol. 36, 2024

  23. [23]

    Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation,

    J. Li, D. Li, C. Xiong, and S. Hoi, “Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation,” in International conference on machine learning. PMLR, 2022, pp. 12 888–12 900

  24. [24]

    Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,

    J. Li, D. Li, S. Savarese, and S. Hoi, “Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,” in International conference on machine learning. PMLR, 2023, pp. 19 730–19 742

  25. [25]

    Remoteclip: A vision language foundation model for remote sensing,

    F. Liu, D. Chen, Z. Guan, X. Zhou, J. Zhu, Q. Ye, L. Fu, and J. Zhou, “Remoteclip: A vision language foundation model for remote sensing,” IEEE Transactions on Geoscience and Remote Sensing, 2024

  26. [26]

    Rs5m: A large scale vision- language dataset for remote sensing vision-language foundation model,

    Z. Zhang, T. Zhao, Y. Guo, and J. Yin, “Rs5m: A large scale vision- language dataset for remote sensing vision-language foundation model,” arXiv preprint arXiv:2306.11300, 2023

  27. [27]

    Skyscript: A large and semantically diverse vision-language dataset for remote sensing,

    Z. Wang, R. Prabha, T. Huang, J. Wu, and R. Rajagopal, “Skyscript: A large and semantically diverse vision-language dataset for remote sensing,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 6, 2024, pp. 5805–5813

  28. [28]

    H2rsvlm: Towards helpful and honest remote sensing large vision language model,

    C. Pang, J. Wu, J. Li, Y. Liu, J. Sun, W. Li, X. Weng, S. Wang, L. Feng, G.-S. Xia et al., “H2rsvlm: Towards helpful and honest remote sensing large vision language model,” arXiv preprint arXiv:2403.20213, 2024

  29. [29]

    Spectralgpt: Spectral remote sensing foundation model,

    D. Hong, B. Zhang, X. Li, Y. Li, C. Li, J. Yao, N. Yokoya, H. Li, P . Ghamisi, X. Jia et al. , “Spectralgpt: Spectral remote sensing foundation model,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024

  30. [30]

    Hsigene: A foundation model for hyperspectral image generation,

    L. Pang, D. Tang, S. Xu, D. Meng, and X. Cao, “Hsigene: A foundation model for hyperspectral image generation,” arXiv preprint arXiv:2409.12470, 2024

  31. [31]

    Joint bilateral upsampling,

    J. Kopf, M. F. Cohen, D. Lischinski, and M. Uyttendaele, “Joint bilateral upsampling,” ACM Transactions on Graphics (ToG), vol. 26, no. 3, pp. 96–es, 2007

  32. [32]

    U-net: Convolutional networks for biomedical image segmentation,

    O. Ronneberger, P . Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in International Conference on Medical image computing and computer-assisted inter- vention. Springer, 2015, pp. 234–241

  33. [33]

    Unified perceptual parsing for scene understanding,

    T. Xiao, Y. Liu, B. Zhou, Y. Jiang, and J. Sun, “Unified perceptual parsing for scene understanding,” in Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 418–434

  34. [34]

    Panoptic feature pyramid networks,

    A. Kirillov, R. Girshick, K. He, and P . Doll ´ar, “Panoptic feature pyramid networks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 6399–6408

  35. [35]

    Per-pixel classification is not all you need for semantic segmentation,

    B. Cheng, A. Schwing, and A. Kirillov, “Per-pixel classification is not all you need for semantic segmentation,” Advances in Neural Information Processing Systems, vol. 34, pp. 17 864–17 875, 2021

  36. [36]

    Learning to upsample by learning to sample,

    W. Liu, H. Lu, H. Fu, and Z. Cao, “Learning to upsample by learning to sample,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 6027–6037

  37. [37]

    Sapa: Similarity- aware point affiliation for feature upsampling,

    H. Lu, W. Liu, Z. Ye, H. Fu, Y. Liu, and Z. Cao, “Sapa: Similarity- aware point affiliation for feature upsampling,” Advances in Neural Information Processing Systems, vol. 35, pp. 20 889–20 901, 2022

  38. [38]

    A refreshed similarity-based upsampler for direct high-ratio feature upsam- pling,

    M. Zhou, H. Wang, Y. Zheng, and D. Meng, “A refreshed similarity-based upsampler for direct high-ratio feature upsam- pling,” arXiv preprint arXiv:2407.02283, 2024

  39. [39]

    Featup: A model-agnostic framework for features at any resolution,

    S. Fu, M. Hamilton, L. E. Brandt, A. Feldmann, Z. Zhang, and W. T. Freeman, “Featup: A model-agnostic framework for features at any resolution,” in The Twelfth International Conference on Learning Representations , 2024. [Online]. Available: https://openreview.net/forum?id=GkJiNn2QDF

  40. [40]

    Nerf: Representing scenes as neural radiance fields for view synthesis,

    B. Mildenhall, P . P . Srinivasan, M. Tancik, J. T. Barron, R. Ra- mamoorthi, and R. Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,” Communications of the ACM , vol. 65, no. 1, pp. 99–106, 2021

  41. [41]

    Foreground-aware relation network for geospatial object segmentation in high spatial resolution remote sensing imagery,

    Z. Zheng, Y. Zhong, J. Wang, and A. Ma, “Foreground-aware relation network for geospatial object segmentation in high spatial resolution remote sensing imagery,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2020, pp. 4096– 4105

  42. [42]

    Rotated multi-scale interaction network for referring remote sensing image segmentation,

    S. Liu, Y. Ma, X. Zhang, H. Wang, J. Ji, X. Sun, and R. Ji, “Rotated multi-scale interaction network for referring remote sensing image segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 26 658–26 668. 15

  43. [43]

    Loveda: A remote sensing land-cover dataset for domain adaptive semantic segmentation,

    J. Wang, Z. Zheng, A. Ma, X. Lu, and Y. Zhong, “Loveda: A remote sensing land-cover dataset for domain adaptive semantic segmentation,” in Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, J. Vanschoren and S. Yeung, Eds., vol. 1. Curran Associates, Inc., 2021. [Online]. Available: https://datasets-benchmarks-proce...

  44. [44]

    Samrs: Scaling-up remote sensing segmentation dataset with seg- ment anything model,

    D. Wang, J. Zhang, B. Du, M. Xu, L. Liu, D. Tao, and L. Zhang, “Samrs: Scaling-up remote sensing segmentation dataset with seg- ment anything model,” Advances in Neural Information Processing Systems, vol. 36, pp. 8815–8827, 2023

  45. [45]

    Flair: a country-scale land cover semantic segmentation dataset from multi-source optical imagery,

    A. Garioud, N. Gonthier, L. Landrieu, A. De Wit, M. Valette, M. Poup´ee, S. Giordano et al., “Flair: a country-scale land cover semantic segmentation dataset from multi-source optical imagery,” Advances in Neural Information Processing Systems , vol. 36, pp. 16 456–16 482, 2023

  46. [46]

    Weakly supervised semantic segmentation with consistency- constrained multiclass attention for remote sensing scenes,

    J. Zhang, Q. Zhang, Y. Gong, J. Zhang, L. Chen, and D. Zeng, “Weakly supervised semantic segmentation with consistency- constrained multiclass attention for remote sensing scenes,” IEEE Transactions on Geoscience and Remote Sensing , vol. 62, pp. 1–18, 2024

  47. [47]

    Class similarity transition: Decoupling class similarities and imbalance from generalized few-shot segmentation,

    S. Wang, R. Liu, K. Li, J. Jiang, and X. Cao, “Class similarity transition: Decoupling class similarities and imbalance from generalized few-shot segmentation,” in 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) . IEEE, 2024, pp. 2762–2770

  48. [48]

    Generalized few-shot meets remote sensing: Discovering novel classes in land cover mapping via hybrid semantic segmentation framework,

    Z. Li, F. Lu, J. Zou, L. Hu, and H. Zhang, “Generalized few-shot meets remote sensing: Discovering novel classes in land cover mapping via hybrid semantic segmentation framework,” in 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). IEEE, 2024, pp. 2744–2754

  49. [49]

    Scaling open-vocabulary image segmentation with image-level labels,

    G. Ghiasi, X. Gu, Y. Cui, and T.-Y. Lin, “Scaling open-vocabulary image segmentation with image-level labels,” in European Confer- ence on Computer Vision. Springer, 2022, pp. 540–557

  50. [50]

    Segclip: Patch aggregation with learnable centers for open-vocabulary semantic segmenta- tion,

    H. Luo, J. Bao, Y. Wu, X. He, and T. Li, “Segclip: Patch aggregation with learnable centers for open-vocabulary semantic segmenta- tion,” in International Conference on Machine Learning. PMLR, 2023, pp. 23 033–23 044

  51. [51]

    Clip- self: Vision transformer distills itself for open-vocabulary dense prediction,

    S. Wu, W. Zhang, L. Xu, S. Jin, X. Li, W. Liu, and C. C. Loy, “Clip- self: Vision transformer distills itself for open-vocabulary dense prediction,” in The Twelfth International Conference on Learning Representations, 2024

  52. [52]

    Language-driven semantic segmentation,

    B. Li, K. Q. Weinberger, S. Belongie, V . Koltun, and R. Ran- ftl, “Language-driven semantic segmentation,” arXiv preprint arXiv:2201.03546, 2022

  53. [53]

    Decoupling zero-shot semantic segmentation,

    J. Ding, N. Xue, G.-S. Xia, and D. Dai, “Decoupling zero-shot semantic segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 11 583–11 592

  54. [54]

    Side adapter network for open-vocabulary semantic segmentation,

    M. Xu, Z. Zhang, F. Wei, H. Hu, and X. Bai, “Side adapter network for open-vocabulary semantic segmentation,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 2945–2954

  55. [55]

    San: Side adapter network for open-vocabulary semantic segmentation,

    ——, “San: Side adapter network for open-vocabulary semantic segmentation,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023

  56. [56]

    Cat-seg: Cost aggregation for open-vocabulary semantic segmentation,

    S. Cho, H. Shin, S. Hong, A. Arnab, P . H. Seo, and S. Kim, “Cat-seg: Cost aggregation for open-vocabulary semantic segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 4113–4123

  57. [57]

    Open-vocabulary segmentation with semantic-assisted calibration,

    Y. Liu, S. Bai, G. Li, Y. Wang, and Y. Tang, “Open-vocabulary segmentation with semantic-assisted calibration,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 3491–3500

  58. [58]

    Extract free dense labels from clip,

    C. Zhou, C. C. Loy, and B. Dai, “Extract free dense labels from clip,” in European Conference on Computer Vision. Springer, 2022, pp. 696–712

  59. [59]

    Clip surgery for better explainability with enhancement in open-vocabulary tasks,

    Y. Li, H. Wang, Y. Duan, and X. Li, “Clip surgery for better explainability with enhancement in open-vocabulary tasks,” arXiv preprint arXiv:2304.05653, 2023

  60. [60]

    Ground- ing everything: Emerging localization properties in vision- language transformers,

    W. Bousselham, F. Petersen, V . Ferrari, and H. Kuehne, “Ground- ing everything: Emerging localization properties in vision- language transformers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 3828–3837

  61. [61]

    Clearclip: Decomposing clip representations for dense vision- language inference,

    M. Lan, C. Chen, Y. Ke, X. Wang, L. Feng, and W. Zhang, “Clearclip: Decomposing clip representations for dense vision- language inference,” arXiv preprint arXiv:2407.12442, 2024

  62. [62]

    Explore the potential of clip for training-free open vocabulary semantic segmentation,

    T. Shao, Z. Tian, H. Zhao, and J. Su, “Explore the potential of clip for training-free open vocabulary semantic segmentation,” in European Conference on Computer Vision. Springer, 2024

  63. [63]

    In defense of lazy visual grounding for open-vocabulary semantic segmentation,

    D. Kang and M. Cho, “In defense of lazy visual grounding for open-vocabulary semantic segmentation,” in European Conference on Computer Vision and Pattern Recognition (ECCV), 2024

  64. [64]

    Clip as rnn: Segment count- less visual concepts without training endeavor,

    S. Sun, R. Li, P . Torr, X. Gu, and S. Li, “Clip as rnn: Segment count- less visual concepts without training endeavor,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 13 171–13 182

  65. [65]

    Training-free open-vocabulary segmentation with offline diffusion-augmented prototype generation,

    L. Barsellotti, R. Amoroso, M. Cornia, L. Baraldi, and R. Cuc- chiara, “Training-free open-vocabulary segmentation with offline diffusion-augmented prototype generation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 3689–3698

  66. [66]

    Segment anything,

    A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo et al. , “Segment anything,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 4015–4026

  67. [67]

    High-resolution image synthesis with latent diffusion models,

    R. Rombach, A. Blattmann, D. Lorenz, P . Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” 2021

  68. [68]

    Proxyclip: Proxy attention improves clip for open-vocabulary seg- mentation,

    M. Lan, C. Chen, Y. Ke, X. Wang, L. Feng, and W. Zhang, “Proxyclip: Proxy attention improves clip for open-vocabulary seg- mentation,” in European Conference on Computer Vision. Springer, 2025, pp. 70–88

  69. [69]

    Diffusion model is secretly a training-free open vocabulary semantic segmenter,

    J. Wang, X. Li, J. Zhang, Q. Xu, Q. Zhou, Q. Yu, L. Sheng, and D. Xu, “Diffusion model is secretly a training-free open vocabulary semantic segmenter,” arXiv preprint arXiv:2309.02773, 2023

  70. [70]

    Open-vocabulary re- mote sensing image semantic segmentation,

    Q. Cao, Y. Chen, C. Ma, and X. Yang, “Open-vocabulary re- mote sensing image semantic segmentation,” arXiv preprint arXiv:2409.07683, 2024

  71. [71]

    Towards open-vocabulary remote sensing image semantic segmentation,

    C. Ye, Y. Zhuge, and P . Zhang, “Towards open-vocabulary remote sensing image semantic segmentation,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 39, no. 9, 2025, pp. 9436– 9444

  72. [72]

    Skysense-o: Towards open-world remote sensing interpretation with vision-centric visual-language modeling,

    Q. Zhu, J. Lao, D. Ji, J. Luo, K. Wu, Y. Zhang, L. Ru, J. Wang, J. Chen, M. Yang et al., “Skysense-o: Towards open-world remote sensing interpretation with vision-centric visual-language modeling,” in Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 14 733–14 744

  73. [73]

    Spacenet 6: Multi-sensor all weather mapping dataset,

    J. Shermeyer, D. Hogan, J. Brown, A. Van Etten, N. Weir, F. Pacifici, R. Hansch, A. Bastidas, S. Soenen, T. Bacastow et al., “Spacenet 6: Multi-sensor all weather mapping dataset,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops, 2020, pp. 196–197

  74. [74]

    Multi-resolution sar and optical remote sensing image registration methods: A review, datasets, and future perspectives,

    W. Zhang, R. Zhao, Y. Yao, Y. Wan, P . Wu, J. Li, Y. Li, and Y. Zhang, “Multi-resolution sar and optical remote sensing image registration methods: A review, datasets, and future perspectives,” arXiv preprint arXiv:2502.01002, 2025

  75. [75]

    The qxs-saropt dataset for deep learning in sar-optical data fusion,

    M. Huang, Y. Xu, L. Qian, W. Shi, Y. Zhang, W. Bao, N. Wang, X. Liu, and X. Xiang, “The qxs-saropt dataset for deep learning in sar-optical data fusion,” arXiv preprint arXiv:2103.08259, 2021

  76. [76]

    A comparative analysis of gan-based methods for sar-to-optical image translation,

    Y. Zhao, T. Celik, N. Liu, and H.-C. Li, “A comparative analysis of gan-based methods for sar-to-optical image translation,” IEEE Geoscience and Remote Sensing Letters, vol. 19, pp. 1–5, 2022

  77. [77]

    Emerging properties in self-supervised vision transformers,

    M. Caron, H. Touvron, I. Misra, H. J ´egou, J. Mairal, P . Bojanowski, and A. Joulin, “Emerging properties in self-supervised vision transformers,” in Proceedings of the IEEE/CVF international confer- ence on computer vision, 2021, pp. 9650–9660

  78. [78]

    Dinov2: Learning robust visual features without supervision,

    M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V . Khalidov, P . Fernandez, D. Haziza, F. Massa, A. El-Noubyet al., “Dinov2: Learning robust visual features without supervision,” arXiv preprint arXiv:2304.07193, 2023

  79. [79]

    Auto-encoding variational bayes,

    D. P . Kingma, “Auto-encoding variational bayes,”arXiv preprint arXiv:1312.6114, 2013

  80. [80]

    Understanding intermediate layers using linear classi- fier probes,

    G. Alain, “Understanding intermediate layers using linear classi- fier probes,” arXiv preprint arXiv:1610.01644, 2016

Showing first 80 references.

This paper was first reviewed by deepseek-v4-flash on August 5, 2026.