Pith. sign in

REVIEW 2 cited by

Towards Open-Vocabulary Remote Sensing Image Semantic Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.19492 v1 pith:VDHIM64M submitted 2024-12-27 cs.CV cs.MM

classification cs.CVcs.MM
keywords semanticimageremotesensingclassesmethodsmodelsfeatures
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, deep learning based methods have revolutionized remote sensing image segmentation. However, these methods usually rely on a pre-defined semantic class set, thus needing additional image annotation and model training when adapting to new classes. More importantly, they are unable to segment arbitrary semantic classes. In this work, we introduce Open-Vocabulary Remote Sensing Image Semantic Segmentation (OVRSISS), which aims to segment arbitrary semantic classes in remote sensing images. To address the lack of OVRSISS datasets, we develop LandDiscover50K, a comprehensive dataset of 51,846 images covering 40 diverse semantic classes. In addition, we propose a novel framework named GSNet that integrates domain priors from special remote sensing models and versatile capabilities of general vision-language models. Technically, GSNet consists of a Dual-Stream Image Encoder (DSIE), a Query-Guided Feature Fusion (QGFF), and a Residual Information Preservation Decoder (RIPD). DSIE first captures comprehensive features from both special models and general models in dual streams. Then, with the guidance of variable vocabularies, QGFF integrates specialist and generalist features, enabling them to complement each other. Finally, RIPD is proposed to aggregate multi-source features for more accurate mask predictions. Experiments show that our method outperforms other methods by a large margin, and our proposed LandDiscover50K improves the performance of OVRSISS methods. The proposed dataset and method will be made publicly available at https://github.com/yecy749/GSNet.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Vision-Language Model Purified Semi-Supervised Semantic Segmentation for Remote Sensing Images

    cs.CV 2026-01 reject novelty 5.0 of 10

    A remote-sensing semi-supervised segmentation method that uses a vision-language model to purify and correct teacher-generated pseudo-labels reports large mIoU gains over prior SOTA.

  2. SCORE: Scene Context Matters in Open-Vocabulary Remote Sensing Instance Segmentation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    SCORE improves open-vocabulary remote sensing instance segmentation by injecting regional and global scene context from RemoteCLIP into CLIP-based class and text embeddings.

Pith tools