Pith. sign in

REVIEW 3 cited by

SegEarth-OV: Towards Training-Free Open-Vocabulary Segmentation for Remote Sensing Images

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.01768 v2 pith:4WGEV7YF submitted 2024-10-02 cs.CV

classification cs.CV
keywords remotesensingsegmentationdetectionextensivefeatureshoweverimage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Remote sensing image plays an irreplaceable role in fields such as agriculture, water resources, military, and disaster relief. Pixel-level interpretation is a critical aspect of remote sensing image applications; however, a prevalent limitation remains the need for extensive manual annotation. For this, we try to introduce open-vocabulary semantic segmentation (OVSS) into the remote sensing context. However, due to the sensitivity of remote sensing images to low-resolution features, distorted target shapes and ill-fitting boundaries are exhibited in the prediction mask. To tackle this issue, we propose a simple and general upsampler, SimFeatUp, to restore lost spatial information in deep features in a training-free style. Further, based on the observation of the abnormal response of local patch tokens to [CLS] token in CLIP, we propose to execute a straightforward subtraction operation to alleviate the global bias in patch tokens. Extensive experiments are conducted on 17 remote sensing datasets spanning semantic segmentation, building extraction, road detection, and flood detection tasks. Our method achieves an average of 5.8%, 8.2%, 4.0%, and 15.3% improvement over state-of-the-art methods on 4 tasks. All codes are released. \url{https://earth-insights.github.io/SegEarth-OV}

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RemoteSAM: Towards Segment Anything for Earth Observation

    cs.CV 2025-05 reject novelty 6.0 of 10

    RemoteSAM unifies remote sensing classification, detection, segmentation, and grounding through a single referring expression segmentation model trained on 270K VLM-generated image-text-mask triplets.

  2. SCORE: Scene Context Matters in Open-Vocabulary Remote Sensing Instance Segmentation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    SCORE improves open-vocabulary remote sensing instance segmentation by injecting regional and global scene context from RemoteCLIP into CLIP-based class and text embeddings.

  3. Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives

    cs.CV 2025-05 conditional novelty 3.0 of 10

    A structured review of remote sensing vision-language models, organizing contrastive, instruction-tuned, and generative approaches alongside their datasets and benchmarks.

Pith tools