Pith. sign in

REVIEW 3 cited by

UrbanSAM: Learning Invariance-Inspired Adapters for Segment Anything Models in Urban Construction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.15199 v1 pith:5U5MTGKW submitted 2025-02-21 cs.CV

classification cs.CV
keywords urbanobjectssegmentationurbansamcomplexperformanceaccurateadapter
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Object extraction and segmentation from remote sensing (RS) images is a critical yet challenging task in urban environment monitoring. Urban morphology is inherently complex, with irregular objects of diverse shapes and varying scales. These challenges are amplified by heterogeneity and scale disparities across RS data sources, including sensors, platforms, and modalities, making accurate object segmentation particularly demanding. While the Segment Anything Model (SAM) has shown significant potential in segmenting complex scenes, its performance in handling form-varying objects remains limited due to manual-interactive prompting. To this end, we propose UrbanSAM, a customized version of SAM specifically designed to analyze complex urban environments while tackling scaling effects from remotely sensed observations. Inspired by multi-resolution analysis (MRA) theory, UrbanSAM incorporates a novel learnable prompter equipped with a Uscaling-Adapter that adheres to the invariance criterion, enabling the model to capture multiscale contextual information of objects and adapt to arbitrary scale variations with theoretical guarantees. Furthermore, features from the Uscaling-Adapter and the trunk encoder are aligned through a masked cross-attention operation, allowing the trunk encoder to inherit the adapter's multiscale aggregation capability. This synergy enhances the segmentation performance, resulting in more powerful and accurate outputs, supported by the learned adapter. Extensive experimental results demonstrate the flexibility and superior segmentation performance of the proposed UrbanSAM on a global-scale dataset, encompassing scale-varying urban objects such as buildings, roads, and water.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CLDTracker: A Comprehensive Language Description for Visual Tracking

    cs.CV 2025-05 conditional novelty 6.0 of 10

    CLDTracker improves language-guided visual tracking by building a bag of diverse textual descriptions from CLIP and GPT-4V and updating them across frames, achieving top normalized precision on five benchmarks and bes...

  2. Multimodal Feature Fusion Network with Text Difference Enhancement for Remote Sensing Change Detection

    cs.CV 2025-09 conditional novelty 5.0 of 10

    MMChange fuses image features with VLM-generated text descriptions of bitemporal remote sensing images, reporting state-of-the-art IoU/F1 on LEVIR-CD, WHU-CD, and SYSU-CD.

  3. Hyper-spectral Unmixing algorithms for remote compositional surface mapping: a review of the state of the art

    astro-ph.IM 2025-07 accept

    The paper provides an updated survey of hyper-spectral unmixing methods, publicly available spectral libraries and image datasets, and future research directions such as uncertainty quantification and transfer learning.

Pith tools