Pith. sign in

REVIEW 2 cited by

GeoSAM: Fine-tuning SAM with Multi-Modal Prompts for Mobility Infrastructure Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.11319 v4 pith:W35B7K7F submitted 2023-11-19 cs.CV cs.AI

classification cs.CVcs.AI
keywords geosaminfrastructuregeographicalimagesmobilitymodelpromptssegmentation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In geographical image segmentation, performance is often constrained by the limited availability of training data and a lack of generalizability, particularly for segmenting mobility infrastructure such as roads, sidewalks, and crosswalks. Vision foundation models like the Segment Anything Model (SAM), pre-trained on millions of natural images, have demonstrated impressive zero-shot segmentation performance, providing a potential solution. However, SAM struggles with geographical images, such as aerial and satellite imagery, due to its training being confined to natural images and the narrow features and textures of these objects blending into their surroundings. To address these challenges, we propose Geographical SAM (GeoSAM), a SAM-based framework that fine-tunes SAM using automatically generated multi-modal prompts. Specifically, GeoSAM integrates point prompts from a pre-trained task-specific model as primary visual guidance, and text prompts generated by a large language model as secondary semantic guidance, enabling the model to better capture both spatial structure and contextual meaning. GeoSAM outperforms existing approaches for mobility infrastructure segmentation in both familiar and completely unseen regions by at least 5\% in mIoU, representing a significant leap in leveraging foundation models to segment mobility infrastructure, including both road and pedestrian infrastructure in geographical images. The source code can be found in this GitHub Repository: https://github.com/rafiibnsultan/GeoSAM.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Online Location Planning for AI-Defined Vehicles: Optimizing Joint Tasks of Order Serving and Spatio-Temporal Heterogeneous Model Fine-Tuning

    cs.LG 2025-02 conditional novelty 5.0 of 10

    A multi-agent RL framework with graph neural networks lets ride-hailing vehicles jointly serve orders and collect fresh data for foundation model fine-tuning, improving a combined utility metric in simulation.

  2. fabSAM: A Farmland Boundary Delineation Method Based on the Segment Anything Model

    cs.CV 2025-01 conditional novelty 5.0 of 10

    fabSAM couples a Deeplabv3+ prompter with fine-tuned SAM decoders, improving mIOU on AI4Boundaries and AI4SmallFarms over zero-shot SAM and Deeplabv3+ by 4.9 to 23.5 percentage points.

Pith tools