Pith. sign in

REVIEW 2 cited by

GeoDecoder: Empowering Multimodal Map Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.15118 v2 pith:XX7WRQEI submitted 2024-01-26 cs.CV cs.AI

classification cs.CVcs.AI
keywords geodecodermodeltasksgeospatialtextimagemarkersmultimodal
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This paper presents GeoDecoder, a dedicated multimodal model designed for processing geospatial information in maps. Built on the BeitGPT architecture, GeoDecoder incorporates specialized expert modules for image and text processing. On the image side, GeoDecoder utilizes GaoDe Amap as the underlying base map, which inherently encompasses essential details about road and building shapes, relative positions, and other attributes. Through the utilization of rendering techniques, the model seamlessly integrates external data and features such as symbol markers, drive trajectories, heatmaps, and user-defined markers, eliminating the need for extra feature engineering. The text module of GeoDecoder accepts various context texts and question prompts, generating text outputs in the style of GPT. Furthermore, the GPT-based model allows for the training and execution of multiple tasks within the same model in an end-to-end manner. To enhance map cognition and enable GeoDecoder to acquire knowledge about the distribution of geographic entities in Beijing, we devised eight fundamental geospatial tasks and conducted pretraining of the model using large-scale text-image samples. Subsequently, rapid fine-tuning was performed on three downstream tasks, resulting in significant performance improvements. The GeoDecoder model demonstrates a comprehensive understanding of map elements and their associated operations, enabling efficient and high-quality application of diverse geospatial tasks in different business scenarios.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GeoBridge: Decoupled Semantic Conditioning for Generative Image Geolocalization

    cs.CV 2026-08 conditional novelty 6.0 of 10

    GeoBridge, a role-decoupled conditioning mechanism, connects a frozen MLLM to a frozen spherical flow-matching head and improves geolocation accuracy at 25/200/750 km on IM2GPS3K.

  2. Towards Interactive Global Geolocation Assistant

    cs.CV 2024-12 conditional novelty 6.0 of 10

    GaGA is an MLLM-based interactive geolocation system that improves country-level accuracy by 4.57% and city-level accuracy by 2.92% over OSV-5M-Baseline on a reproduced GWS15k benchmark.

Pith tools