REVIEW 18 cited by
SkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language Model
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Large language models (LLMs) have recently been extended to the vision-language realm, obtaining impressive general multi-modal capabilities. However, the exploration of multi-modal large language models (MLLMs) for remote sensing (RS) data is still in its infancy, and the performance is not satisfactory. In this work, we introduce SkyEyeGPT, a unified multi-modal large language model specifically designed for RS vision-language understanding. To this end, we meticulously curate an RS multi-modal instruction tuning dataset, including single-task and multi-task conversation instructions. After manual verification, we obtain a high-quality RS instruction-following dataset with 968k samples. Our research demonstrates that with a simple yet effective design, SkyEyeGPT works surprisingly well on considerably different tasks without the need for extra encoding modules. Specifically, after projecting RS visual features to the language domain via an alignment layer, they are fed jointly with task-specific instructions into an LLM-based RS decoder to predict answers for RS open-ended tasks. In addition, we design a two-stage tuning method to enhance instruction-following and multi-turn dialogue ability at different granularities. Experiments on 8 datasets for RS vision-language tasks demonstrate SkyEyeGPT's superiority in image-level and region-level tasks, such as captioning and visual grounding. In particular, SkyEyeGPT exhibits encouraging results compared to GPT-4V in some qualitative tests. The online demo, code, and dataset will be released in https://github.com/ZhanYang-nwpu/SkyEyeGPT.
Forward citations
Cited by 18 Pith papers
-
DARAD: Dual Adapters and Ranking-Aware Distillation for Continual Remote Sensing Image-Text Retrieval
A dual-adapter method with ranking-aware distillation preserves historical image-text ranking in continual remote sensing retrieval, achieving positive historical change scores across six semantic stages.
-
Enhancing Remote Sensing Vision-Language Models Through MLLM and LLM-Based High-Quality Image-Text Dataset Generation
A two-stage method (MpGI) produces a 210K-image, 1.26M-caption remote sensing dataset and state-of-the-art CLIP and CoCa models.
-
A Satellite-Ground Synergistic Large Vision-Language Model System for Earth Observation
SpaceVerse jointly decides where to run vision-language inference in LEO satellite networks and compresses task-irrelevant image regions before downlink, improving accuracy and cutting latency versus baselines.
-
VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs
VectorLLM, a multimodal LLM that regresses building contour vertices token by token, reports gains of 5.6 to 13.6 AP over prior polygon extraction methods on WHU, WHU-Mix, and CrowdAI.
-
UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence with Spatial Reasoning and Understanding
A fine-tuned small multimodal LLM outperforms much larger general models on urban tasks in a new benchmark, with caveats about benchmark overlap with training data.
-
LISAT: Language-Instructed Segmentation Assistant for Satellite Imagery
A remote-sensing vision-language model combining LLaVA-style instruction tuning with a SAM decoder, trained on a new semi-synthetic geospatial reasoning-segmentation dataset, GRES.
-
GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing
GeoPixel brings pixel-level grounding to remote sensing large multimodal models for the first time, with a new dataset and benchmark built from iSAID.
-
GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing
GeoPix extends multimodal language models for remote sensing to pixel-level referring segmentation through a mask predictor, a class-wise learnable memory, and a new 65,463-image instruction dataset.
-
EarthDial: Turning Multi-sensory Earth Observations to Interactive Dialogues
EarthDial is a 4B-parameter remote sensing chatbot trained on 11.11M instruction pairs to handle multi-resolution, multi-spectral, and multi-temporal satellite imagery, and it reports gains over prior VLMs on dozens o...
-
RSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mixture of Experts
RSUniVLM is a 1-billion-parameter remote sensing vision-language model that unifies image-, region-, and pixel-level tasks plus multi-image change analysis, achieving state-of-the-art visual grounding on VRSBench and ...
-
GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks
State-of-the-art vision-language models answer only 41.7% of GEOBench-VLM multiple-choice questions correctly, roughly double random chance.
-
Multi-Agent Geospatial Copilots for Remote Sensing Workflows
A hybrid multi-agent orchestrator (composition plus iterative reassessment) reports 60.3% agentic correctness on generated remote sensing workflows, about 17 percentage points above the single-agent GeoLLM-Engine baseline.
-
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models
A single vision-language model is fine-tuned to handle three remote sensing input types and reports state-of-the-art results on VQA, change captioning, and video classification benchmarks.
-
LHRS-Bot-Nova: Improved Multimodal Large Language Model for Remote Sensing Vision-Language Interpretation
A remote sensing chatbot with recaptioned image-text data and a mixture-of-experts visual bridge reports gains over prior general and remote sensing models, but no code or data are released yet.
-
Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes
The authors argue that spatial reasoning in multimodal LLMs requires dedicated new recipes in data, architecture, and training objectives, and will not emerge from scaling alone.
-
REO-VLM: Transforming VLM to Meet Regression Challenges in Earth Observation
A vision-language model for Earth observation that adds numeric biomass regression to generative question answering, with R^2 up to 0.36 and patch counting still failing.
-
Visual Large Language Models for Generalized and Specialized Applications
This paper reviews and taxonomizes VLLM applications into vision-to-text, vision-to-action, and text-to-vision, adding ethics and future-work discussion.
-
Large Vision-Language Models for Remote Sensing Visual Question Answering
A generative vision-language model finetuned on remote-sensing data is claimed to beat baseline models on the RSVQAxBEN benchmark, but the supporting experimental details are missing.
Discussion (0). Continue with ORCID to comment.