REVIEW 25 cited by
SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Remote Sensing Large Multi-Modal Models (RSLMMs) are developing rapidly and showcase significant capabilities in remote sensing imagery (RSI) comprehension. However, due to the limitations of existing datasets, RSLMMs have shortcomings in understanding the rich semantic relations among objects in complex remote sensing scenes. To unlock RSLMMs' complex comprehension ability, we propose a large-scale instruction tuning dataset FIT-RS, containing 1,800,851 instruction samples. FIT-RS covers common interpretation tasks and innovatively introduces several complex comprehension tasks of escalating difficulty, ranging from relation reasoning to image-level scene graph generation. Based on FIT-RS, we build the FIT-RSFG benchmark. Furthermore, we establish a new benchmark to evaluate the fine-grained relation comprehension capabilities of LMMs, named FIT-RSRC. Based on combined instruction data, we propose SkySenseGPT, which achieves outstanding performance on both public datasets and FIT-RSFG, surpassing existing RSLMMs. We hope the FIT-RS dataset can enhance the relation comprehension capability of RSLMMs and provide a large-scale fine-grained data source for the remote sensing community. The dataset will be available at https://github.com/Luo-Z13/SkySenseGPT
Forward citations
Cited by 25 Pith papers
-
TESSERA v2: Scaling Pixel-wise Earth Foundation Models
Downstream-driven scaling of pixel-wise Barlow Twins EO models favors large encoders and matched data over projectors, and distillation yields compact Matryoshka students that lead multi-task embedding benchmarks.
-
NeSy-Route: A Neuro-Symbolic Benchmark for Constrained Route Planning in Remote Sensing
NeSy-Route supplies 10,821 optimally labeled remote-sensing route-planning tasks plus a three-level neuro-symbolic protocol that reveals major perception and planning deficits in current MLLMs.
-
Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing
Training on multi-tool visual reasoning trajectories (zoom, grounding, lines) with an attention-focused RL objective improves UHR remote-sensing VQA accuracy over single-tool zoom-in and larger base models.
-
Filling Before Advancing: Capability-Gap-Driven Post-Training for Scenario-Specialized Remote Sensing MLLMs
An ordered three-stage post-training route (FBA) improves harbor-scenario performance of RS-MLLMs over direct and collapsed fine-tuning on the authors' HarborEval benchmark.
-
Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?
Under one zero-shot protocol, general-purpose MLLMs match or outperform remote-sensing-specific MLLMs on several RS benchmarks, while RS-MLLMs keep advantages in visual grounding, RS-VQA, and ultra-high-resolution und...
-
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos
A new pixel-level aerial streaming referring-segmentation dataset (DroneEyes) and an MLLM (SkyAnchor) with learned token routing and hierarchical memory report SOTA on DroneEyes and large zero-shot gains on SkyFind.
-
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing
SkyVLaM introduces a temporal basis perceiver and adaptive dense selection to improve language-conditioned video segmentation in UAV scenes, and contributes the SkyVid dataset.
-
GeoChrono: Benchmarking and Rethinking Long-Term Temporal Understanding in Remote Sensing
ChronoBench decomposes long-term remote sensing understanding into four cognitive levels, and the GeoChrono model, using per-location temporal trajectories, achieves 78.34% accuracy—over 20 points above prior MLLMs—bu...
-
OSMDA: OpenStreetMap-based Domain Adaptation for Remote Sensing VLMs
A remote-sensing VLM can be adapted by having it read rendered OpenStreetMap maps paired with satellite images, then fine-tuning it on satellite images alone.
-
OpenEarthAgent: A Unified Framework for Tool-Augmented Geospatial Agents
A 4B model fine-tuned on tool-augmented geospatial reasoning traces outperforms larger general-purpose models on executable GIS/spectral tool-use benchmarks and matches frontier models on trajectory fidelity.
-
VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs
VectorLLM, a multimodal LLM that regresses building contour vertices token by token, reports gains of 5.6 to 13.6 AP over prior polygon extraction methods on WHU, WHU-Mix, and CrowdAI.
-
GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields
GeoProg3D combines a georeferenced hierarchical 3D language field, geographic vision APIs, and LLM-generated programs to answer natural-language queries about city-scale 3D scenes, and includes a new 952-query benchma...
-
UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence with Spatial Reasoning and Understanding
A fine-tuned small multimodal LLM outperforms much larger general models on urban tasks in a new benchmark, with caveats about benchmark overlap with training data.
-
GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing
GeoPix extends multimodal language models for remote sensing to pixel-level referring segmentation through a mask predictor, a class-wise learnable memory, and a new 65,463-image instruction dataset.
-
EarthDial: Turning Multi-sensory Earth Observations to Interactive Dialogues
EarthDial is a 4B-parameter remote sensing chatbot trained on 11.11M instruction pairs to handle multi-resolution, multi-spectral, and multi-temporal satellite imagery, and it reports gains over prior VLMs on dozens o...
-
RSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mixture of Experts
RSUniVLM is a 1-billion-parameter remote sensing vision-language model that unifies image-, region-, and pixel-level tasks plus multi-image change analysis, achieving state-of-the-art visual grounding on VRSBench and ...
-
GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks
State-of-the-art vision-language models answer only 41.7% of GEOBench-VLM multiple-choice questions correctly, roughly double random chance.
-
More with Less: a Large Scale Remote Sensing VLM with a Simple Recipe
An unmodified general-purpose VLM trained with multi-task RL and a SAM3 tool reaches top results on most remote sensing zero-shot benchmarks, with gains the paper attributes to training-data diversity rather than arch...
-
Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling
A remote sensing LVLM that augments visual features with retrieved captions and routes them through level-specific experts improves performance on several RS vision-language benchmarks.
-
A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features
A lightweight projection layer trained to align frozen SpectralGPT multispectral features with LLaMA-3 text embeddings markedly improves EuroSAT classification and enables multispectral scene description.
-
A Simple Aerial Detection Baseline of Multimodal Language Models
Fine-tuned Florence-2 detects rotated aerial objects in text form and matches conventional detectors under the paper's confidence-free mAPnc and mF1 metrics.
-
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models
A single vision-language model is fine-tuned to handle three remote sensing input types and reports state-of-the-art results on VQA, change captioning, and video classification benchmarks.
-
LHRS-Bot-Nova: Improved Multimodal Large Language Model for Remote Sensing Vision-Language Interpretation
A remote sensing chatbot with recaptioned image-text data and a mixture-of-experts visual bridge reports gains over prior general and remote sensing models, but no code or data are released yet.
-
REO-VLM: Transforming VLM to Meet Regression Challenges in Earth Observation
A vision-language model for Earth observation that adds numeric biomass regression to generative question answering, with R^2 up to 0.36 and patch counting still failing.
-
Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives
A structured review of remote sensing vision-language models, organizing contrastive, instruction-tuned, and generative approaches alongside their datasets and benchmarks.
Discussion (0). Continue with ORCID to comment.