Pith. sign in

Paper Citation Record · LEDGER

RSGPT: A Remote Sensing Vision Language Model and Benchmark

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2307.15266.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2307.15266 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:09:08.528718Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T17:05:51.453799Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a9774f71-66ff-4897-97db-3c0f802a7730 · inbound

LHRS-Bot-Nova: Improved Multimodal Large Language Model for Remote Sensing Vision-Language Interpretation cites this paper.

LHRS-Bot-Nova: Improved Multimodal Large Language Model for Remote Sensing Vision-Language Interpretation RSGPT: A Remote Sensing Vision Language Model and Benchmark

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T20:54:45.651757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:54:45.651757Z digest=sha256:d52186adff0d66b3484145a58b184cda2b3d758f3873a847994f35b5dee23fce

Observation 8f5e3d6f-4ffe-41f9-8388-d5f0bf429695 · inbound

Large Vision-Language Models for Remote Sensing Visual Question Answering cites this paper.

Large Vision-Language Models for Remote Sensing Visual Question Answering RSGPT: A Remote Sensing Vision Language Model and Benchmark

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T19:15:40.732785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:15:40.732785Z digest=sha256:f6ce6ccfccfe304cb9778a9ac97a76e7fee1eea2b3ef52591daa10be904bfd95

Observation ba962f70-11da-476b-bbda-9c765070f5a4 · inbound

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs cites this paper.

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs RSGPT: A Remote Sensing Vision Language Model and Benchmark

Reference 138

Resolution
unresolved
no resolver link, observed 2026-08-12T14:31:37.266735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:31:37.266735Z digest=sha256:920b028208fe80b0e4db76e32b0bde5690b6003c166c8bf4e49d56aac8f2cdbc

Observation 2e5381a7-c03e-4aa2-8af6-376c51d4163c · inbound

GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks cites this paper.

GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks RSGPT: A Remote Sensing Vision Language Model and Benchmark

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T10:21:22.491536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:21:22.491536Z digest=sha256:06296617884fcfc5a5af97ec6a50f0d855ca32c680cca225dc7a7cf2db39d03f

Observation 1632509a-76ec-42c0-b2ad-5f63ca53df58 · inbound

RSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mixture of Experts cites this paper.

RSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mixture of Experts RSGPT: A Remote Sensing Vision Language Model and Benchmark

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T20:32:31.516181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:32:31.516181Z digest=sha256:294350b55b2c3e57e7189afb8750f6d52a9c1f810b59045e6781f26fcadf36c0

Observation 0f3179b1-9005-4bdd-8ab2-f8ec1cf0bc58 · inbound

EarthDial: Turning Multi-sensory Earth Observations to Interactive Dialogues cites this paper.

EarthDial: Turning Multi-sensory Earth Observations to Interactive Dialogues RSGPT: A Remote Sensing Vision Language Model and Benchmark

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:35.263849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:36:35.263849Z digest=sha256:436df7e127b653668974af0a3bc97ef44608f49d9fa772a9c31ee086d0db5dad

Observation e1fa586c-e91f-4078-82c5-69d5a2d4deed · inbound

REO-VLM: Transforming VLM to Meet Regression Challenges in Earth Observation cites this paper.

REO-VLM: Transforming VLM to Meet Regression Challenges in Earth Observation RSGPT: A Remote Sensing Vision Language Model and Benchmark

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T10:30:49.600576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:30:49.600576Z digest=sha256:fa7984f8cee82c1b69f0c8b31930e8f42bef92ba5eb935504a70176345a6e260

Observation 0221ae97-370c-4d17-9822-1dfbd82cf01e · inbound

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models cites this paper.

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models RSGPT: A Remote Sensing Vision Language Model and Benchmark

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T23:18:21.294085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:18:21.294085Z digest=sha256:4be2886db671b66069e92974f8cc9bbba35418f43a992b42ed01fb46575e52a2

Observation 73413d2c-6b90-4240-b15c-a994ab71d390 · inbound

Visual Large Language Models for Generalized and Specialized Applications cites this paper.

Visual Large Language Models for Generalized and Specialized Applications RSGPT: A Remote Sensing Vision Language Model and Benchmark

Reference 124

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.363894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.363894Z digest=sha256:e0aa80c228d8a1187b92aa2519d8171c123c3468af05bb34c78f3916bd340c32

Observation b57f24db-3cc5-4524-b60a-e8a3273e966d · inbound

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing cites this paper.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing RSGPT: A Remote Sensing Vision Language Model and Benchmark

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:47.960477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:47.960477Z digest=sha256:e111b487fa8aa8f7119d4c625b41440a97c18a05ba3d595e20342a7ddfcc426c

Observation 76172d4a-8c94-4492-81bd-414cd2f905d0 · inbound

GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing cites this paper.

GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing RSGPT: A Remote Sensing Vision Language Model and Benchmark

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:12.314531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:12.314531Z digest=sha256:a009b795e207d64971d3958d2940ecf3b50fdb0b901041b3a3405bc68439d966

Observation 230bbcc1-05db-493b-b1c1-5f64b364b39c · inbound

Multi-Agent Geospatial Copilots for Remote Sensing Workflows cites this paper.

Multi-Agent Geospatial Copilots for Remote Sensing Workflows RSGPT: A Remote Sensing Vision Language Model and Benchmark

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T13:41:49.043349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:41:49.043349Z digest=sha256:114bd30e345460076e1e60544fdc54039c66f74479312b559080bee8a3400d5e

Observation 52a68107-ddfe-483d-9ca3-3aee66aa6a90 · inbound

SARChat-Bench-2M: A Multi-Task Vision-Language Benchmark for SAR Image Interpretation cites this paper.

SARChat-Bench-2M: A Multi-Task Vision-Language Benchmark for SAR Image Interpretation RSGPT: A Remote Sensing Vision Language Model and Benchmark

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T10:14:30.133950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:14:30.133950Z digest=sha256:1a7b818ad37de863851e9d89e283bf0c7095a4c5b636fe55a9d55c39f3e8c63a

Observation a590ed7e-41a6-4a79-b50d-2dbdb13faf58 · inbound

FrogDogNet: Fourier frequency Retained visual prompt Output Guidance for Domain Generalization of CLIP in Remote Sensing cites this paper.

FrogDogNet: Fourier frequency Retained visual prompt Output Guidance for Domain Generalization of CLIP in Remote Sensing RSGPT: A Remote Sensing Vision Language Model and Benchmark

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T11:09:08.528718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:09:08.528718Z digest=sha256:4f552feadc16ae0dbc3e72406f4c96fcabc4f2b96553f446f508441c01615198

Observation 4ea8d0c7-8626-4ed6-97ff-c88b562bda53 · inbound

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions cites this paper.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions RSGPT: A Remote Sensing Vision Language Model and Benchmark

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:54.615881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:09:54.615881Z digest=sha256:c4fe4f534d185dd1bae781d0233d11462962cedc7a1d37d767a445a279464de4

Observation 89386da8-4e18-4d7c-801e-a4d5714a0262 · inbound

UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence with Spatial Reasoning and Understanding cites this paper.

UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence with Spatial Reasoning and Understanding RSGPT: A Remote Sensing Vision Language Model and Benchmark

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:21.928081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:21.928081Z digest=sha256:d60195af849762597166285276ed12e7f6647145cfca7aa2ba2c93137a8c985c

Observation 3edb0cda-0e2e-45c5-b5c1-bc48b5709486 · inbound

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields cites this paper.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields RSGPT: A Remote Sensing Vision Language Model and Benchmark

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.574175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.574175Z digest=sha256:b3696867924dbe45d7c071bfb1cabaca917978efb1ffe8b23359a6f9cea5e77e

Observation d91cd1ef-41d9-4f02-9fbf-215499192c50 · inbound

A Satellite-Ground Synergistic Large Vision-Language Model System for Earth Observation cites this paper.

A Satellite-Ground Synergistic Large Vision-Language Model System for Earth Observation RSGPT: A Remote Sensing Vision Language Model and Benchmark

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T19:24:45.014869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:24:45.014869Z digest=sha256:0d9942b0fa17ff0a0951d20eea81c62549552bc0b8c317c5e3c86da60291e9fb

Observation 4e9ac6fd-9126-4258-8354-b240f88382ba · inbound

GeoMag: A Vision-Language Model for Pixel-level Fine-Grained Remote Sensing Image Parsing cites this paper.

GeoMag: A Vision-Language Model for Pixel-level Fine-Grained Remote Sensing Image Parsing RSGPT: A Remote Sensing Vision Language Model and Benchmark

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:13.189322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:13.189322Z digest=sha256:aeee9409b412dc5e675d5177201c1e8c7a0d541bf1557aa9688703100b3bcb4a

Observation 88f2835b-5272-4196-83ea-6c24caab3b65 · inbound

Enhancing Remote Sensing Vision-Language Models Through MLLM and LLM-Based High-Quality Image-Text Dataset Generation cites this paper.

Enhancing Remote Sensing Vision-Language Models Through MLLM and LLM-Based High-Quality Image-Text Dataset Generation RSGPT: A Remote Sensing Vision Language Model and Benchmark

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:07:32.895752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:07:32.895752Z digest=sha256:10ab5463fafefab9155fcf606ffc78c863539f97ad688aeb504ad3e232806245

Observation 4fc331a4-3056-4123-b880-b41c7de012e1 · inbound

Few-Shot Vision-Language Reasoning for Satellite Imagery via Verifiable Rewards cites this paper.

Few-Shot Vision-Language Reasoning for Satellite Imagery via Verifiable Rewards RSGPT: A Remote Sensing Vision Language Model and Benchmark

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T12:29:37.161294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:29:37.161294Z digest=sha256:36751612a6ab24288fed537b5604b29e8a75bc25745c3b8a4da82590c6d01004

Observation 84425ee0-fc77-48b2-ba75-c0e776e709ec · inbound

WildfireVLM: AI-powered Analysis for Early Wildfire Detection and Risk Assessment Using Satellite Imagery cites this paper.

WildfireVLM: AI-powered Analysis for Early Wildfire Detection and Risk Assessment Using Satellite Imagery RSGPT: A Remote Sensing Vision Language Model and Benchmark

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:17:22.437778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T05:14:25.778239Z digest=sha256:816a7bbd2cbfa452e61149cc251f09c7b790f0e3aa6ff653b794cf7cd41ec8f6

Observation fdae9e8b-69a3-4b7a-b2d3-c0fbc8a65ea0 · inbound

Geo2Sound: A Scalable Geo-Aligned Framework for Soundscape Generation from Satellite Imagery cites this paper.

Geo2Sound: A Scalable Geo-Aligned Framework for Soundscape Generation from Satellite Imagery RSGPT: A Remote Sensing Vision Language Model and Benchmark

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:59:03.522206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T09:52:35.741400Z digest=sha256:99113c56bce71943caac71920a62cbbeaf1582887be30a2680ce9d8a2878bc69

Observation 149d635b-41a2-4a53-be7b-bcbc09121d74 · inbound

ChangeQuery: Advancing Remote Sensing Change Analysis for Natural and Human-Induced Disasters from Visual Detection to Semantic Understanding cites this paper.

ChangeQuery: Advancing Remote Sensing Change Analysis for Natural and Human-Induced Disasters from Visual Detection to Semantic Understanding RSGPT: A Remote Sensing Vision Language Model and Benchmark

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:06:09.807361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T12:41:38.571997Z digest=sha256:ff5cfeb3fd13474fc5f8997b8b756ae739719614cf82e129b60209ee9561f040

Observation fcf196cd-2b09-4b70-bd31-e1c570e8d72e · inbound

Beyond GSD-as-Token: Continuous Scale Conditioning for Remote Sensing VLMs cites this paper.

Beyond GSD-as-Token: Continuous Scale Conditioning for Remote Sensing VLMs RSGPT: A Remote Sensing Vision Language Model and Benchmark

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:15:56.985551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T01:52:51.777228Z digest=sha256:b3b74be9a5505bae07010c91b1599ea2a996f39ca3270799b94520d4bcc516bd

Observation 8dcd6151-2a25-462d-a736-66e18c306fc0 · inbound

GeoVista: Visually Grounded Active Perception for Ultra-High-Resolution Remote Sensing Understanding cites this paper.

GeoVista: Visually Grounded Active Perception for Ultra-High-Resolution Remote Sensing Understanding RSGPT: A Remote Sensing Vision Language Model and Benchmark

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:53:33.974344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T02:39:56.666424Z digest=sha256:6551911728be3ecbfb4a7f3f084eed9ab35df0b737ff99c54eb2280c0760d4c3

Observation 4419f07d-7ece-4962-831b-b018f9877d56 · inbound

OmniCD: A Foundational Framework for Remote Sensing Image Change Detection Guided by Multimodal Semantics cites this paper.

OmniCD: A Foundational Framework for Remote Sensing Image Change Detection Guided by Multimodal Semantics RSGPT: A Remote Sensing Vision Language Model and Benchmark

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T08:13:16.011552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T08:03:27.451288Z digest=sha256:4fae287d66ae819a3ea1bab38e5d4650b824a3b49053d1538c5eb56899e75407

Observation 905bb316-54d7-40f5-a2d2-cd856d52b5c4 · inbound

RSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioning cites this paper.

RSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioning RSGPT: A Remote Sensing Vision Language Model and Benchmark

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T17:05:51.455236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T04:09:32.397341Z digest=sha256:48f58202cb1b6af4d7f9e32f2dc1c8e5c6a25287e88f95c8603e87604f4c941f

Observation 702be6a7-ebeb-4edb-9787-cc5e61a184f9 · inbound

WeaveEarth: Structured Evidence Construction and Reasoning for Training-Free UHR Remote Sensing Understanding cites this paper.

WeaveEarth: Structured Evidence Construction and Reasoning for Training-Free UHR Remote Sensing Understanding RSGPT: A Remote Sensing Vision Language Model and Benchmark

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T14:09:30.395518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:09:30.395518Z digest=sha256:8346cff47033f76c62ea75f6f5eeccbc8ba7af2bf8206c48ad7c9d714ed3ac89