Pith. sign in

Paper Citation Record · LEDGER

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery

As of 21 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2602.19190.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.19190 v4

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T21:47:04.669204Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved52
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4a1d9cd5-0a6d-4593-9242-8ac4b7c4089f · outbound

This paper cites Qwen3-VL Technical Report.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Qwen3-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:00.159441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:00.159441Z digest=sha256:189d87b2542a0cb0f4f99bb09b705a976f69691537b4fb4afb8c495f56bc2c35

Observation c1827d55-dcf3-4c0b-943e-13b6e65721fa · outbound

This paper cites Qwen2.5-vl technical report, 2025.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Qwen2.5-vl technical report, 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:00.243122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:00.243122Z digest=sha256:624a2b3b35a25761f085e86f769079608fa996e4472f6bec2b244edf6a76a96e

Observation 3b652e5c-b10f-48b8-936a-634a3b9bbb24 · outbound

This paper cites Multi-spectral remote sensing image retrieval using geospatial foundation models.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Multi-spectral remote sensing image retrieval using geospatial foundation models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:00.336448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:00.336448Z digest=sha256:2fb75098deeb77adc32f1e3bdd337353b110fcb00f255d3684075de56a0b84cd

Observation b273c935-9a05-44a3-8c12-5d0dbca40fbc · outbound

This paper cites Brown, Michal R.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Brown, Michal R

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:00.410631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:00.410631Z digest=sha256:eafa4bbfd50a124db93b85cf695448620b95d625b3808ec1528327334ddc99c6

Observation 87f84470-04b8-4af2-bd58-cd47c4e25988 · outbound

This paper cites Changeclip: Remote sensing change detection with multi- modal vision-language representation learning.ISPRS Jour- nal of Photogrammetry and Remote Sensing, 208:53–69,.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Changeclip: Remote sensing change detection with multi- modal vision-language representation learning.ISPRS Jour- nal of Photogrammetry and Remote Sensing, 208:53–69,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:00.522548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:00.522548Z digest=sha256:aeab693a5337d6c33b5dcc8f844e9f582f07540e148047f11b60f6059da5a6ad

Observation 68c87035-719c-4166-99ad-cb66105ee3b2 · outbound

This paper cites Combining sam with limited data for change detection in remote sensing.IEEE Transactions on Geoscience and Remote Sensing, 2025.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Combining sam with limited data for change detection in remote sensing.IEEE Transactions on Geoscience and Remote Sensing, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:00.660623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:00.660623Z digest=sha256:04e779a5348f1d338dc1e272766fed54b8aa013fcfcca14761f93cb659eb8d82

Observation 8e6191f6-3429-40ae-b4ac-d04e8106f2f8 · outbound

This paper cites Deep- learning for radar: A survey.IEEE Access, 9:141800– 141818, 2021.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Deep- learning for radar: A survey.IEEE Access, 9:141800– 141818, 2021

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:00.765883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:00.765883Z digest=sha256:f237953b2ce8a319bd167d1f19621c9a35d590f71c6da6b6c4422eaaaee45c6d

Observation ff677ab4-c154-48ea-98f6-b7fb1b3fe55f · outbound

This paper cites Unpaired image-text match- ing via multimodal aligned conceptual knowledge.IEEE Transactions on Pattern Analysis and Machine Intelligence, 47(7):5160–5176, 2024.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Unpaired image-text match- ing via multimodal aligned conceptual knowledge.IEEE Transactions on Pattern Analysis and Machine Intelligence, 47(7):5160–5176, 2024

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:00.835884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:00.835884Z digest=sha256:fb3e486840f21c240517368cbbfb7001aeab0b75ec1ec1b4bcde0bd4e812eeb4

Observation 1333c661-a96d-48ed-8c37-e926d8c7622a · outbound

This paper cites A fast progressive ship detection method for very large full-scene sar images.IEEE Transactions on Geo- science and Remote Sensing, 62:1–15, 2024.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery A fast progressive ship detection method for very large full-scene sar images.IEEE Transactions on Geo- science and Remote Sensing, 62:1–15, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:00.913150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:00.913150Z digest=sha256:4a7bb692827847b44bb84ae5d89917f19b9d47536036b8672948145c726e2ebc

Observation f658c8b6-1846-4ea6-91b4-655424f225e8 · outbound

This paper cites Sarclip: a multimodal foundation framework for sar imagery via con- trastive language-image pre-training.ISPRS Journal of Pho- togrammetry and Remote Sensing, 231:17–34, 2026.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Sarclip: a multimodal foundation framework for sar imagery via con- trastive language-image pre-training.ISPRS Journal of Pho- togrammetry and Remote Sensing, 231:17–34, 2026

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:01.005862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:01.005862Z digest=sha256:29025d3a6a81d8072367db4c70214721f093bc4ae1903496cad238f95a7a349f

Observation 1097abd9-5ccd-4918-8392-7e27724164c9 · outbound

This paper cites Em-yolo: A fine-grained recognition model for aircraft targets in sar im- ages.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Em-yolo: A fine-grained recognition model for aircraft targets in sar im- ages

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:01.094296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:01.094296Z digest=sha256:0d5d4dc1d09bcf17dd395faccf6fdc8c40334affa56e78e9675fc9dd5ee5bbaf

Observation 6fbf4849-38ec-4ddc-91a9-68f123c00731 · outbound

This paper cites Sar ship detection based on an improved faster r-cnn using deformable convolution.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Sar ship detection based on an improved faster r-cnn using deformable convolution

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:01.191491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:01.191491Z digest=sha256:4cecd0575751fac58c10b69189fcb88bcb53624f7dd14b9535db390b5cba4cc1

Observation dd608eeb-097d-4140-91d0-b0131a9f4319 · outbound

This paper cites Geochat: Grounded large vision-language model for remote sensing.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Geochat: Grounded large vision-language model for remote sensing

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:01.320843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:01.320843Z digest=sha256:737bd66a17e768feab8d257114d45d09327f8163be27792e703be7451448cea9

Observation 9751cc24-2318-45f1-81d5-29457cb8fdae · outbound

This paper cites Llava-onevision: Easy visual task transfer, 2024.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Llava-onevision: Easy visual task transfer, 2024

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:01.407673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:01.407673Z digest=sha256:9cb291e7a48a7efef98865667b0ff5bff32b4a07d82d5f5d23f099dde87c9009

Observation e5d53f27-2f39-481a-86ea-8ac61c154189 · outbound

This paper cites BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:01.472873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:01.472873Z digest=sha256:4a4b8ee0755b16e9f73fa7a0d0dd687bc7c9152815ec7756a4e51730e977893f

Observation 7311518f-8bce-4cb2-bcc6-c30ccc2b0cb1 · outbound

This paper cites A new learn- ing paradigm for foundation model-based remote-sensing change detection.IEEE Transactions on Geoscience and Re- mote Sensing, 62:1–12, 2024.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery A new learn- ing paradigm for foundation model-based remote-sensing change detection.IEEE Transactions on Geoscience and Re- mote Sensing, 62:1–12, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:01.532051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:01.532051Z digest=sha256:4b512eab0a0737fbcba75d88bece66d43f5fa56a701a5db27662e2252854ce28

Observation e879ceb5-1487-4b10-8b70-c2a7e1ecd95a · outbound

This paper cites Co-training vision-language models for remote sensing multi-task learning.Remote Sensing, 18(2): 222, 2026.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Co-training vision-language models for remote sensing multi-task learning.Remote Sensing, 18(2): 222, 2026

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:01.587363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:01.587363Z digest=sha256:264ad6c58804adf503cd994c88ad89e7a4eb374561d7182db4706bf447b48d83

Observation 847853ec-32d5-4941-8a1a-5d055eace7ee · outbound

This paper cites an unresolved cited work.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:01.671693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:01.671693Z digest=sha256:4889bb7505f2899dd3a32e4f34b930497c3279e17d2253fb8ca0cc3e3e6a4d2c

Observation 0e0e6b94-402e-40df-ba6a-7a596b1b1ff3 · outbound

This paper cites Star: A first-ever dataset and a large-scale benchmark for scene graph generation in large-size satellite imagery.IEEE Trans.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Star: A first-ever dataset and a large-scale benchmark for scene graph generation in large-size satellite imagery.IEEE Trans

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:01.763737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:01.763737Z digest=sha256:1be861a28d238336cc0fa9cddae4502e4bf8e399b1e9620051062dbe880beafa

Observation c9fdb410-9f36-48a5-9c1a-b7a390f4f9c4 · outbound

This paper cites Re- moteclip: A vision language foundation model for remote sensing.IEEE Transactions on Geoscience and Remote Sensing, 62:1–16, 2024.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Re- moteclip: A vision language foundation model for remote sensing.IEEE Transactions on Geoscience and Remote Sensing, 62:1–16, 2024

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:01.872284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:01.872284Z digest=sha256:97f8e301b21eb0b0e2a6d5f160c3fd261eefea9bf90be8be328b5eca96a8646b

Observation 03b5fdf9-576d-4099-9e96-56769a1dc7b1 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Improved baselines with visual instruction tuning, 2023

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:01.967690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:01.967690Z digest=sha256:d8a73f4d2ea8cf78a969183683d5fea23e0b33d4deecf620da481e18e18634cb

Observation 5032543d-858d-44c0-973a-38b5391247a7 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:02.043624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:02.043624Z digest=sha256:0297744c32a373007b0a40550fdbb79eafbb527ea5dc6f5f5ae13872c680da2f

Observation ddc3ecea-cfcf-48ed-b64b-d11515d66ad4 · outbound

This paper cites Texture classification from random features.IEEE transactions on pattern analysis and machine intelligence, 34(3):574–586, 2012.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Texture classification from random features.IEEE transactions on pattern analysis and machine intelligence, 34(3):574–586, 2012

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:02.107226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:02.107226Z digest=sha256:9821e582cffe39c31b2384d72a4ab558004241043365d1df87be37c51173f4df

Observation e0cd5172-886c-4a57-81e2-e6c552f1d6c2 · outbound

This paper cites A causal adjustment module for debiasing scene graph generation.IEEE Trans- actions on Pattern Analysis and Machine Intelligence, 47(5): 4024–4043, 2025.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery A causal adjustment module for debiasing scene graph generation.IEEE Trans- actions on Pattern Analysis and Machine Intelligence, 47(5): 4024–4043, 2025

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:02.222057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:02.222057Z digest=sha256:f3d3feaad1fec153ab7f75877f548ee837042c27330e9540342609bed5ca0b20

Observation c1ff220c-0a5e-48d5-a3a8-aa6cdda9d966 · outbound

This paper cites Atrnet-star: A large dataset and benchmark to- wards remote sensing object recognition in the wild.IEEE Transactions on Pattern Analysis and Machine Intelligence,.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Atrnet-star: A large dataset and benchmark to- wards remote sensing object recognition in the wild.IEEE Transactions on Pattern Analysis and Machine Intelligence,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:02.350777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:02.350777Z digest=sha256:521869bf9eb7351989fecd2cfb3d2dac4819f2a74daec9410c61a40593a61f69

Observation 94f33e6d-503a-4e36-8815-f7a98e71e93f · outbound

This paper cites Vhm: Versatile and honest vision language model for remote sensing image analysis.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Vhm: Versatile and honest vision language model for remote sensing image analysis

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:02.475420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:02.475420Z digest=sha256:07eef33ed713149cf00648442c84fb5e74cb2763feaadc4a3ea1eea2b8d2e622

Observation a8237ffe-93fd-4ffe-80e4-be3551e3f696 · outbound

This paper cites Learning transferable visual models from natural language supervision.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Learning transferable visual models from natural language supervision

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:02.578235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:02.578235Z digest=sha256:5c73bee9281bc6e9254a1b86feedcda39dfe4627cc6dd5d00f4bbbca7335dfc7

Observation 30ee7653-332c-433a-8396-7340f9a605da · outbound

This paper cites Grounding Everything in Tokens for Multimodal Large Language Models.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Grounding Everything in Tokens for Multimodal Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:02.670792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:02.670792Z digest=sha256:f6db209ae37c28e7488e727ceb310b684556eeaf49aea34a07bf570bbee72d30

Observation 2636d552-4451-4f70-91ad-669b5defe258 · outbound

This paper cites Earthdial: Turning multi-sensory earth observations to interactive dialogues.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Earthdial: Turning multi-sensory earth observations to interactive dialogues

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:02.720036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:02.720036Z digest=sha256:7da14f1bd807337c6ede5026f8795ad849dc7f4820eb116c294a8fff9c8613ff

Observation 96b05514-31a5-43ba-9864-c4080fbe7bfb · outbound

This paper cites Fully polsar image reconstruction for enhanced land cover mapping.Pattern Recognition, 169:111895, 2026.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Fully polsar image reconstruction for enhanced land cover mapping.Pattern Recognition, 169:111895, 2026

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:02.806763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:02.806763Z digest=sha256:0aefff3a1a543724e8b8cd538b7e4473b890b519fd79458774e0479d6bc04ab7

Observation 65ae49c1-3117-4d2d-8a93-1474b238facf · outbound

This paper cites Vi- sual position prompt for mllm based visual grounding.IEEE Transactions on Multimedia, 2026.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Vi- sual position prompt for mllm based visual grounding.IEEE Transactions on Multimedia, 2026

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:02.865615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:02.865615Z digest=sha256:5af9b5be6d736de94d53850028de169e65bd4557692d2b0f3a5c3c827cb56a89

Observation 82e796d5-c171-4dcc-957d-5ec0b4323b50 · outbound

This paper cites Re- ferring expressions as a lens into spatial language grounding in vision-language models.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Re- ferring expressions as a lens into spatial language grounding in vision-language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:02.926297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:02.926297Z digest=sha256:90a5fb4b71e181db5dcc855a275ccbd7e0262b24305cff827a51e18c70ad03f7

Observation e5f1c877-975e-4832-8af8-9d7fa80254c8 · outbound

This paper cites Group equivariant u-net for the semantic segmentation of sar im- ages.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Group equivariant u-net for the semantic segmentation of sar im- ages

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:03.020406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:03.020406Z digest=sha256:46c2a88a5375b5411c54738810c66b0613dfd0ed331788be9b1c702b23aec016

Observation 5f93fb2f-fb8d-416b-b3ab-ebae82e92f7b · outbound

This paper cites Annotation-free, high-fidelity sar oil-spill image synthesis via classification-guided diffusion model.IEEE Transactions on Geoscience and Remote Sens- ing, 63:1–11, 2025.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Annotation-free, high-fidelity sar oil-spill image synthesis via classification-guided diffusion model.IEEE Transactions on Geoscience and Remote Sens- ing, 63:1–11, 2025

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:03.077134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:03.077134Z digest=sha256:5261aad22c298dcd46f7b7d0dcc6a326a8526da02ca0c61033ed970044cfe556

Observation c106161f-313b-42c0-af97-81130340cb03 · outbound

This paper cites Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution, 2024.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution, 2024

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:03.159165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:03.159165Z digest=sha256:94cd4bdf9960ef3d9e2e25282d223d2c7146e20faa85e244f6df0af5b8b6d831

Observation 7219751a-a3fc-4116-9c8a-9f13a82b205e · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:03.238957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:03.238957Z digest=sha256:b6a4d85adef312e6a5f50e55dedb6aeed66cc4083d511d1f5901f005fb49b564

Observation 40e3341b-a7dd-43d5-bcae-e3e217765fad · outbound

This paper cites Skyscript: A large and seman- tically diverse vision-language dataset for remote sensing.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Skyscript: A large and seman- tically diverse vision-language dataset for remote sensing

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:03.364687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:03.364687Z digest=sha256:9c35d8b2ce72f594e32aa5c0e5d1c8f652f838c0823b97f276e51e1ed86d71c6

Observation f4c42a91-38f8-494f-850e-962fcfcf5088 · outbound

This paper cites Sarlang-1m: A benchmark for vision-language modeling in sar image un- derstanding.IEEE Transactions on Geoscience and Remote Sensing, 2026.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Sarlang-1m: A benchmark for vision-language modeling in sar image un- derstanding.IEEE Transactions on Geoscience and Remote Sensing, 2026

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:03.430443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:03.430443Z digest=sha256:e9d1f83e5ed6d2dbb89d692278c57def2414dc3dd80b0f2a14d3d4878c8671e3

Observation 8e2d6445-020a-4461-a9a0-9fe298599595 · outbound

This paper cites Bootstrapping interactive image–text alignment for remote sensing image captioning.IEEE Transactions on Geoscience and Remote Sensing, 62:1–12, 2024.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Bootstrapping interactive image–text alignment for remote sensing image captioning.IEEE Transactions on Geoscience and Remote Sensing, 62:1–12, 2024

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:03.497183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:03.497183Z digest=sha256:6c70def0d9b302cf99541f5cfc2eddd39201c90686165af6d79b6214502fd6dd

Observation 7278e134-5a14-46b5-9afc-5348ab33b4c4 · outbound

This paper cites R3det: Refined single-stage detector with feature refinement for ro- tating object.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery R3det: Refined single-stage detector with feature refinement for ro- tating object

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:03.564279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:03.564279Z digest=sha256:1964ebb8f0b53bc1de8de63a496f18378cc64f8b4e33ca3d4774efab2c5fec1a

Observation 65a717dc-5683-4db1-9176-8ce0a9c40e60 · outbound

This paper cites an unresolved cited work.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:03.659956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:03.659956Z digest=sha256:c20d014dce28ecc24f910cf226e7a96ede0ffca8a59f43b0ace0159125ff565f

Observation 3f3da96d-90ac-4bee-86a2-de6914c4fedf · outbound

This paper cites Fusar-klip: Towards multimodal foundation models for remote sensing, 2025.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Fusar-klip: Towards multimodal foundation models for remote sensing, 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:03.766569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:03.766569Z digest=sha256:daa89056b3d00c98ac70d0e482dbbe0dd8ac63a66689ef1139f2f7c53ba42394

Observation 8ae3e397-d85e-4bf1-99c4-058d6ad952f2 · outbound

This paper cites Object fidelity diffu- sion for remote sensing image generation.arXiv preprint arXiv:2508.10801, 2025.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Object fidelity diffu- sion for remote sensing image generation.arXiv preprint arXiv:2508.10801, 2025

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:03.820210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:03.820210Z digest=sha256:60d5b0388b6816176388b22d6266daf6f479b209555d794645a3bdbaa0dca095

Observation ae798bc3-940e-4d09-a463-f9e2b0589c94 · outbound

This paper cites an unresolved cited work.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:03.888503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:03.888503Z digest=sha256:f2a487ac581e72aa8f6d9d75e8c8d5e900921a3539fc7686863c9069950947cd

Observation 51e2950d-11f2-497b-9fd0-a9c76a72eac4 · outbound

This paper cites Earthgpt-x: A spatial mllm for multi-level multi-source re- mote sensing imagery understanding with visual prompting,.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Earthgpt-x: A spatial mllm for multi-level multi-source re- mote sensing imagery understanding with visual prompting,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:03.977426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:03.977426Z digest=sha256:1d43fd677e06aa2e5a4392381d3c5ce1cdf6791ea7950ac315bdf030278b1f30

Observation a035c27e-c4e5-4ddd-9e5c-80ce78f2d639 · outbound

This paper cites A fast training method for sar large scale samples based on cnn for targets recognition.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery A fast training method for sar large scale samples based on cnn for targets recognition

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:04.061253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:04.061253Z digest=sha256:44cd3654cb6cd0f22e2fa552379b1ba81d617f822f68b70fa40bbc0e874f1c3c

Observation affe2e7f-db1b-4d46-b862-806b358bf36c · outbound

This paper cites Rs5m and georsclip: A large-scale vision-language dataset and a large vision-language model for remote sensing.IEEE Transactions on Geoscience and Remote Sensing, 62:1–23,.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Rs5m and georsclip: A large-scale vision-language dataset and a large vision-language model for remote sensing.IEEE Transactions on Geoscience and Remote Sensing, 62:1–23,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:04.164672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:04.164672Z digest=sha256:a7b46e6eebb34234bbd1f0140b644d199e037d27320681242fd1180e3e39e85e

Observation 4b761e81-d0b8-4935-9497-958a3d42c5fd · outbound

This paper cites Geo-r1: Improving few-shot geospatial referring expression understanding with reinforcement fine- tuning, 2025.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Geo-r1: Improving few-shot geospatial referring expression understanding with reinforcement fine- tuning, 2025

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:04.259014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:04.259014Z digest=sha256:3215f190885dcf6eda43c2e1b743beae71de58a0208a51d7e504b38da1e53b2c

Observation c7ca8b52-d712-4670-a9e3-8dd4b11d2891 · outbound

This paper cites Towards vision- language geo-foundation model: A survey.arXiv preprint arXiv:2406.09385, 2024.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Towards vision- language geo-foundation model: A survey.arXiv preprint arXiv:2406.09385, 2024

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:04.363440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:04.363440Z digest=sha256:c5d681ef85bd8b68b925285337717d6cc02c8efdc34b61ba4389ecd3ab15919c

Observation ca862e42-4fd8-44d8-98e6-2d9ab1192f37 · outbound

This paper cites 6, SAR imagery exhibits inherent lim- itations that constrain visual–semantic understanding.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery 6, SAR imagery exhibits inherent lim- itations that constrain visual–semantic understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:04.441518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:04.441518Z digest=sha256:973cda4807dae0734df7b0da3b2c1ab84057fee54eb3476ba8135e4cd6352bff

Observation 1110d89e-08a1-4980-a8e7-9da1ed3e5c91 · outbound

This paper cites FUSAR-GPT consistently outperforms all competing methods by a signif- icant margin across counting, grid-based localization, and classification tasks.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery FUSAR-GPT consistently outperforms all competing methods by a signif- icant margin across counting, grid-based localization, and classification tasks

Reference 51

Resolution
malformed identifier
no resolver link, observed 2026-08-02T21:47:04.531930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:04.531930Z digest=sha256:274f3d6aef832734dd80bc15daaec503afc98e70602b08e5ac82d7df44cbf2c8

Observation b4d696b9-b1ea-42fe-a303-415949e6328e · outbound

This paper cites Ablation results on the target counting task.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Ablation results on the target counting task

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:04.613602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:04.613602Z digest=sha256:62e42013bdf49b62c0f94463b19ba3b1a9ef2458822a42e3b3fceb9038a09b85

Observation bbf26c0e-e3ff-4fd3-a6e4-530933cdbcb8 · outbound

This paper cites an unresolved cited work.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:04.669204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:04.669204Z digest=sha256:3c52ab2d029a1bdf5cb3a53207c783392c2bd184aa76b5ad8c65e01ef363a63a

Pith citing papers

No inbound Pith citation observations are available.