Pith. sign in

Paper Citation Record · LEDGER

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery

As of 7 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2602.19190.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.19190 v4

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T21:47:04.669204Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved52
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4a1d9cd5-0a6d-4593-9242-8ac4b7c4089f · outbound

This paper cites Qwen3-VL Technical Report.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Qwen3-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:00.159441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:00.159441Z digest=sha256:13d3c091330b77a0b4e8c9008cf3024b078b640584c63d7dc568cd42bf1e421a

Observation c1827d55-dcf3-4c0b-943e-13b6e65721fa · outbound

This paper cites Qwen2.5-vl technical report, 2025.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Qwen2.5-vl technical report, 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:00.243122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:00.243122Z digest=sha256:ee88d0b9b1e6cba7ec1dbd90cbf3cd70d9967104eae90385e2af0ca769dc34d3

Observation 3b652e5c-b10f-48b8-936a-634a3b9bbb24 · outbound

This paper cites Multi-spectral remote sensing image retrieval using geospatial foundation models.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Multi-spectral remote sensing image retrieval using geospatial foundation models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:00.336448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:00.336448Z digest=sha256:9cc73159ed8cad29b8b6898c46a117ccee95d1d545d7f7eec40bb6a6f33c5651

Observation b273c935-9a05-44a3-8c12-5d0dbca40fbc · outbound

This paper cites Brown, Michal R.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Brown, Michal R

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:00.410631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:00.410631Z digest=sha256:017e16500346bd0ed39117fa907b7e201bc4972ecc1c031f0228452be780c741

Observation 87f84470-04b8-4af2-bd58-cd47c4e25988 · outbound

This paper cites Changeclip: Remote sensing change detection with multi- modal vision-language representation learning.ISPRS Jour- nal of Photogrammetry and Remote Sensing, 208:53–69,.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Changeclip: Remote sensing change detection with multi- modal vision-language representation learning.ISPRS Jour- nal of Photogrammetry and Remote Sensing, 208:53–69,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:00.522548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:00.522548Z digest=sha256:5baf75e3539329b4c01ed86c90272c7a4f135166273111aab3904919fe1dd869

Observation 68c87035-719c-4166-99ad-cb66105ee3b2 · outbound

This paper cites Combining sam with limited data for change detection in remote sensing.IEEE Transactions on Geoscience and Remote Sensing, 2025.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Combining sam with limited data for change detection in remote sensing.IEEE Transactions on Geoscience and Remote Sensing, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:00.660623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:00.660623Z digest=sha256:b36ae0344459278fb80d581289c7decd6c0b8912e296fbfef782b01dffeef666

Observation 8e6191f6-3429-40ae-b4ac-d04e8106f2f8 · outbound

This paper cites Deep- learning for radar: A survey.IEEE Access, 9:141800– 141818, 2021.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Deep- learning for radar: A survey.IEEE Access, 9:141800– 141818, 2021

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:00.765883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:00.765883Z digest=sha256:18e0ba4ee00e7e5d6f9ea89ba060c6a292abdce0a1e9f1a2e7a9c567898f0c84

Observation ff677ab4-c154-48ea-98f6-b7fb1b3fe55f · outbound

This paper cites Unpaired image-text match- ing via multimodal aligned conceptual knowledge.IEEE Transactions on Pattern Analysis and Machine Intelligence, 47(7):5160–5176, 2024.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Unpaired image-text match- ing via multimodal aligned conceptual knowledge.IEEE Transactions on Pattern Analysis and Machine Intelligence, 47(7):5160–5176, 2024

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:00.835884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:00.835884Z digest=sha256:54b16d800bc81165d8e531d22343773521aa3d61d3612a53e5b164aba2b82a2d

Observation 1333c661-a96d-48ed-8c37-e926d8c7622a · outbound

This paper cites A fast progressive ship detection method for very large full-scene sar images.IEEE Transactions on Geo- science and Remote Sensing, 62:1–15, 2024.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery A fast progressive ship detection method for very large full-scene sar images.IEEE Transactions on Geo- science and Remote Sensing, 62:1–15, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:00.913150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:00.913150Z digest=sha256:cdd80cae0b2a0a7b6328e3771e8dab754e2a896cfc433b872281c06e87a67f78

Observation f658c8b6-1846-4ea6-91b4-655424f225e8 · outbound

This paper cites Sarclip: a multimodal foundation framework for sar imagery via con- trastive language-image pre-training.ISPRS Journal of Pho- togrammetry and Remote Sensing, 231:17–34, 2026.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Sarclip: a multimodal foundation framework for sar imagery via con- trastive language-image pre-training.ISPRS Journal of Pho- togrammetry and Remote Sensing, 231:17–34, 2026

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:01.005862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:01.005862Z digest=sha256:4b13a9dfe14ead05ae4d53413c3f1a050f1562a37706e2de4e04c57faf23fac2

Observation 1097abd9-5ccd-4918-8392-7e27724164c9 · outbound

This paper cites Em-yolo: A fine-grained recognition model for aircraft targets in sar im- ages.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Em-yolo: A fine-grained recognition model for aircraft targets in sar im- ages

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:01.094296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:01.094296Z digest=sha256:47702264386f1e018b97b04a79a5d1533e4fc6fb04bbaf5e409de782a55b1c7d

Observation 6fbf4849-38ec-4ddc-91a9-68f123c00731 · outbound

This paper cites Sar ship detection based on an improved faster r-cnn using deformable convolution.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Sar ship detection based on an improved faster r-cnn using deformable convolution

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:01.191491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:01.191491Z digest=sha256:76ff712d183e4b7cd879d0daf31b462b5e9e9bde2494a12ee41aed33d6cee48f

Observation dd608eeb-097d-4140-91d0-b0131a9f4319 · outbound

This paper cites Geochat: Grounded large vision-language model for remote sensing.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Geochat: Grounded large vision-language model for remote sensing

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:01.320843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:01.320843Z digest=sha256:077dff7c9c3f37556af21db5c5410533d21220b4acff96677ad52a9ff01e2365

Observation 9751cc24-2318-45f1-81d5-29457cb8fdae · outbound

This paper cites Llava-onevision: Easy visual task transfer, 2024.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Llava-onevision: Easy visual task transfer, 2024

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:01.407673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:01.407673Z digest=sha256:a08582b07f29a57689081b1855d232e77e5376fd1f7ba2dc4d70d3d46210083d

Observation e5d53f27-2f39-481a-86ea-8ac61c154189 · outbound

This paper cites BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:01.472873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:01.472873Z digest=sha256:c8b94eb142aeb42bd509b29bc7870a1e7eb77852bee2293dd655e257ca9d3f72

Observation 7311518f-8bce-4cb2-bcc6-c30ccc2b0cb1 · outbound

This paper cites A new learn- ing paradigm for foundation model-based remote-sensing change detection.IEEE Transactions on Geoscience and Re- mote Sensing, 62:1–12, 2024.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery A new learn- ing paradigm for foundation model-based remote-sensing change detection.IEEE Transactions on Geoscience and Re- mote Sensing, 62:1–12, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:01.532051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:01.532051Z digest=sha256:8af21b29aa56f9ad9958e95e678277a7f982b355bbf755207dfe908ec3cbfa64

Observation e879ceb5-1487-4b10-8b70-c2a7e1ecd95a · outbound

This paper cites Co-training vision-language models for remote sensing multi-task learning.Remote Sensing, 18(2): 222, 2026.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Co-training vision-language models for remote sensing multi-task learning.Remote Sensing, 18(2): 222, 2026

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:01.587363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:01.587363Z digest=sha256:05480e8553cf3fffe5e1bcc93ee52f9fb21adb519cff1c44f576f63241003990

Observation 847853ec-32d5-4941-8a1a-5d055eace7ee · outbound

This paper cites an unresolved cited work.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:01.671693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:01.671693Z digest=sha256:d92ad43fb6c6bc85f4e3b02ca625a54eaa0e4462f4a117571a293b2963f5bbb9

Observation 0e0e6b94-402e-40df-ba6a-7a596b1b1ff3 · outbound

This paper cites Star: A first-ever dataset and a large-scale benchmark for scene graph generation in large-size satellite imagery.IEEE Trans.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Star: A first-ever dataset and a large-scale benchmark for scene graph generation in large-size satellite imagery.IEEE Trans

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:01.763737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:01.763737Z digest=sha256:5d0a165fb22d13be45b627b2c15721fbb291c7a22ee8f5cf066560660b8968de

Observation c9fdb410-9f36-48a5-9c1a-b7a390f4f9c4 · outbound

This paper cites Re- moteclip: A vision language foundation model for remote sensing.IEEE Transactions on Geoscience and Remote Sensing, 62:1–16, 2024.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Re- moteclip: A vision language foundation model for remote sensing.IEEE Transactions on Geoscience and Remote Sensing, 62:1–16, 2024

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:01.872284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:01.872284Z digest=sha256:e13a478018cc0186376c76a5c3a220805b305750506febb8fa33558fef1aaf8b

Observation 03b5fdf9-576d-4099-9e96-56769a1dc7b1 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Improved baselines with visual instruction tuning, 2023

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:01.967690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:01.967690Z digest=sha256:d539813dbb8cec32e15a1ea06f7770d076d2ed78fd8945a455608670492d0660

Observation 5032543d-858d-44c0-973a-38b5391247a7 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:02.043624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:02.043624Z digest=sha256:073c19542bfab7cea8f369faa17d70e0d08d97724687ef77a128d4e99cfb8e63

Observation ddc3ecea-cfcf-48ed-b64b-d11515d66ad4 · outbound

This paper cites Texture classification from random features.IEEE transactions on pattern analysis and machine intelligence, 34(3):574–586, 2012.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Texture classification from random features.IEEE transactions on pattern analysis and machine intelligence, 34(3):574–586, 2012

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:02.107226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:02.107226Z digest=sha256:4ade35e3df229e2694cf74325bc19d19a6f4a7bc22e7b88fc7d922d8d828f259

Observation e0cd5172-886c-4a57-81e2-e6c552f1d6c2 · outbound

This paper cites A causal adjustment module for debiasing scene graph generation.IEEE Trans- actions on Pattern Analysis and Machine Intelligence, 47(5): 4024–4043, 2025.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery A causal adjustment module for debiasing scene graph generation.IEEE Trans- actions on Pattern Analysis and Machine Intelligence, 47(5): 4024–4043, 2025

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:02.222057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:02.222057Z digest=sha256:b5f56c157385c86e7ce4dbaba73f34795944244d787554f7478fc5822b7834c0

Observation c1ff220c-0a5e-48d5-a3a8-aa6cdda9d966 · outbound

This paper cites Atrnet-star: A large dataset and benchmark to- wards remote sensing object recognition in the wild.IEEE Transactions on Pattern Analysis and Machine Intelligence,.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Atrnet-star: A large dataset and benchmark to- wards remote sensing object recognition in the wild.IEEE Transactions on Pattern Analysis and Machine Intelligence,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:02.350777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:02.350777Z digest=sha256:352edbc5b8286d0749bafeedd661259154f4b72eb40b9f221f26b8f105bcb380

Observation 94f33e6d-503a-4e36-8815-f7a98e71e93f · outbound

This paper cites Vhm: Versatile and honest vision language model for remote sensing image analysis.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Vhm: Versatile and honest vision language model for remote sensing image analysis

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:02.475420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:02.475420Z digest=sha256:6cb89dc93eeeaccf1e2b88265677eaf5def050137b3d8aecfcbb07b6a2104dac

Observation a8237ffe-93fd-4ffe-80e4-be3551e3f696 · outbound

This paper cites Learning transferable visual models from natural language supervision.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Learning transferable visual models from natural language supervision

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:02.578235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:02.578235Z digest=sha256:da3093ef2071705ba207f5c3e61d8901fec43f24f31abef39f5914e5bb986bf7

Observation 30ee7653-332c-433a-8396-7340f9a605da · outbound

This paper cites Grounding Everything in Tokens for Multimodal Large Language Models.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Grounding Everything in Tokens for Multimodal Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:02.670792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:02.670792Z digest=sha256:bdeb30646dc417ce2f6d5ebebc3e0ea74fd4931babb4fb9401f8a1c7232c39e3

Observation 2636d552-4451-4f70-91ad-669b5defe258 · outbound

This paper cites Earthdial: Turning multi-sensory earth observations to interactive dialogues.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Earthdial: Turning multi-sensory earth observations to interactive dialogues

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:02.720036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:02.720036Z digest=sha256:59f69c4122b7174884ec1933be13c715639c40baf2eb739cfbca291e5d0abad5

Observation 96b05514-31a5-43ba-9864-c4080fbe7bfb · outbound

This paper cites Fully polsar image reconstruction for enhanced land cover mapping.Pattern Recognition, 169:111895, 2026.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Fully polsar image reconstruction for enhanced land cover mapping.Pattern Recognition, 169:111895, 2026

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:02.806763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:02.806763Z digest=sha256:c2e7dfbf05dbbab435a4ee80f31e311608c0b914cea2d395c9e01174b03d668e

Observation 65ae49c1-3117-4d2d-8a93-1474b238facf · outbound

This paper cites Vi- sual position prompt for mllm based visual grounding.IEEE Transactions on Multimedia, 2026.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Vi- sual position prompt for mllm based visual grounding.IEEE Transactions on Multimedia, 2026

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:02.865615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:02.865615Z digest=sha256:39eff386b4ed2b91072be82dd46b872b2e503281a83c5088229c22785eaef952

Observation 82e796d5-c171-4dcc-957d-5ec0b4323b50 · outbound

This paper cites Re- ferring expressions as a lens into spatial language grounding in vision-language models.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Re- ferring expressions as a lens into spatial language grounding in vision-language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:02.926297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:02.926297Z digest=sha256:7a15c0f3b2a75293c75303f434d33a85cc46b4e8b58169fb0f5ea601069b8b90

Observation e5f1c877-975e-4832-8af8-9d7fa80254c8 · outbound

This paper cites Group equivariant u-net for the semantic segmentation of sar im- ages.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Group equivariant u-net for the semantic segmentation of sar im- ages

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:03.020406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:03.020406Z digest=sha256:7f11ddd5ee1059217c3013cbf2b2cabc415f50e57373065d5d040b6b2bfcfeac

Observation 5f93fb2f-fb8d-416b-b3ab-ebae82e92f7b · outbound

This paper cites Annotation-free, high-fidelity sar oil-spill image synthesis via classification-guided diffusion model.IEEE Transactions on Geoscience and Remote Sens- ing, 63:1–11, 2025.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Annotation-free, high-fidelity sar oil-spill image synthesis via classification-guided diffusion model.IEEE Transactions on Geoscience and Remote Sens- ing, 63:1–11, 2025

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:03.077134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:03.077134Z digest=sha256:c69720c729e1a518b9dc71a42bcb19c762f2d545614d4c70f43a0fbf6ede7ae3

Observation c106161f-313b-42c0-af97-81130340cb03 · outbound

This paper cites Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution, 2024.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution, 2024

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:03.159165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:03.159165Z digest=sha256:4fcc42d8d956f20ecac8f4ea17bc450aa5e68bbc7bccfdf02684190e8f439b41

Observation 7219751a-a3fc-4116-9c8a-9f13a82b205e · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:03.238957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:03.238957Z digest=sha256:ecdc6ebcbbf538dd2bb335ba571fa4301a1198c1efd36884f257b3fb676e9423

Observation 40e3341b-a7dd-43d5-bcae-e3e217765fad · outbound

This paper cites Skyscript: A large and seman- tically diverse vision-language dataset for remote sensing.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Skyscript: A large and seman- tically diverse vision-language dataset for remote sensing

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:03.364687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:03.364687Z digest=sha256:f2d1dc9fb3ecb81ee0c8b3140b92e152fe472ef7e0c7e1c6e506f705767c196a

Observation f4c42a91-38f8-494f-850e-962fcfcf5088 · outbound

This paper cites Sarlang-1m: A benchmark for vision-language modeling in sar image un- derstanding.IEEE Transactions on Geoscience and Remote Sensing, 2026.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Sarlang-1m: A benchmark for vision-language modeling in sar image un- derstanding.IEEE Transactions on Geoscience and Remote Sensing, 2026

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:03.430443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:03.430443Z digest=sha256:1611d2c2eb1fe7ee7167efeefb94e9f3ec9186d86e26c001f5b607891ca01f13

Observation 8e2d6445-020a-4461-a9a0-9fe298599595 · outbound

This paper cites Bootstrapping interactive image–text alignment for remote sensing image captioning.IEEE Transactions on Geoscience and Remote Sensing, 62:1–12, 2024.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Bootstrapping interactive image–text alignment for remote sensing image captioning.IEEE Transactions on Geoscience and Remote Sensing, 62:1–12, 2024

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:03.497183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:03.497183Z digest=sha256:6bce5695011e8d095fe1d4d91f3f9da511c056deab58c14165fc5d54f40fc4a6

Observation 7278e134-5a14-46b5-9afc-5348ab33b4c4 · outbound

This paper cites R3det: Refined single-stage detector with feature refinement for ro- tating object.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery R3det: Refined single-stage detector with feature refinement for ro- tating object

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:03.564279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:03.564279Z digest=sha256:f646e2682ef60dc765861986befb0c629c34fd6d21d8289e87a736fe2f9b54e9

Observation 65a717dc-5683-4db1-9176-8ce0a9c40e60 · outbound

This paper cites an unresolved cited work.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:03.659956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:03.659956Z digest=sha256:c58904befd0fc1aa618cbf3289838d77d683910256345ccf06d868e32c3344c1

Observation 3f3da96d-90ac-4bee-86a2-de6914c4fedf · outbound

This paper cites Fusar-klip: Towards multimodal foundation models for remote sensing, 2025.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Fusar-klip: Towards multimodal foundation models for remote sensing, 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:03.766569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:03.766569Z digest=sha256:5dd926e74b46794af1774c7651eb406d9357e8c1ca8b9db574dccf1036ae6db6

Observation 8ae3e397-d85e-4bf1-99c4-058d6ad952f2 · outbound

This paper cites Object fidelity diffu- sion for remote sensing image generation.arXiv preprint arXiv:2508.10801, 2025.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Object fidelity diffu- sion for remote sensing image generation.arXiv preprint arXiv:2508.10801, 2025

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:03.820210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:03.820210Z digest=sha256:16a0b082da786b63cbf26236377e7e02599ee9ad40e1ec7178899412150b9601

Observation ae798bc3-940e-4d09-a463-f9e2b0589c94 · outbound

This paper cites an unresolved cited work.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:03.888503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:03.888503Z digest=sha256:62247b2b8a64597784cb1d7c8611c3bc7bf4e9a10596bed271bfaab55083c1fb

Observation 51e2950d-11f2-497b-9fd0-a9c76a72eac4 · outbound

This paper cites Earthgpt-x: A spatial mllm for multi-level multi-source re- mote sensing imagery understanding with visual prompting,.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Earthgpt-x: A spatial mllm for multi-level multi-source re- mote sensing imagery understanding with visual prompting,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:03.977426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:03.977426Z digest=sha256:4a52f29c34eeaf72128f32397d84331808a8fe90b6bca63076544973ee87b85a

Observation a035c27e-c4e5-4ddd-9e5c-80ce78f2d639 · outbound

This paper cites A fast training method for sar large scale samples based on cnn for targets recognition.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery A fast training method for sar large scale samples based on cnn for targets recognition

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:04.061253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:04.061253Z digest=sha256:ebd578d921f3e1f38481871ba3acf5924b6ee81f33d26bda49521a736bdaeeb4

Observation affe2e7f-db1b-4d46-b862-806b358bf36c · outbound

This paper cites Rs5m and georsclip: A large-scale vision-language dataset and a large vision-language model for remote sensing.IEEE Transactions on Geoscience and Remote Sensing, 62:1–23,.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Rs5m and georsclip: A large-scale vision-language dataset and a large vision-language model for remote sensing.IEEE Transactions on Geoscience and Remote Sensing, 62:1–23,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:04.164672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:04.164672Z digest=sha256:7d6aad168d9da07629bb9f05c1561d25b4bea93375678d15b2ae538a7023dc9d

Observation 4b761e81-d0b8-4935-9497-958a3d42c5fd · outbound

This paper cites Geo-r1: Improving few-shot geospatial referring expression understanding with reinforcement fine- tuning, 2025.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Geo-r1: Improving few-shot geospatial referring expression understanding with reinforcement fine- tuning, 2025

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:04.259014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:04.259014Z digest=sha256:9f8f7d032374b2bf919f19009e98f1068dd6d93e8c64383d6085e0b16bc9e1c3

Observation c7ca8b52-d712-4670-a9e3-8dd4b11d2891 · outbound

This paper cites Towards vision- language geo-foundation model: A survey.arXiv preprint arXiv:2406.09385, 2024.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Towards vision- language geo-foundation model: A survey.arXiv preprint arXiv:2406.09385, 2024

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:04.363440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:04.363440Z digest=sha256:7e32b5099b6eaee864b1e5ec67cabd410a845a8eff81cce846a95588e623ff2b

Observation ca862e42-4fd8-44d8-98e6-2d9ab1192f37 · outbound

This paper cites 6, SAR imagery exhibits inherent lim- itations that constrain visual–semantic understanding.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery 6, SAR imagery exhibits inherent lim- itations that constrain visual–semantic understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:04.441518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:04.441518Z digest=sha256:0ab259e29129d91b1af3a38a5b658d08fb651d716d1ccfb9fe611749a6414e28

Observation 1110d89e-08a1-4980-a8e7-9da1ed3e5c91 · outbound

This paper cites FUSAR-GPT consistently outperforms all competing methods by a signif- icant margin across counting, grid-based localization, and classification tasks.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery FUSAR-GPT consistently outperforms all competing methods by a signif- icant margin across counting, grid-based localization, and classification tasks

Reference 51

Resolution
malformed identifier
no resolver link, observed 2026-08-02T21:47:04.531930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:04.531930Z digest=sha256:a9c841fa6da7ead1d3615fe178159a68b76f86e488d64abf2439033fe5c02758

Observation b4d696b9-b1ea-42fe-a303-415949e6328e · outbound

This paper cites Ablation results on the target counting task.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Ablation results on the target counting task

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:04.613602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:04.613602Z digest=sha256:dbf1c895a3ccb621055f411234cadc60b48c0a0e20f79c99c59a1265f268f0a8

Observation bbf26c0e-e3ff-4fd3-a6e4-530933cdbcb8 · outbound

This paper cites an unresolved cited work.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:04.669204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:04.669204Z digest=sha256:cf42b9d81a9d802fae587b75e3af75eea7d2aba2ecc302fd45c332232b2ede69

Pith citing papers

No inbound Pith citation observations are available.