Pith. sign in

Paper Citation Record · LEDGER

Visual Question Answering on Multiple Remote Sensing Image Modalities

As of 9 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2505.15401.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15401 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:21:57.827375Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact2
  • verified fuzzy38
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ee5f9f08-db8f-49f3-aa17-aae05f21407d · outbound

This paper cites Visual Question An- swering for Wishart H-Alpha Classification of Polarimetric SAR Images.

Visual Question Answering on Multiple Remote Sensing Image Modalities Visual Question An- swering for Wishart H-Alpha Classification of Polarimetric SAR Images

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:58.579321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.619777Z digest=sha256:204569a9736f897d93422b0b33ca9e3c67e46a42324077eb6e6f99d1878a4a8d

Observation 72cd52bd-54da-406e-a04e-f8f67f362add · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering.

Visual Question Answering on Multiple Remote Sensing Image Modalities Bottom-up and top-down attention for image captioning and visual question answering

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:58.565503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.626013Z digest=sha256:e0ed08398eb5f48c43d0e5ab3fd85c0673873c859e9d15cf778a9ed6d8e36321

Observation cbb3dcdb-0f39-4df0-af7e-50773d900597 · outbound

This paper cites VQA: Visual question answering.

Visual Question Answering on Multiple Remote Sensing Image Modalities VQA: Visual question answering

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:58.552413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.631059Z digest=sha256:423f9de0ef2c253416fdff760a6753e073d4d493bdfccb3e3230d7b1d6508d4d

Observation 14b93957-a845-4c1d-901a-f47bb9a90d79 · outbound

This paper cites Language trans- formers for remote sensing visual question answering.

Visual Question Answering on Multiple Remote Sensing Image Modalities Language trans- formers for remote sensing visual question answering

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:58.538376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.635584Z digest=sha256:6e54556a89624f3c98d1ca968b36163a7055d451ca20e50aceb3300b49d051e0

Observation 1569f4e1-b7de-4fde-a3df-03e7d78eda25 · outbound

This paper cites Prompt-RSVQA: Prompting visual context to a language model for remote sensing visual question answering.

Visual Question Answering on Multiple Remote Sensing Image Modalities Prompt-RSVQA: Prompting visual context to a language model for remote sensing visual question answering

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:58.525109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.641455Z digest=sha256:b45a1ab3c9269b0b27ecbc96fee97b13614e0bc26d42468b0a0a887f9e6ad7c9

Observation 3a3acf35-426a-4159-b21a-8a5961d67822 · outbound

This paper cites Multi-task prompt-RSVQA to explicitly count objects on aerial images.

Visual Question Answering on Multiple Remote Sensing Image Modalities Multi-task prompt-RSVQA to explicitly count objects on aerial images

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:58.507890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.646217Z digest=sha256:93f0695fa1c3855875112e7b75a9c96a17a6f967d3ff64fb48240ea693522cee

Observation 87075da6-15b3-43bc-a8d6-88ade9d3c2be · outbound

This paper cites The curse of language biases in remote sensing VQA: the role of spatial attributes, language diversity, and the need for clear evaluation.

Visual Question Answering on Multiple Remote Sensing Image Modalities The curse of language biases in remote sensing VQA: the role of spatial attributes, language diversity, and the need for clear evaluation

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:21:57.948776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.650897Z digest=sha256:60446f69b0e25e904b408a5c6e3b135c0a68d4c5f1602b3afcf0630b68669262

Observation 9c8c442e-917f-41f3-a6d5-0f50ec553a95 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Visual Question Answering on Multiple Remote Sensing Image Modalities BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:57.656270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:57.656270Z digest=sha256:39945712427b097614e7ce94a8128101de066c6cf9d1cd22ba84e7c4cbb03e13

Observation bad9fba5-c881-435c-bcba-c4c1ea3913c5 · outbound

This paper cites PubMedCLIP: How Much Does CLIP Benefit Visual Ques- tion Answering in the Medical Domain? InEACL, pages 1151–1163, 2023.

Visual Question Answering on Multiple Remote Sensing Image Modalities PubMedCLIP: How Much Does CLIP Benefit Visual Ques- tion Answering in the Medical Domain? InEACL, pages 1151–1163, 2023

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:58.492397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.661019Z digest=sha256:6ba58024a9c1fb210d7e861e3cc3590c8c44181b26e7094746e9093ff10e94c1

Observation 04287ee4-337b-413e-ab54-6e83bea5363a · outbound

This paper cites S2 missionhttps : / / sentiwiki.copernicus.eu/web/s2- mission# S2Mission - RadiometricPerformanceS2 - Mission - Radiometric - Performancetrue.

Visual Question Answering on Multiple Remote Sensing Image Modalities S2 missionhttps : / / sentiwiki.copernicus.eu/web/s2- mission# S2Mission - RadiometricPerformanceS2 - Mission - Radiometric - Performancetrue

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:58.476357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.665264Z digest=sha256:989f67b493f6dfbcb180de5b4b5916c1d13c9d0ec49172433acc47b38b003d02

Observation ce95719a-c126-4d54-9a28-a069b0e37db5 · outbound

This paper cites Cross- modal visual question answering for remote sensing data.

Visual Question Answering on Multiple Remote Sensing Image Modalities Cross- modal visual question answering for remote sensing data

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:58.462321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.669545Z digest=sha256:534394c7098c1a95dfc491d48b6a193536b13ac17c3633562e4b2189f1e9889a

Observation 6167abf0-df47-4bd7-b9bb-470696cd188d · outbound

This paper cites Making the V in VQA matter: El- evating the role of image understanding in visual question answering.

Visual Question Answering on Multiple Remote Sensing Image Modalities Making the V in VQA matter: El- evating the role of image understanding in visual question answering

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:58.447398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.673681Z digest=sha256:bbc67906738db17d94f1bbc1a10ad9f4572d364125efee8ed12392d4761d08f6

Observation b14f041d-5c20-43d0-a7f6-b1d2c604712d · outbound

This paper cites Overview of image- CLEF 2018 medical domain visual question answering task.

Visual Question Answering on Multiple Remote Sensing Image Modalities Overview of image- CLEF 2018 medical domain visual question answering task

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:58.433192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.677840Z digest=sha256:8a42ad89064cdcd27c8c2e50b28ab8d525280301d4c7044348c95937473b382b

Observation acab82ca-f225-44be-82a5-af778643655f · outbound

This paper cites PromptCap: Prompt-guided image captioning for SAR with GPT-3.

Visual Question Answering on Multiple Remote Sensing Image Modalities PromptCap: Prompt-guided image captioning for SAR with GPT-3

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:58.416960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.681918Z digest=sha256:e06cce3720c9408818fa2e54e149cd2fc02293a9a78f86369aa87bb6acc5292c

Observation 6b9b9f92-e9f7-404f-9574-b7ea41712f52 · outbound

This paper cites CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning.

Visual Question Answering on Multiple Remote Sensing Image Modalities CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:58.395061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.686919Z digest=sha256:4cb74aab151dfc9ba64fe879bf37e47e528634b04108094cf5c3d4432c837edc

Observation 9b8a1a1a-3e0b-4894-a8e1-7a382063846c · outbound

This paper cites Q: How to specialize large vision-language models to data-scarce VQA tasks? a: Self-train on unlabeled images! InCVPR Proceedings, pages 15005–15015, 2023.

Visual Question Answering on Multiple Remote Sensing Image Modalities Q: How to specialize large vision-language models to data-scarce VQA tasks? a: Self-train on unlabeled images! InCVPR Proceedings, pages 15005–15015, 2023

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:58.379262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.692176Z digest=sha256:c18599c8637e8e22bce8b0448bf394f73cbdf171ceb2be261f7faf6e6d857163

Observation 158f617a-7fbd-4cdd-ab8d-5c937bae040b · outbound

This paper cites Deep learning in multi- modal remote sensing data fusion: A comprehensive review.

Visual Question Answering on Multiple Remote Sensing Image Modalities Deep learning in multi- modal remote sensing data fusion: A comprehensive review

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:58.363362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.696242Z digest=sha256:cafd8585d2ce406f0e10e5edd82ea5b37728ae38e3a5bed6ab5e7147ba956b0d

Observation 86e211b8-017c-46b3-899a-a9ae42449976 · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

Visual Question Answering on Multiple Remote Sensing Image Modalities VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:57.700445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:57.700445Z digest=sha256:dcf23a1c277a3373a18f73444c2fdc8b233f1fca7276cdca3b98f82f7857fbb7

Observation 1301e4cd-21d6-46a0-b1e2-3cfff6b3f5d0 · outbound

This paper cites A comprehensive study of GPT-4V’s multimodal capabilities in medical imaging.

Visual Question Answering on Multiple Remote Sensing Image Modalities A comprehensive study of GPT-4V’s multimodal capabilities in medical imaging

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:58.348262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.706239Z digest=sha256:a5ff68a46addfb144f6290708b44a003393e2e150a698dd667aa5b9791a614d0

Observation 12c92640-5b84-4a23-b599-6aafacffe9af · outbound

This paper cites Medical visual question answering: A survey.Artificial In- telligence in Medicine, page 102611, 2023.

Visual Question Answering on Multiple Remote Sensing Image Modalities Medical visual question answering: A survey.Artificial In- telligence in Medicine, page 102611, 2023

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:58.333948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.711085Z digest=sha256:c72582692ebd5a0da30c8e92b2d9e3ff971ea44a30c1c6533f4aa8e3b40952de

Observation 01596c57-bf5b-485a-ba05-ac31a85d6fbd · outbound

This paper cites RSVQA: Visual question answering for remote sensing data.

Visual Question Answering on Multiple Remote Sensing Image Modalities RSVQA: Visual question answering for remote sensing data

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:58.320721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.715448Z digest=sha256:b464d147a4abd07c7bad61c058a6109e3752539d429da37849c3e19f26ac3f99

Observation ea858d96-e9e7-4634-85fd-bdbb2f754d66 · outbound

This paper cites RSVQA meets BigEarthNet: a new, large-scale, visual question an- swering dataset for remote sensing.

Visual Question Answering on Multiple Remote Sensing Image Modalities RSVQA meets BigEarthNet: a new, large-scale, visual question an- swering dataset for remote sensing

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:58.305404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.720420Z digest=sha256:fa6e70e0daa05a248d8f8154c702210b13b7220e1ece9faba4d473975d522720

Observation a031490f-64a6-4c73-95d2-fff894e5a46c · outbound

This paper cites Deep learning and earth observation to support the sustainable development goals: Current approaches, open challenges, and future opportunities.GRS, 10(2):172–200,.

Visual Question Answering on Multiple Remote Sensing Image Modalities Deep learning and earth observation to support the sustainable development goals: Current approaches, open challenges, and future opportunities.GRS, 10(2):172–200,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:58.288909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.725510Z digest=sha256:7e4d0693732ea8331b1dd3836f8d00ff722483bdee84d01e5fa0e7f4776786b6

Observation e1a65670-4db2-4f13-850d-ef3ddd8017a3 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Visual Question Answering on Multiple Remote Sensing Image Modalities Learn- ing transferable visual models from natural language super- vision

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:57.730126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:57.730126Z digest=sha256:f542e2a77f4889987b553bee2a42924f02018d37821e4bc530badc9065937cf6

Observation b2d6b8fa-8ecb-491b-ab41-ac35a409a5b3 · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

Visual Question Answering on Multiple Remote Sensing Image Modalities DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:57.734380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:57.734380Z digest=sha256:3667dc6da439c0aa2122fb39ede6427ba946d660805fd52d7466b49a4caf4d30

Observation 7f51676e-46bc-46ab-a39e-c9b364ed7332 · outbound

This paper cites How Much Can CLIP Benefit Vision-and-Language Tasks?.

Visual Question Answering on Multiple Remote Sensing Image Modalities How Much Can CLIP Benefit Vision-and-Language Tasks?

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:57.738692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:57.738692Z digest=sha256:fd4334f34f4c2fc38f0d20f76c46b5ca795a177445b55edeea971baa2e95c15a

Observation 33e86e99-5d9c-4be6-92c2-07820ee4a957 · outbound

This paper cites BigEarthNet: A Large-Scale Benchmark Archive For Remote Sensing Image Understanding.

Visual Question Answering on Multiple Remote Sensing Image Modalities BigEarthNet: A Large-Scale Benchmark Archive For Remote Sensing Image Understanding

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:58.265233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.742849Z digest=sha256:b7adabcb17c8b5efece2ecc98df0e97cadab0b471ac1371402b90c07fb5e1c82

Observation 5414bcfe-b7c1-4eb3-9991-cbed8248c692 · outbound

This paper cites BigEarthNet-MM: A large-scale, multimodal, multilabel benchmark archive for remote sensing image classification and retrieval [software and data sets].GRS, 9(3):174–180, 2021.

Visual Question Answering on Multiple Remote Sensing Image Modalities BigEarthNet-MM: A large-scale, multimodal, multilabel benchmark archive for remote sensing image classification and retrieval [software and data sets].GRS, 9(3):174–180, 2021

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:58.249885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.746878Z digest=sha256:a6a558157e7db9e8a9efecc1fd60ba9dcd7b0a305a3813cf36890d70a571a4a9

Observation ac91fa7a-49c9-428c-bace-78be154dbe7d · outbound

This paper cites LXMERT: Learning cross- modality encoder representations from transformers.

Visual Question Answering on Multiple Remote Sensing Image Modalities LXMERT: Learning cross- modality encoder representations from transformers

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:58.234529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.752529Z digest=sha256:86e527c1b7f3fa24c65d520e427d67743675196a960d0db582163a96f6757998

Observation 479d1f48-f656-44c2-9d25-704d89862741 · outbound

This paper cites Segmentation-guided attention for visual question answering from remote sensing images.

Visual Question Answering on Multiple Remote Sensing Image Modalities Segmentation-guided attention for visual question answering from remote sensing images

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:58.220015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.757403Z digest=sha256:35be44ab30daf6003970ff02cdd119ae1e733133e4afdbfc774d2aae1393ee83

Observation 33ff6f2d-ca6c-4bc6-8303-f71d08550ef9 · outbound

This paper cites Can SAR improve RSVQA performance? In EUSAR, pages 1287–1292.

Visual Question Answering on Multiple Remote Sensing Image Modalities Can SAR improve RSVQA performance? In EUSAR, pages 1287–1292

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:58.204385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.762055Z digest=sha256:cd96d41dbd21266a34ed494506743cea9c64725b14d3b9f19a92a49226de2b46

Observation 28086c65-172e-49b7-a07b-3d98ad54f09c · outbound

This paper cites A Visual Question Answering Method for SAR Ship: Breaking the Requirement for Multimodal Dataset Construction and Model Fine-Tuning.

Visual Question Answering on Multiple Remote Sensing Image Modalities A Visual Question Answering Method for SAR Ship: Breaking the Requirement for Multimodal Dataset Construction and Model Fine-Tuning

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:21:57.871156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.767100Z digest=sha256:b88927ab5447c878d2c0b544c3ffda9d836d43e49378b9edd68b4d4119d31d56

Observation 190efb43-6d6a-4629-a6e1-d16bb1bc55ff · outbound

This paper cites Labsar, a one- gcp coregistration tool for sar–insar local analysis in high- mountain regions.Frontiers in Remote Sensing, 3:935137,.

Visual Question Answering on Multiple Remote Sensing Image Modalities Labsar, a one- gcp coregistration tool for sar–insar local analysis in high- mountain regions.Frontiers in Remote Sensing, 3:935137,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:58.185744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.772554Z digest=sha256:7aef2480287eb6f19c2a5812c1248e472e4db2444fd2d15aa4019f1699df21fa

Observation eefacd51-dccc-4db1-a41f-169c64f1d508 · outbound

This paper cites LabSAR, a one-GCP coregistration tool for SAR–InSAR local analysis in high- mountain regions.Frontiers in Remote Sensing, 3, 2022.

Visual Question Answering on Multiple Remote Sensing Image Modalities LabSAR, a one-GCP coregistration tool for SAR–InSAR local analysis in high- mountain regions.Frontiers in Remote Sensing, 3, 2022

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:58.165300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.776892Z digest=sha256:4e480630e9c99eb9d08a6e130018f425bafce0b2e118aa3c132121dc079b706b

Observation 4370f8d1-0e16-41e1-bde3-84d33c886329 · outbound

This paper cites Stacked attention networks for image question answering.

Visual Question Answering on Multiple Remote Sensing Image Modalities Stacked attention networks for image question answering

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:58.147493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.781219Z digest=sha256:75ce269c8319dfa9a4bb64c425dc642fa1264dc518664260818bafdd6f8a30f3

Observation 481cc563-b99c-4592-8e44-1758d5739415 · outbound

This paper cites Self- paced curriculum learning for visual question answering on remote sensing data.

Visual Question Answering on Multiple Remote Sensing Image Modalities Self- paced curriculum learning for visual question answering on remote sensing data

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:58.131387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.785581Z digest=sha256:43f050d51970341e7c54ccc2619f3698d370510d728e2680023f4abc635329f6

Observation 4922c276-5ea5-4815-b7dc-0e8fc436f18a · outbound

This paper cites Multi- lingual augmentation for robust visual question answering in remote sensing images.

Visual Question Answering on Multiple Remote Sensing Image Modalities Multi- lingual augmentation for robust visual question answering in remote sensing images

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:58.113860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.790056Z digest=sha256:5c3e9d4d129909e3423f849ecd8e87e89df319cdbbde0618963a47ddc4e6a3f5

Observation 3639b5e5-987d-42e6-a45c-1158e4d87256 · outbound

This paper cites Frequency domain transfer learning for remote sensing visual question answering.

Visual Question Answering on Multiple Remote Sensing Image Modalities Frequency domain transfer learning for remote sensing visual question answering

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:58.090347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.794006Z digest=sha256:a06c3f66137b5a29e2def19caf023d6987468e7d0daeade283dbb6f9b7a0cf5f

Observation d2f071f0-c906-4382-8fdf-342f25dfc8a7 · outbound

This paper cites Exploring data and models in SAR ship image captioning.IEEE Access, 10:pp.

Visual Question Answering on Multiple Remote Sensing Image Modalities Exploring data and models in SAR ship image captioning.IEEE Access, 10:pp

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:58.074940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.798095Z digest=sha256:5e874d1d11feaac0cce99e0ffaafcdeff5cccd5c73e453bac685d17ee9b63337

Observation 404125af-856a-44de-885a-48d7e81ab912 · outbound

This paper cites Mutual attention inception network for remote sensing visual question answering.TGRS, 60:1–14, 2021.

Visual Question Answering on Multiple Remote Sensing Image Modalities Mutual attention inception network for remote sensing visual question answering.TGRS, 60:1–14, 2021

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:58.058588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.802499Z digest=sha256:557b7253a428dc9f20028099e7b63e711398ef8038a361d52f4c1db258b4bc4c

Observation 2702e120-efee-4a32-aa7d-5874ef75d4fe · outbound

This paper cites TRAR: Routing the attention spans in transformer for visual question answering.

Visual Question Answering on Multiple Remote Sensing Image Modalities TRAR: Routing the attention spans in transformer for visual question answering

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:58.040297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.806834Z digest=sha256:ece5f3a607f5bdec0765cd8f0ec3f2edbdbb09048679e037eeee3abcdbdd4b37

Observation 08bba89e-9e50-447d-9c45-c26b9122cbbc · outbound

This paper cites To do so, we need the geographical position of the center of the VHR patch.

Visual Question Answering on Multiple Remote Sensing Image Modalities To do so, we need the geographical position of the center of the VHR patch

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:58.025421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.811034Z digest=sha256:62b5bc5b2e3247ed7ed9ce7c378e077ba0e30c0cf9b281ce080bfddf97aaf19b

Observation 06147c63-591a-457c-8970-996923b93856 · outbound

This paper cites an unresolved cited work.

Visual Question Answering on Multiple Remote Sensing Image Modalities Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:21:58.010460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.815022Z digest=sha256:e7b64a1e0d3f9bdd8dd91c94864aabf79d8070edf4b43d3e875cacbdecb696c7

Observation aa5e0a4b-d6f6-4b0c-ab9b-3e3eb4239863 · outbound

This paper cites To find the correct swath, the projection of the geographical point is applied, using the meta-data linked to each swath.

Visual Question Answering on Multiple Remote Sensing Image Modalities To find the correct swath, the projection of the geographical point is applied, using the meta-data linked to each swath

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:57.996401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.818952Z digest=sha256:146dd1656e473e10e4631fe359ebcc0fdb44e4cbd2a63a864d481df7bbcb1b07

Observation 784ca467-5426-4cdd-b736-b85e603c9cbd · outbound

This paper cites The S1 images need to be debursted (removing of the black line and of the overlap) to get a continuous image before to extract the S1 patch that is inputted in the model.

Visual Question Answering on Multiple Remote Sensing Image Modalities The S1 images need to be debursted (removing of the black line and of the overlap) to get a continuous image before to extract the S1 patch that is inputted in the model

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:57.979691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.823227Z digest=sha256:c2eb74646f9fa6568144bf6d171d46617673a23f56f8e7c23a0152b60c84ce59

Observation 6b0b2b99-4316-4d9d-9902-264e0b7a7d45 · outbound

This paper cites A tail- value elimination procedure is performed on each chan- nel separately using statistics information extracted over the whole dataset.

Visual Question Answering on Multiple Remote Sensing Image Modalities A tail- value elimination procedure is performed on each chan- nel separately using statistics information extracted over the whole dataset

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:57.965667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:57.827375Z digest=sha256:404742f61e430fde2900551ff4673c9d2f48ddfa5d3a9813926a82cb53f95adb

Pith citing papers

No inbound Pith citation observations are available.