Pith. sign in

Paper Citation Record · LEDGER

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks

As of 12 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2412.02531.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.02531 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:25:10.376115Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact1
  • verified fuzzy35
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7071abf1-6c2f-486f-9014-4107d615da74 · outbound

This paper cites Deep learning for remote sensing image scene classification: A review and meta-analysis,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Deep learning for remote sensing image scene classification: A review and meta-analysis,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T23:25:10.229124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:25:10.229124Z digest=sha256:bcf22c961df886301098daa0f1dce4286450fcd278f619cab0e93708bf5f48bc

Observation 527e1dd8-255b-4880-bccb-718267df7479 · outbound

This paper cites Scene and environment monitoring using aerial imagery and deep learning,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Scene and environment monitoring using aerial imagery and deep learning,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.877827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.233407Z digest=sha256:4a9ad9daaf72616c86f8a598cf17d97a7c873e43e5fbbc24c433c20eb19b08d6

Observation 7212442e-d9ac-4b65-abb3-523425eee87a · outbound

This paper cites Spatial context-aware method for urban land use classification using street view images,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Spatial context-aware method for urban land use classification using street view images,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.868141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.237070Z digest=sha256:7f87d759657795ec30839a9fd44e620ab194d027f3ecfdc66cdc6216ce7186db

Observation 05b9dae7-3a75-4c1a-8286-27dd85fd1598 · outbound

This paper cites Urban scene understanding based on semantic and socioeconomic features: From high-resolution remote sensing imagery to multi-source geographic datasets,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Urban scene understanding based on semantic and socioeconomic features: From high-resolution remote sensing imagery to multi-source geographic datasets,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.857464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.240960Z digest=sha256:fd7581a86d7942f4a56a59d95edef445be062c840fa90984342adc56a0ead0ee

Observation 28541f4f-d411-41b0-b11e-a7d2d2c02d7e · outbound

This paper cites A survey on deep learning-driven remote sensing image scene understanding: Scene classification, scene retrieval and scene-guided object detection,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks A survey on deep learning-driven remote sensing image scene understanding: Scene classification, scene retrieval and scene-guided object detection,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.847730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.244642Z digest=sha256:922980110d8d9cf2671fd32c8fc75c3e54e58d4702b4b5528ea413df4f0a7e17

Observation 2410f6e5-b6ba-4c04-ac03-4d9cfec9ad72 · outbound

This paper cites Evaluating the potential of texture and color descriptors for remote sensing image retrieval and classification,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Evaluating the potential of texture and color descriptors for remote sensing image retrieval and classification,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.838031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.248249Z digest=sha256:509ad17f7090ba405ee225348d79a5462b3fa171881ec4a472f683ee68683ad8

Observation 010b672d-a5bd-4dbf-9510-83834df5fac8 · outbound

This paper cites Indexing of remote sensing images with different resolutions by multiple features,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Indexing of remote sensing images with different resolutions by multiple features,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.828162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.252196Z digest=sha256:60b4d3b68432af8c8035d967ccecae7a6822ef60f0e4be81cb4dc3a3e987e0cf

Observation 2a2bf74e-c86c-4c47-86cd-481d6a9a096f · outbound

This paper cites High-resolution satellite scene classification using a sparse coding based multiple feature combination,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks High-resolution satellite scene classification using a sparse coding based multiple feature combination,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.818269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.255697Z digest=sha256:5502d40b5d6aac3887e6575e87b4e4daddb16a813d8ed1490470891d0daa23c8

Observation ba22cb26-5dba-4745-ac5f-ac9a67248489 · outbound

This paper cites Remote sensing image scene classification meets deep learning: Challenges, methods, benchmarks, and opportunities,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Remote sensing image scene classification meets deep learning: Challenges, methods, benchmarks, and opportunities,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T23:25:10.259030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:25:10.259030Z digest=sha256:a98a74331912fddbd3fa4e3ae0c6427a5dfd9614741df12f3f774f958dd5544d

Observation 49046a35-f649-43e9-8e76-6bfea94693c2 · outbound

This paper cites Scene classification based on multiscale convolutional neural network,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Scene classification based on multiscale convolutional neural network,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.801839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.262380Z digest=sha256:60a87550061d46253362cd9b2ff9087797a58bcd20f4a238a512991774941ecf

Observation 0a0b01d1-c752-4d27-bf41-b941df70e6f3 · outbound

This paper cites Gan-based semisupervised scene classifi- cation of remote sensing image,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Gan-based semisupervised scene classifi- cation of remote sensing image,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.792178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.265554Z digest=sha256:4982b32963e8e3ade4a1b47056cc4c8d6aae927c17e69ff07de502d1ad539fe1

Observation e7b264d0-4fd4-4d44-8d17-6190cd718c63 · outbound

This paper cites Scvit: A spatial-channel feature preserving vision transformer for remote sensing image scene classification,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Scvit: A spatial-channel feature preserving vision transformer for remote sensing image scene classification,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.782374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.269031Z digest=sha256:42d08f3257f35a934809dfbb5a76603056bc52d1341cd4ac7f03886ba29b15b2

Observation e7619c2e-677c-4029-a957-1d1e4a755059 · outbound

This paper cites Vision transformer with contrastive learning for remote sensing image scene classification,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Vision transformer with contrastive learning for remote sensing image scene classification,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.772943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.272302Z digest=sha256:8944ee5101cd214ff337b0a155d343475a5fcda8562077cfdbc851c9cffdff15

Observation 17e4a056-52c8-47bd-a3ed-22c767a01698 · outbound

This paper cites Deep learning-a new approach for multi-label scene classification in plan- etscope and sentinel-2 imagery,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Deep learning-a new approach for multi-label scene classification in plan- etscope and sentinel-2 imagery,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.763313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.275625Z digest=sha256:6adb41531267d185fb83ae831a77685b400f9b3a0df335c1d989fccb957f3190

Observation 441c68e4-96bc-4e07-911f-356a7cf9cf4b · outbound

This paper cites A multi- level label-aware semi-supervised framework for remote sensing scene classification,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks A multi- level label-aware semi-supervised framework for remote sensing scene classification,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T23:25:10.278716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:25:10.278716Z digest=sha256:5458c571dcb9ae913028937b1e2c668ed378766af368d0181809cf2545e7d91a

Observation 2a250829-f0c9-41c7-ac3e-bfda466284fc · outbound

This paper cites He and Q.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks He and Q

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.747640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.281890Z digest=sha256:b7555d891bfc1c4d68e09a19ced9c93c329e0e56b3fa2fb8b04bd944f99456c1

Observation 711cc84a-e56f-4dbd-bfac-3865f06ea4dc · outbound

This paper cites A lightweight multi-scale crossmodal text-image retrieval method in remote sensing,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks A lightweight multi-scale crossmodal text-image retrieval method in remote sensing,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.737549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.285092Z digest=sha256:6ddb207d845d69f623aaba47f9563f2474116663485ab8373b0aac9f1dd804c4

Observation 48b40467-48a4-495b-9d03-6b880c100835 · outbound

This paper cites Multi-label semantic feature fusion for remote sensing image captioning,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Multi-label semantic feature fusion for remote sensing image captioning,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.727851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.288191Z digest=sha256:1aa748f956a06b22485cfad3c41e7d92bb2a461f17ec1e7d1145c509ea57145b

Observation 58169b2b-4533-4cab-ac3d-e44861e1151f · outbound

This paper cites Teaw: Text-aware few-shot remote sensing image scene classification,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Teaw: Text-aware few-shot remote sensing image scene classification,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.718106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.291364Z digest=sha256:1b24cbe432a6adb0fa7c544526e1768207628170f102934b8abb386a387b2cdc

Observation 4e715f2b-7443-4163-9a93-995601a9a425 · outbound

This paper cites Remoteclip: A vision language foundation model for remote sensing,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Remoteclip: A vision language foundation model for remote sensing,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T23:25:10.294762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:25:10.294762Z digest=sha256:c36f42992233dae9f741d25a5c39cbe5bd9084d976bdf6a80d075b37d4d5c363

Observation bbccb5ff-034d-402f-884d-77d00a688a07 · outbound

This paper cites Land-cover classification with high-resolution remote sensing images using transferable deep models,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Land-cover classification with high-resolution remote sensing images using transferable deep models,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.702609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.297929Z digest=sha256:6d59565e6b257dc10ab91e0b9c8cf654c518e62c6e5c538045add8cc6fc9f95a

Observation 4f9966a2-fa2b-4011-864c-199deac4b070 · outbound

This paper cites A dual- model architecture with grouping-attention-fusion for remote sensing scene classification,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks A dual- model architecture with grouping-attention-fusion for remote sensing scene classification,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.693529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.301005Z digest=sha256:fc7e33e6d7286773eabb85ba0cec8bede57d9c2aa8b5ca03ca486e357541f247

Observation bd3c34a4-8aa0-4bef-9741-1be01d262657 · outbound

This paper cites Scene classification with recurrent attention of vhr remote sensing images,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Scene classification with recurrent attention of vhr remote sensing images,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.684148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.304137Z digest=sha256:2404f3a7fdd46b0f7203aaaced5fb10f282123620872e94f21a56bea47d0d545

Observation 2dc92c41-72be-498f-8556-d556561b435b · outbound

This paper cites Classification of remote sensing images using efficientnet-b3 cnn model with attention,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Classification of remote sensing images using efficientnet-b3 cnn model with attention,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.675233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.307138Z digest=sha256:14806970c8927c00191904ef5a2d1dc7021a40962b6547cb1129f447c79056fb

Observation 7fb639c1-91dc-49ac-8e32-69e74a0c8722 · outbound

This paper cites Rs-deepsuperlearner: fusion of cnn ensemble for remote sensing scene classification,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Rs-deepsuperlearner: fusion of cnn ensemble for remote sensing scene classification,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.666022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.310320Z digest=sha256:617acd2c951825d93b6c2d36375db8f55063356a2d8fccfb52e18b6589eb3428

Observation f157f96c-6b35-445a-b4c4-d9fb71c36280 · outbound

This paper cites Contextual spatial-channel attention network for remote sensing scene classification,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Contextual spatial-channel attention network for remote sensing scene classification,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.656970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.313382Z digest=sha256:f3f9dc76728235a57f85dc5b48a83bd9225c4c3f62e6fc6c232592b358273f22

Observation 78461957-93eb-4f2d-a992-370049753e55 · outbound

This paper cites A local–global interactive vision trans- former for aerial scene classification,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks A local–global interactive vision trans- former for aerial scene classification,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.647236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.316600Z digest=sha256:21fe2da191b9d2dfaed2cb15b27c7350307045d3be5547b21f06e8a7a74d0604

Observation 753af8c8-49a9-410e-b950-ef4ead80cb6c · outbound

This paper cites Nir/rgb image fusion for scene classification using deep neural networks,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Nir/rgb image fusion for scene classification using deep neural networks,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.637260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.319967Z digest=sha256:bf93c0ddfdfa5940646e2bcfe59657bad316f5ffc7567845076dedb94b662f99

Observation c746d677-d954-478a-a1d1-2aa7ff562e09 · outbound

This paper cites Mrssc: a benchmark dataset for multimodal remote sensing scene classification,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Mrssc: a benchmark dataset for multimodal remote sensing scene classification,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.627429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.323038Z digest=sha256:5fbd8c04822b5a8cf0940ddf98c7b12a9b7a2280684894cc179d513f7375378d

Observation 63606dcf-7ef1-48ca-96c8-42c6ef88bc77 · outbound

This paper cites An optical image-aided approach for zero-shot sar image scene classifica- tion,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks An optical image-aided approach for zero-shot sar image scene classifica- tion,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.617723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.326298Z digest=sha256:bc8c14544266b579665ce729d60b82e7c6f036ae983a48ac597bb234a32312cd

Observation a1a2c3d8-c511-4dae-a7ec-1a7bdb13eaf7 · outbound

This paper cites Alignment and fusion using distinct sensor data for multimodal aerial scene classification,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Alignment and fusion using distinct sensor data for multimodal aerial scene classification,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.607260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.329375Z digest=sha256:f430a8f7dac79d1c5c842f944629040b02f42192f767265570731486bb17f0d0

Observation a75a814f-958d-4879-9e2a-90e740063b27 · outbound

This paper cites Text2seg: Remote sensing image semantic segmentation via text-guided visual foundation models,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Text2seg: Remote sensing image semantic segmentation via text-guided visual foundation models,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T23:25:10.332515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:25:10.332515Z digest=sha256:98323beb04832e9501d795615fddd0939d106addeba0153c82f25644d3994635

Observation 4de625c1-270d-4dbe-b0df-17d64887ab11 · outbound

This paper cites Parameter-efficient transfer learning for remote sensing image-text retrieval,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Parameter-efficient transfer learning for remote sensing image-text retrieval,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.597502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.335765Z digest=sha256:1b6ecc87f0a9a120054a049739304697ed7650322818b66dee1f78d00e345f80

Observation f7b8332f-69e6-4264-b83f-52d170207ec3 · outbound

This paper cites Interacting- enhancing feature transformer for cross-modal remote-sensing image and text retrieval,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Interacting- enhancing feature transformer for cross-modal remote-sensing image and text retrieval,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.587501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.338847Z digest=sha256:4ea840079c203d4810f7e78394ef7dc8292fea1be8c3c8492db2b83ef90785c7

Observation 1f16899c-85ef-45dd-a398-b31376b6ccbb · outbound

This paper cites Few-shot remote sensing image scene classification: Recent advances, new baselines, and future trends,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Few-shot remote sensing image scene classification: Recent advances, new baselines, and future trends,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.577824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.341993Z digest=sha256:245f9e4640f891c2d87e05a11eef43959563ad1edd5f78807f7e28a6d0768a83

Observation 18f3fbe3-8b87-4d28-a26f-8ac5abeffb9c · outbound

This paper cites Few-shot medical image classification with simple shape and texture text descriptors using vision-language models.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Few-shot medical image classification with simple shape and texture text descriptors using vision-language models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T23:25:10.345426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:25:10.345426Z digest=sha256:93f7e5fd1895540be1c6abb874fef624bda92267b534bbd5b7181ab454a3182a

Observation c94b1f13-b0d8-48a9-860f-f6d705d5e643 · outbound

This paper cites Vision-language model for generating textual descriptions from clinical images: model development and validation study,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Vision-language model for generating textual descriptions from clinical images: model development and validation study,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.567331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.349243Z digest=sha256:c0fb612f00502da17c3b5562be177c30e9f0da81ef864198917be6a2c043e31d

Observation da38a411-eb5e-4bb6-9795-dbdbf2926c31 · outbound

This paper cites Medical image understanding with pretrained vision language models: A comprehensive study,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Medical image understanding with pretrained vision language models: A comprehensive study,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.557266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.352529Z digest=sha256:a0c52768fd7b2306489bdf0f064b44d4bd5b26599899950bc65a73e2373add36

Observation 64a15a57-9d45-4872-9e6a-b2d36ecd8e68 · outbound

This paper cites VLM-CPL: Consensus Pseudo Labels from Vision-Language Models for Annotation-Free Pathological Image Classification.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks VLM-CPL: Consensus Pseudo Labels from Vision-Language Models for Annotation-Free Pathological Image Classification

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-11T23:25:10.419574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.355724Z digest=sha256:6abbbe3de21b7e188855c66dd03ca58a424d088cb512ad60e3979ee745d60777

Observation 42cf6449-0ff5-4167-a4f4-40965193b14d · outbound

This paper cites Improved zero-shot classification by adapting vlms with text descriptions,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Improved zero-shot classification by adapting vlms with text descriptions,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.547199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.359331Z digest=sha256:71e7b03e5883b6c7f7666fdd90f4619c128557cf32e969e81a5b5d736515045d

Observation 33fd1be0-9798-49a7-b681-d2b250db52ac · outbound

This paper cites Exploiting lmm-based knowledge for image classification tasks,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Exploiting lmm-based knowledge for image classification tasks,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.536941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.362556Z digest=sha256:50479ed3318fbfe7e6fb3eb145a7855b8d709300848c0c65bfa3e21ed073e7a9

Observation 72d59773-5176-4f51-8408-cf63cff6aafa · outbound

This paper cites Visual instruction tuning,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Visual instruction tuning,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:25:10.526594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T23:25:10.366014Z digest=sha256:cf2c1d44de3e49b4500133128a4841b36609e1a39f5766e5bc3a3c8b7e57a5df

Observation 80dc399b-2693-4174-ad2f-2f4381ea9fdf · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T23:25:10.369305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:25:10.369305Z digest=sha256:383de5cbcdb45c8492d8f0513161b6a24ada36f9eace7e151cda1bfb499e0fc5

Observation 6a9a74f2-f280-466b-a6cb-68413ec2a2f5 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Learning transferable visual models from natural language supervision,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T23:25:10.372826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:25:10.372826Z digest=sha256:b597a1939d6b00e2e24e1f3a03777302e5f8a1beb7ecba3c842dcdbe3225f108

Observation f43a904c-1605-4580-8050-66cb49da067f · outbound

This paper cites Segment anything,.

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks Segment anything,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T23:25:10.376115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:25:10.376115Z digest=sha256:9526f62a7b2be9290a33c1ea0802c64769268198c6547b827e702377d5b4c928

Pith citing papers

No inbound Pith citation observations are available.