Pith. sign in

Paper Citation Record · LEDGER

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives

As of 9 August 2026, this Paper Citation Record lists 100 of 299 outbound references and 1 inbound Pith citation observation for arXiv:2505.14361.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14361 v1

Coverage vector

measured 100 of 299 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:38:55.619271Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T05:11:58.947649Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T10:45:42.279098Z

Reference resolution

100 of 299 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved99
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3c13e333-f01a-4f75-a0a5-affae0186252 · outbound

This paper cites Rsgpt: A remote sensing vision language model and benchmark,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Rsgpt: A remote sensing vision language model and benchmark,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:43.064939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:43.064939Z digest=sha256:38cc99c4b826f352ccd6ff815744860f34fc81263b7bec6aeeb81b04f17540f1

Observation 4d061b49-c907-4ddc-8d88-e98eff39cea3 · outbound

This paper cites Remoteclip: A vision language foundation model for remote sensing,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Remoteclip: A vision language foundation model for remote sensing,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:43.156146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:43.156146Z digest=sha256:8fff5b5962f4e2bfafbf668f95c03cdefc0e2fb56f5276bd268564fa7e627db9

Observation 28528bda-18a9-4da0-9c19-e9106bba1398 · outbound

This paper cites Geochat: Grounded large vision-language model for remote sensing,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Geochat: Grounded large vision-language model for remote sensing,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:43.283004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:43.283004Z digest=sha256:ae22316bfced4f79f699d3489df4b0d0d6563f6f122396e083e3e6b086faa5fd

Observation 7edb7424-6f49-4b1a-844d-b3c0ca45e78f · outbound

This paper cites Remote sensing vision-language foundation models without annotations via ground remote alignment,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Remote sensing vision-language foundation models without annotations via ground remote alignment,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:43.452270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:43.452270Z digest=sha256:ea48647fb2a2fc72eba97307a889a5f1bc36b97cf76e9e6d10a404a143d043d0

Observation 80c7a8fa-64c7-4355-a09c-ba621b435a22 · outbound

This paper cites Skyeyegpt: Unifying remote sens- ing vision-language tasks via instruction tuning with large language model,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Skyeyegpt: Unifying remote sens- ing vision-language tasks via instruction tuning with large language model,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:43.608411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:43.608411Z digest=sha256:17fcdc4413793bb3cd934c4bccc25100e44c2ad84db3d7a2fbe4a3ed28c71eaf

Observation 8317d75c-4ee6-4c4c-aa23-a0b09e186326 · outbound

This paper cites Earthgpt: A universal multimodal large language model for multisensor image com- prehension in remote sensing domain,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Earthgpt: A universal multimodal large language model for multisensor image com- prehension in remote sensing domain,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:43.778323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:43.778323Z digest=sha256:e799af010081221d22036d2c5c241d8fe1490fe401f26041431223724fc7e1e1

Observation 9f968775-1c81-4dba-934d-ea48443625d5 · outbound

This paper cites Rs5m and georsclip: A large scale vision-language dataset and a large vision-language model for remote sensing,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Rs5m and georsclip: A large scale vision-language dataset and a large vision-language model for remote sensing,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:43.900412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:43.900412Z digest=sha256:1167a3432764cf40311ba34c737abf7875f2d92369817780d285c1e409164ada

Observation 7991c0ca-874f-437a-bf88-f02088e1cb2d · outbound

This paper cites Skyscript: A large and semantically diverse vision-language dataset for remote sensing,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Skyscript: A large and semantically diverse vision-language dataset for remote sensing,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:44.024239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:44.024239Z digest=sha256:2fdb5415f0824fc71cf5a7e61c59f73cfcb8258cdf5c6c5c99aa6605163bfe92

Observation 85e592d4-8a4d-4a9b-8181-be5747a519a5 · outbound

This paper cites Lhrs-bot: Empowering remote sensing with vgi-enhanced large multimodal language model,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Lhrs-bot: Empowering remote sensing with vgi-enhanced large multimodal language model,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:44.181611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:44.181611Z digest=sha256:6882e378c86d9bceaeee3e3f38d08453c961e9b3b86a0026fe2da26aa2718646

Observation 6cf05be3-18b1-49f7-9990-60f465189842 · outbound

This paper cites Rs- llava: A large vision-language model for joint captioning and question answering in remote sensing imagery,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Rs- llava: A large vision-language model for joint captioning and question answering in remote sensing imagery,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:44.289917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:44.289917Z digest=sha256:6d2db404a4e026a735176e2cd811dbd5a4b92535e7345f9d46ef3f61e2c8ed9d

Observation 9ae29f85-5d1f-4171-9a1e-bd4620c1b678 · outbound

This paper cites SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:44.401424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:44.401424Z digest=sha256:461ec1612233b5d61dbc8ff0b14224a4d2204006bdc75890f4d4669ab46b9bfb

Observation ff8d6127-e1ca-42d2-a67d-78004064ca42 · outbound

This paper cites Vrsbench: A versatile vision- language benchmark dataset for remote sensing image understanding,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Vrsbench: A versatile vision- language benchmark dataset for remote sensing image understanding,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:44.510141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:44.510141Z digest=sha256:d88be4443c2103978faf145b85eca78facdd59445c0efda9aae17980421a36c1

Observation 7e6441a9-928e-4ed8-ab1a-d9411385a790 · outbound

This paper cites Changeclip: Remote sens- ing change detection with multimodal vision-language representation learning,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Changeclip: Remote sens- ing change detection with multimodal vision-language representation learning,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:44.622345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:44.622345Z digest=sha256:7dc94068165ee79f36a31d3935f06763299807b2c34e44e04870ee68e8baf578

Observation 21ec66ba-65ee-4a29-8c3a-e81b9062f7dc · outbound

This paper cites Rs-clip: Zero shot remote sensing scene classification via contrastive vision-language supervision,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Rs-clip: Zero shot remote sensing scene classification via contrastive vision-language supervision,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:44.787635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:44.787635Z digest=sha256:fe3c49e45ad679bf736482cb0735e281f6478a910b572a2bb479bd151cdfffa2

Observation cb51dcff-43d7-4a42-8371-a7f4ee1c1383 · outbound

This paper cites Mind the Modality Gap: Towards a Remote Sensing Vision-Language Model via Cross-modal Alignment.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Mind the Modality Gap: Towards a Remote Sensing Vision-Language Model via Cross-modal Alignment

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:44.914015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:44.914015Z digest=sha256:882cab424a2a6f2e5b92c4eb0f7279e615a389663d53c99526e96199d59dd3aa

Observation c5713cf3-7683-4486-8685-24b06828bc24 · outbound

This paper cites Towards vision-language geo-foundation model: A survey,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Towards vision-language geo-foundation model: A survey,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:45.014750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:45.014750Z digest=sha256:04a69b0bda8192da85088af7845ef9875b86237a44cccd44e464378ae0059bae

Observation 58a77bdf-717b-4601-bf47-2ba6fece9281 · outbound

This paper cites Vision-language models in remote sensing: Current progress and future trends,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Vision-language models in remote sensing: Current progress and future trends,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:45.170587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:45.170587Z digest=sha256:c42b38ed24e7fbc197ecda9e0d6aa9c1f0c9a5d78ce703532c1ddee9586999f5

Observation 2ddba138-6449-4a73-a96c-46fe659847aa · outbound

This paper cites S-clip: Semi-supervised vision- language learning using few specialist captions,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives S-clip: Semi-supervised vision- language learning using few specialist captions,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:45.235768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:45.235768Z digest=sha256:360f86d1c45cdf62a214f8bd2c55fd34486f51e38aee6f546ce347979685a419

Observation 962fc3d3-1b70-4987-8969-f0e56291152f · outbound

This paper cites Multi-view feature fusion and visual prompt for remote sensing image captioning,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Multi-view feature fusion and visual prompt for remote sensing image captioning,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:45.332966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:45.332966Z digest=sha256:65ae9977792d05d22f63a0574d73aa16a0179dc9291af50cb6097c664b470bed

Observation 34fe4cf5-db90-46af-aa20-a628f1c79d1a · outbound

This paper cites Vision- language models for zero-shot classification of remote sensing images,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Vision- language models for zero-shot classification of remote sensing images,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:45.468804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:45.468804Z digest=sha256:b7b58fe699b4b7e664282f8e4386245e93e51ad0203cd192bed935b33f457c9f

Observation 4591f1d6-c747-4afb-a06f-509911022cc7 · outbound

This paper cites Detecting cloud presence in satellite images using the rgb-based clip vision-language model,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Detecting cloud presence in satellite images using the rgb-based clip vision-language model,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:45.646020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:45.646020Z digest=sha256:c897ead8ff5f14f49c2e1e0f05c8eba544cbe54f55afac852c50ade443624a81

Observation 8fa01820-6df0-413f-8c19-17a775d5a98f · outbound

This paper cites Chatearthnet: a global- scale image-text dataset empowering vision-language geo-foundation models,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Chatearthnet: a global- scale image-text dataset empowering vision-language geo-foundation models,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:45.809267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:45.809267Z digest=sha256:c28b81cdbff1c93998b3d863dc6d9f55a898d4f4c3ac87f9b5bb847c5ac6a6a1

Observation 9e13967b-f96d-4e1f-8a9b-3c50c95f9267 · outbound

This paper cites Bi-modal transformer-based approach for visual question answering in remote sensing imagery,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Bi-modal transformer-based approach for visual question answering in remote sensing imagery,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:45.933746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:45.933746Z digest=sha256:54cd6431a580add597dd65516f47aff658f1e1d437da27bb5884993155d38853

Observation e010fd27-cd52-4b66-9022-b49a6f7b66c6 · outbound

This paper cites Popeye: A unified visual-language model for multi-source ship detection from remote sensing imagery,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Popeye: A unified visual-language model for multi-source ship detection from remote sensing imagery,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:46.052643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:46.052643Z digest=sha256:80dcc3003f18bcc54067c91e1382ba7b9187a8d1c4f918fcfd72d1f2187a1908

Observation f5d64ec1-2d08-49b0-9e2b-c884b2e27c38 · outbound

This paper cites Deep semantic-visual alignment for zero-shot remote sensing image scene classification,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Deep semantic-visual alignment for zero-shot remote sensing image scene classification,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:46.198276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:46.198276Z digest=sha256:5625e4878df37564c0b7c825f128b034baa1351bcb6a8788dde810e51cd0a42a

Observation 50a34e47-9fbc-4ba5-b65b-6eb24e18de1f · outbound

This paper cites Set-CLIP: Exploring Aligned Semantic From Low-Alignment Multimodal Data Through A Distribution View.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Set-CLIP: Exploring Aligned Semantic From Low-Alignment Multimodal Data Through A Distribution View

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:39:23.164372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:38:46.336526Z digest=sha256:40b1d7e0f85c2c1a4f9b94589b95c781360dbcfd9a94b972fd9d5813b8925f97

Observation 069b0439-49c8-441e-9d6e-9e43c7c134ad · outbound

This paper cites Segment change model (scm) for unsupervised change detection in vhr remote sensing images: a case study of buildings,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Segment change model (scm) for unsupervised change detection in vhr remote sensing images: a case study of buildings,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:46.502942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:46.502942Z digest=sha256:ef449b48db903a30b46f043f232613c4b56592444618e7e6cd27392682d1931f

Observation dc1962f1-fe34-4f97-86b0-e2b6d62e2dde · outbound

This paper cites A new learning paradigm for foundation model-based remote-sensing change detection,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives A new learning paradigm for foundation model-based remote-sensing change detection,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:46.641308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:46.641308Z digest=sha256:4c6e073e1cd2277f8c4b6971216b80d8045d8ab7a2b10f7cf99d1fc07e94a745

Observation 2cc03d95-bcea-4853-aaac-ce83c706e8de · outbound

This paper cites Vlca: vision-language aligning model with cross-modal attention for bilingual remote sensing image captioning,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Vlca: vision-language aligning model with cross-modal attention for bilingual remote sensing image captioning,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:46.750482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:46.750482Z digest=sha256:42c69902228bab0e8ad51bdcb57eb8a68604567c4a0f9b9324ee82c2f8c782bb

Observation 1c232cf3-5270-4340-81cf-8ad1dd03c2ae · outbound

This paper cites RS-GPT4V: A Unified Multimodal Instruction-Following Dataset for Remote Sensing Image Understanding.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives RS-GPT4V: A Unified Multimodal Instruction-Following Dataset for Remote Sensing Image Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:46.888234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:46.888234Z digest=sha256:8b578c618ee77a884c3f03fdb076b3404cd1bc8f90fe28005567bf2034b96c95

Observation 0ddae0ae-3251-4bcb-a412-0779d4648ae9 · outbound

This paper cites Luojiahog: A hierarchy oriented geo-aware image caption dataset for remote sensing image–text retrieval,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Luojiahog: A hierarchy oriented geo-aware image caption dataset for remote sensing image–text retrieval,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:47.040530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:47.040530Z digest=sha256:c6d65ffda7d6e18107ccf2f0bf181c87cb74133684ef166fca0dff6853f17ef2

Observation 2df3449b-f06b-4bfa-9f43-30aabb12c296 · outbound

This paper cites SATIN: A Multi-Task Metadataset for Classifying Satellite Imagery using Vision-Language Models.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives SATIN: A Multi-Task Metadataset for Classifying Satellite Imagery using Vision-Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:47.152096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:47.152096Z digest=sha256:186a6c639653b70ee12fc8b314e3e237c49443279d08df70eda8903ffd50349a

Observation 5f9f3b49-df9d-41cb-84d2-ad5c436dd404 · outbound

This paper cites Rsvg: Exploring data and models for visual grounding on remote sensing data,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Rsvg: Exploring data and models for visual grounding on remote sensing data,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:47.311652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:47.311652Z digest=sha256:86fc9ea62a0ee76d13062a4e9aa7383073c574165a7592286a986ca79380e5fd

Observation 7fca31d7-7a71-41d2-8110-f68cd29bebeb · outbound

This paper cites Good at captioning bad at counting: Bench- marking gpt-4v on earth observation data,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Good at captioning bad at counting: Bench- marking gpt-4v on earth observation data,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:47.462892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:47.462892Z digest=sha256:a6c849ab8d777fdab76779da704400f303df9e22cfa76f24f62dea0cf408c5c7

Observation 240f4edb-64b2-4d99-8446-17152f569de6 · outbound

This paper cites Satclip: Global, general-purpose location embeddings with satellite imagery,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Satclip: Global, general-purpose location embeddings with satellite imagery,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:47.610998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:47.610998Z digest=sha256:21e5454d0440a0a8fb0ae485bb23ebb17371488436f428e162c94cee2f0ecb63

Observation c9e19baa-94ff-4442-b24d-75d4b146fd5b · outbound

This paper cites GeoCLIP: Clip-inspired alignment between locations and images for effective worldwide geo- localization,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives GeoCLIP: Clip-inspired alignment between locations and images for effective worldwide geo- localization,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:47.742008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:47.742008Z digest=sha256:a1e737e165a92667cfee0b4dba312afc008ce027edf66363fe31eb99a79f490d

Observation b222e3fc-c3be-4be3-9492-1c3235c67cf6 · outbound

This paper cites Csp: Self-supervised contrastive spatial pre-training for geospatial-visual representations,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Csp: Self-supervised contrastive spatial pre-training for geospatial-visual representations,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:47.857469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:47.857469Z digest=sha256:388ea39d6529bfebb94e08e0899b9f040d5aabc25d55d7eb323530885229d496

Observation 5d3d29a9-925a-4789-a9e6-abb55d56fc23 · outbound

This paper cites Diffusionsat: A generative foundation model for satellite imagery,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Diffusionsat: A generative foundation model for satellite imagery,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:47.942387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:47.942387Z digest=sha256:c72b29c78bdfc6642819eb606099f93675074baa484815cc7f42abfb2bfc361c

Observation d0624a66-64f5-40d7-a41d-597b684d9ab1 · outbound

This paper cites Crs-diff: Controllable remote sensing image generation with diffusion model,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Crs-diff: Controllable remote sensing image generation with diffusion model,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:48.144141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:48.144141Z digest=sha256:10a96de4756c37befe1b64031b602143e63781ff8052f7efccd3163ac68704ec

Observation 431fd769-708a-4b4e-a89a-6f1e0aa122f9 · outbound

This paper cites Practical techniques for vision- language segmentation model in remote sensing,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Practical techniques for vision- language segmentation model in remote sensing,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:48.276207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:48.276207Z digest=sha256:294c6ed096b1fd276e559cc4dad994f40346b87d959eab6a6cf2dc03860280c6

Observation 023dc923-3566-4b26-b689-2de62c972118 · outbound

This paper cites Text2seg: Zero-shot remote sensing image semantic segmentation via text-guided visual foundation models,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Text2seg: Zero-shot remote sensing image semantic segmentation via text-guided visual foundation models,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:48.421970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:48.421970Z digest=sha256:740c9a6f0c2bc1a73f9923320d79805d544621e0b84764c6b42fc560798e074a

Observation 261f59c7-267d-4a7e-96e2-a68940a51f15 · outbound

This paper cites A decoupling paradigm with prompt learning for remote sensing image change captioning,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives A decoupling paradigm with prompt learning for remote sensing image change captioning,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:48.564888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:48.564888Z digest=sha256:5c604983c8ef4866d8751041548fb96b9abac206d1aa833967a0c36a63a1aad1

Observation 97d41679-a69e-41ca-81bb-4d27abffb58a · outbound

This paper cites Bootstrapping interactive image-text alignment for remote sensing image captioning,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Bootstrapping interactive image-text alignment for remote sensing image captioning,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:48.738237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:48.738237Z digest=sha256:28afe55d0e91c55bff398f4e2b5f7d4da70b86197c146d1a9a7da8dddd69a7da

Observation 92996c50-923f-42e7-8516-296d00d0c1d0 · outbound

This paper cites Text- guided diverse image synthesis for long-tailed remote sensing object classification,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Text- guided diverse image synthesis for long-tailed remote sensing object classification,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:48.945772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:48.945772Z digest=sha256:2a39dbff90cfcfbfab8aecc270be1ee0b47982ab64fdb361f11ca71402268316

Observation 16f4bd2c-0ffd-4064-b3b6-752a3595ac07 · outbound

This paper cites Object detection in aerial images: A large-scale benchmark and challenges,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Object detection in aerial images: A large-scale benchmark and challenges,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:49.158228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:49.158228Z digest=sha256:c451941573a98e971894dab1ef9a8fc2d5d7e6becf21f73c12a71d4f626673b0

Observation 977c3415-671e-43ac-87f8-48948426a08b · outbound

This paper cites InstructBLIP: Towards general-purpose vision-language models with instruction tuning,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives InstructBLIP: Towards general-purpose vision-language models with instruction tuning,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:49.305788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:49.305788Z digest=sha256:6a74ae382df086d98765642a22354b918a1c89d517cbb7d95289939403fe424e

Observation 2865e876-251a-47f7-8822-0dba7d917d7c · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:49.424549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:49.424549Z digest=sha256:c2ff51d5009ec748643d4d6396ea2d1bd0bb13d4dfeb2c74ee22652d77966a5a

Observation dd525767-374d-41b4-90ae-70d878c23db1 · outbound

This paper cites Learning to rank question answer pairs with holographic dual lstm architecture,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Learning to rank question answer pairs with holographic dual lstm architecture,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:49.623916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:49.623916Z digest=sha256:6dad1c1ae0888e4f34a396bb2a1859d9e062b56f25a644a8bfc987f7e1bd2e87

Observation 83d2cb6a-951f-4d5b-a5ad-3acb6ec24634 · outbound

This paper cites Improved baselines with visual instruction tuning,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Improved baselines with visual instruction tuning,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:49.777690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:49.777690Z digest=sha256:a6bf6b22aa4254ec50b6eea415d3d2ae4ae2e1c8736b0c3fc2f418a1acbd5fa4

Observation 9c5386c5-479e-4fd6-b4ee-5584bb0148ef · outbound

This paper cites Lora: Low-rank adaptation of large language models,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Lora: Low-rank adaptation of large language models,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:49.897162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:49.897162Z digest=sha256:bf5dcd2f7db4fa47fa8b1b412877b3936a577b2d7ebdc4d07876fb38e3295698

Observation e2d88973-a7c8-495d-926d-da6955896f03 · outbound

This paper cites Object detection in optical remote sensing images: A survey and a new benchmark,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Object detection in optical remote sensing images: A survey and a new benchmark,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:49.957361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:49.957361Z digest=sha256:2fd71dd224e28da424d011c3da0febdb04a30ff41ce933193f2476ca04702c5e

Observation c585a361-a717-4c94-a5f9-7885947955af · outbound

This paper cites Fair1m: A benchmark dataset for fine-grained object recognition in high-resolution remote sensing imagery,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Fair1m: A benchmark dataset for fine-grained object recognition in high-resolution remote sensing imagery,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:50.035586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:50.035586Z digest=sha256:8a6ba985965ede763fe68c09ab5030c6ed003b20b92b67e7297f2ba7ddf293b3

Observation 1929c793-e415-4b4d-9a32-587ad25b5520 · outbound

This paper cites Remote sensing image scene classifi- cation: Benchmark and state of the art,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Remote sensing image scene classifi- cation: Benchmark and state of the art,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:50.099116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:50.099116Z digest=sha256:ff32f3cf37b8150212e44ed3254903ace1ab65adc1b6c6c52f26c8c09e36e9ab

Observation 0a3af32e-b7a6-4bd1-a80a-7a6cdc3a5518 · outbound

This paper cites Rsvqa: Visual question answering for remote sensing data,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Rsvqa: Visual question answering for remote sensing data,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:50.178928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:50.178928Z digest=sha256:f63bbe3e5da0c85c44a76c1590d74e6dfbea7abfa9227a23ae8578e817b19628

Observation 7dc08abe-90cb-4c6b-b0d0-77e74f7a23d9 · outbound

This paper cites Floodnet: A high resolution aerial imagery dataset for post flood scene understanding,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Floodnet: A high resolution aerial imagery dataset for post flood scene understanding,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:50.251075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:50.251075Z digest=sha256:b413245ab38063324bd045cf8171ebaf3bc076dd11c7206f65ed19b32ab9fab9

Observation 5500eab6-e1e9-4c38-9631-29edbf09d0df · outbound

This paper cites Eva: Exploring the limits of masked visual representation learning at scale,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Eva: Exploring the limits of masked visual representation learning at scale,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:50.331731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:50.331731Z digest=sha256:5014c34c185b0b7423aec45e8b50e9036fd5465251a5c47fb10876b497cc96ed

Observation 1c4f9e10-9e25-444d-86ff-83b58f8829a4 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:50.389463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:50.389463Z digest=sha256:6ab81c480860590277bdba5634962eb88bb1463f2befb46ebe9e245565fb07e3

Observation d4b0ffc7-3a98-4cb9-96ee-0dd05d2b16e0 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives LLaMA: Open and Efficient Foundation Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:50.452910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:50.452910Z digest=sha256:2740e10e464c3cce223bf482326a30299940cad19f24e7ea262bd9251c3867dc

Observation c88bc01b-3d2c-4893-8763-3df3bf9063a1 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:50.540598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:50.540598Z digest=sha256:6315cca09861681db2ae826e033d5519706a685368ce5abb9e90b92474584d68

Observation 31a8632f-e268-4e83-93ab-9ab05bc6b133 · outbound

This paper cites Exploring models and data for remote sensing image caption generation,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Exploring models and data for remote sensing image caption generation,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:50.620624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:50.620624Z digest=sha256:7135fb66583031359791c88d405466661056b43e83c397223e97766f42fb6da0

Observation 1440f466-6029-4b44-b24c-01920da4a6d2 · outbound

This paper cites Exploring a fine-grained multiscale method for cross-modal remote sensing image retrieval,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Exploring a fine-grained multiscale method for cross-modal remote sensing image retrieval,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:50.727487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:50.727487Z digest=sha256:7a304494cf0268a041315fdb318eb2e4ec0e2c7292e1bad0af95778df44862c4

Observation 89e8dec0-2306-41a3-8e16-e1017e737596 · outbound

This paper cites Deep semantic understanding of high resolution remote sensing image,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Deep semantic understanding of high resolution remote sensing image,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:50.886595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:50.886595Z digest=sha256:47edbd56e0a6d9029e28268f7c103c591e2db157d4612416a0ab782b1269d292

Observation d18e7c63-ea7e-4c00-8f97-e88ef95621ac · outbound

This paper cites Nwpu- captions dataset and mlca-net for remote sensing image captioning,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Nwpu- captions dataset and mlca-net for remote sensing image captioning,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:51.023865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:51.023865Z digest=sha256:d24ae55fc6cda317e84e97d134201f906ae18d94334095e67dee12e8cad99604

Observation 7f3341c5-52cf-46e0-8ea0-0256fed5e53e · outbound

This paper cites Capera: Captioning events in aerial videos,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Capera: Captioning events in aerial videos,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:51.137450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:51.137450Z digest=sha256:660af77dd0bf54579b5a59a6f6e516b73a154fa05ea8ff5900b70830f100d221

Observation 25f376d8-5c9c-435d-bc12-7a7ed19908a3 · outbound

This paper cites Mutual attention inception network for remote sensing visual question answering,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Mutual attention inception network for remote sensing visual question answering,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:51.255800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:51.255800Z digest=sha256:a21f0ea06ed1e7931499855cc9bd3c2805bad2c5811be40e80528706d2f6e0ab

Observation b6989cdd-a7e0-46fe-a1e1-60f529be1bc9 · outbound

This paper cites Visual grounding in remote sensing images,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Visual grounding in remote sensing images,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:51.389367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:51.389367Z digest=sha256:e64cc488536322e6d932d06f0a1dc860ad2e40ddc8e35d60a7111192f29ea366

Observation a2e78898-96b9-4583-a0a5-d2aba730a934 · outbound

This paper cites Dinov2: Learning robust visual features without supervision,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Dinov2: Learning robust visual features without supervision,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:51.485572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:51.485572Z digest=sha256:aab57d7252bc6e0fda9e083e9098ddf6f06891c9d42bc58caa3295df01df7377

Observation 9d34bc30-b6d2-4f1d-aebf-8a53983c8299 · outbound

This paper cites Multistep question-driven visual question answering for remote sensing,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Multistep question-driven visual question answering for remote sensing,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:51.603773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:51.603773Z digest=sha256:bce081f12bd102ac3723ead7ea2a40ffd72142ed24ebd0cf10664e1cdf1556c8

Observation ccf42811-0ac5-4cd3-8ef2-9250060d525c · outbound

This paper cites Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:51.743120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:51.743120Z digest=sha256:eeff8f8b3403f178c9fbb0b7a9c18505e0f6c3920967c0442da417a3e9f4905c

Observation 3ea23881-6c26-41c5-9df3-ad95e3888724 · outbound

This paper cites Bag-of-visual-words and spatial extensions for land-use classification,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Bag-of-visual-words and spatial extensions for land-use classification,

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:51.805151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:51.805151Z digest=sha256:f9b8488dc297206018ce7e0b7c70c5a977778faf78e700388dccdbb45fcd0c79

Observation 58fa19a6-1021-4da3-bed6-8d4e35b57de7 · outbound

This paper cites Satellite image classification via two-layer sparse coding with biased image representation,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Satellite image classification via two-layer sparse coding with biased image representation,

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:51.888925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:51.888925Z digest=sha256:d3d2258aa2b87e7d2c531c09c0337181002857a0b1dc9186357c569867a5b211

Observation 7c38ff30-6b01-4a80-b7e1-d6e732b0cb26 · outbound

This paper cites Deep learning based feature selection for remote sensing scene classification,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Deep learning based feature selection for remote sensing scene classification,

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:52.018569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:52.018569Z digest=sha256:77b833ad9d146447435dafdcc879b4482c28e3df17cb09a173c23434a2e8d874

Observation 10335210-ebb5-4d98-9f17-e7c17fa4500d · outbound

This paper cites A public dataset for ship classification in remote sensing images,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives A public dataset for ship classification in remote sensing images,

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:52.137785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:52.137785Z digest=sha256:5c13cd2cf397d13488cef8e3c7c177d022fcce463df7b3c0839ce6262d880116

Observation feb5e806-9622-4061-ad42-36407c2f244d · outbound

This paper cites A public dataset for fine-grained ship classification in optical remote sensing images,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives A public dataset for fine-grained ship classification in optical remote sensing images,

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:52.252457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:52.252457Z digest=sha256:27de71b415aa9318321768a36b74807ab37a420031e7da76a5eff8cc6ec2d3b0

Observation 0d191b95-75dc-4fbe-845b-24ce179a9a42 · outbound

This paper cites Rotation-insensitive and context- augmented object detection in remote sensing images,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Rotation-insensitive and context- augmented object detection in remote sensing images,

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:52.381176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:52.381176Z digest=sha256:3863a1424d0f2a8322ca0193dd36e34260803351bdb53dc44befcd892ba5141a

Observation 987be9d5-082a-450f-a7de-9a3295fa2199 · outbound

This paper cites Hierarchical and robust convolutional neural network for very high-resolution remote sensing object detection,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Hierarchical and robust convolutional neural network for very high-resolution remote sensing object detection,

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:52.474659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:52.474659Z digest=sha256:30f448fec1cd0eee15f16ed9600f9aeca3348823a5b0767357f1865ca142dbdc

Observation 38664c2c-4e0d-44bc-bc9f-1bfaab473bec · outbound

This paper cites Orientation robust object detection in aerial images using deep convolutional neural network,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Orientation robust object detection in aerial images using deep convolutional neural network,

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:52.600351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:52.600351Z digest=sha256:85f919a4b9a0d053fac6ef28e7beb615a5cef99b06733796cf97e70f0c51348b

Observation 81d5eaa9-da89-4d24-89c5-e57d9c5408c1 · outbound

This paper cites Detection and tracking meet drones challenge,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Detection and tracking meet drones challenge,

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:52.772576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:52.772576Z digest=sha256:e18bf4d9d69ec9283d90c2253bc87ce2b5eb23f0799959f9cac5c2d13b8da0a1

Observation d30c5c91-1183-494a-82a4-8f749811cec9 · outbound

This paper cites Sar ship detection dataset (ssdd): Official release and comprehensive data analysis,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Sar ship detection dataset (ssdd): Official release and comprehensive data analysis,

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:52.902513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:52.902513Z digest=sha256:3f3291833f623414b7cd476b52812a1a11b3ed1669e737ba53e11c2a6d26710d

Observation d759d29c-a0ae-45d9-881c-8c582e48f16b · outbound

This paper cites Hit-uav: A high-altitude infrared thermal dataset for unmanned aerial vehicle- based object detection,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Hit-uav: A high-altitude infrared thermal dataset for unmanned aerial vehicle- based object detection,

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:53.061271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:53.061271Z digest=sha256:42e24b38458dc3e773e0b2db59e012e87cd0fc64dbdaf76879689616c0fb5fc3

Observation 208bc1da-2418-4272-9cf4-39f93be24714 · outbound

This paper cites Sea-shipping,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Sea-shipping,

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:53.201727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:53.201727Z digest=sha256:6693eb905944a762fd1ffa8686b40ae1d63bc16c3c07eb101c87ae3eb6132810

Observation 799f5ea0-37b9-4163-8f3e-e1e650e6c49d · outbound

This paper cites Infrared-security,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Infrared-security,

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:53.334467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:53.334467Z digest=sha256:07f684aa912b4870ea85585c30d6101c2d56cf52afab5c1e430cec04b9d6a009

Observation 22d82886-fbba-4fa1-aa18-76e1ab140edd · outbound

This paper cites Aerial-mancar,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Aerial-mancar,

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:53.458005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:53.458005Z digest=sha256:586f7cb798950e6c1392f4d6f3bb4c94b69f029f6e4c61593a9ee9eb0eb216b3

Observation 934f876c-0fa5-459f-9829-7b47cf2efbb0 · outbound

This paper cites Double-light-vehicle,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Double-light-vehicle,

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:53.615236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:53.615236Z digest=sha256:deb4d0399dc52c53747913b674efb67a543724a41a96643b5362cd3ea0016a4d

Observation c9252e02-6cb8-4857-a81d-dabf6f28d5f6 · outbound

This paper cites Oceanic-ship,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Oceanic-ship,

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:53.767902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:53.767902Z digest=sha256:93c5b5ea4f2309e7cec68570d038191df01d0af0cd7fffc708439325e154b441

Observation bd39e153-1de2-4451-b5bf-85dd1427fe7e · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Learning transferable visual models from natural language supervision,

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:53.856194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:53.856194Z digest=sha256:dc572a202f385c57b958ed3f631bec331b95660eedf233972a2da74711c55f3f

Observation 8d669b70-2684-4702-98cd-2f6004dbf9cc · outbound

This paper cites A novel svm-based decoder for remote sensing image captioning,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives A novel svm-based decoder for remote sensing image captioning,

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:53.940975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:53.940975Z digest=sha256:9a64fa25a89843c161ee8dcdc7cb9db9e98623a5676cc6daa0706a8e487bed67

Observation f94161d5-5c71-4a9e-a1f3-0b29eebc6008 · outbound

This paper cites Star: A first- ever dataset and a large-scale benchmark for scene graph generation in large-size satellite imagery,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Star: A first- ever dataset and a large-scale benchmark for scene graph generation in large-size satellite imagery,

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:54.042866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:54.042866Z digest=sha256:7365bb64a53dc7c5584c88832de4f1c6b681e0521b7af5a350f237384b401eb3

Observation 0f6ab84d-fb29-4562-8e98-555a0e36d4b6 · outbound

This paper cites Fine-grained recognition for oriented ship against complex scenes in optical remote sensing images,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Fine-grained recognition for oriented ship against complex scenes in optical remote sensing images,

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:54.193358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:54.193358Z digest=sha256:804b69e34c4cefec89283e95cbc1549c94f7a2b699cbac84a35c8f43c6922fe5

Observation 12e4f184-fb94-4dd7-a226-7c059ac8fdb4 · outbound

This paper cites Hrsid: A high-resolution sar images dataset for ship detection and instance segmentation,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Hrsid: A high-resolution sar images dataset for ship detection and instance segmentation,

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:54.375287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:54.375287Z digest=sha256:194fe8a6f5ceb15aeae5a02ff76d5c65e6d23a4e1eb2a92db0afe14f0d1b32e0

Observation 9767e98b-45c8-4f16-bb92-16656e945f49 · outbound

This paper cites Deepsat: a learning framework for satellite imagery,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Deepsat: a learning framework for satellite imagery,

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:54.541073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:54.541073Z digest=sha256:d0076f3b9adc0b1528a4d28ad2f8b7d8553090f018d0dbd71f60f83040e58b8d

Observation 092404e3-5ee1-4800-9439-a696adec9209 · outbound

This paper cites Nasc-tg2: Natural scene classification with tiangong-2 remotely sensed imagery,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Nasc-tg2: Natural scene classification with tiangong-2 remotely sensed imagery,

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:54.677808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:54.677808Z digest=sha256:c2ce2c55169e0974a9eef388a8053c24b255ea0ebd1f9f0f6594465d35e86c64

Observation 24ddc95e-49e5-4702-afdd-33cfdf9e1ce2 · outbound

This paper cites Feature significance-based multibag- of-visual-words model for remote sensing image scene classification,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Feature significance-based multibag- of-visual-words model for remote sensing image scene classification,

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:54.847291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:54.847291Z digest=sha256:080a3d940abef4dd0fe68cea877e2aa0b17d3463028aa1f2467bb08639e50f99

Observation 6b49a768-8919-4179-afc5-6c7c478e92c9 · outbound

This paper cites Dirichlet-derived multiple topic scene classification model for high spatial resolution remote sensing imagery,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Dirichlet-derived multiple topic scene classification model for high spatial resolution remote sensing imagery,

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:54.995650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:54.995650Z digest=sha256:1f5bd677d98245bd4047205fcd64b66f192e52cf6638b7541241106116a213e3

Observation 25deeedd-14ab-498b-a3d1-7a4f4784a03f · outbound

This paper cites Patternnet: A benchmark dataset for performance evaluation of remote sensing image retrieval,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Patternnet: A benchmark dataset for performance evaluation of remote sensing image retrieval,

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:55.164443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:55.164443Z digest=sha256:2d8dee69088921d3a505dbfb8141cf2693ef874d0c25d4bf97f962bbf1b6482c

Observation 3a7e90a7-a976-4e2f-ba1d-b185f8b532f6 · outbound

This paper cites Accurate object localization in remote sensing images based on convolutional neural networks,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Accurate object localization in remote sensing images based on convolutional neural networks,

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:55.256524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:55.256524Z digest=sha256:c992ffbba3d7b3ed6a39ac8cb3d2caa3f08651644a2355c728bcefe842384970

Observation d7d6f5c8-e6fa-446a-b1cc-fb89c5014327 · outbound

This paper cites Land-cover classification with high-resolution remote sensing images using transferable deep models,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Land-cover classification with high-resolution remote sensing images using transferable deep models,

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:55.326585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:55.326585Z digest=sha256:3de9c325294cb996ee3455fedac9bff26816780e1a042ee27128f6f2892c902f

Observation 101371cc-8551-4285-99eb-cec4523d8062 · outbound

This paper cites Scene classification with recurrent attention of vhr remote sensing images,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Scene classification with recurrent attention of vhr remote sensing images,

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:55.387112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:55.387112Z digest=sha256:0174ff6d0d30c7b2abc24459cb2a1edde5d4c9a134553e1b9e3415f296ca52a4

Observation 3e4482f4-aff9-45e9-b94c-ed82e06fa7af · outbound

This paper cites Clrs: Continual learning benchmark for remote sensing image scene classification,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives Clrs: Continual learning benchmark for remote sensing image scene classification,

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:55.503910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:55.503910Z digest=sha256:65bfc31fed80b5b681c45bffc7707ba001abbc76218d47066252bb6492813483

Observation 29b7051a-932f-48f1-8806-dfd06eb16a7c · outbound

This paper cites On creating benchmark dataset for aerial image interpreta- tion: Reviews, guidances, and million-aid,.

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives On creating benchmark dataset for aerial image interpreta- tion: Reviews, guidances, and million-aid,

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:55.619271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:55.619271Z digest=sha256:f8fe8f695f5e098d085045aa4600f8182853a301f16e98bb170f2401a9237641

Pith citing papers

Observation f489b34b-47d3-4a3c-8fb1-24e0b4a0f13a · inbound

TreeAgent: A Generalizable Multi-Agent Framework for Automated Bias Labeling in Forestry via Compiled Expert Rules and Vision-Language Models cites this paper.

TreeAgent: A Generalizable Multi-Agent Framework for Automated Bias Labeling in Forestry via Compiled Expert Rules and Vision-Language Models Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:45:42.280738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T05:11:58.947649Z digest=sha256:e822c83311bdabac1d04658a7dc3c7eaa299f03d668418747569691f37f0fad0