Pith. sign in

Paper Citation Record · LEDGER

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling

As of 7 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 1 inbound Pith citation observation for arXiv:2506.21863.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21863 v1

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:22:41.831349Z

measured 78 of 78 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T13:48:08.135538Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T14:10:29.814089Z

Reference resolution

77 of 77 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved47
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 26b619e7-1fd1-4cff-a1b7-3a42e454f4de · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling LLaMA: Open and Efficient Foundation Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:35.223058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:35.223058Z digest=sha256:9e19e3a7822f82f73c290748519df6e3c8acd496d0fb02a8afb0c138c07aed1d

Observation a367dd25-5608-4ce3-91e5-54c2a34bcb63 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:36.850304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:36.850304Z digest=sha256:b8be067c59a0c6047b34dcd0a6aabf2b7c6066421192dfd80f01dc83cb7546e5

Observation 011fc25c-32c0-4045-9e49-0680ba162714 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:36.905458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:36.905458Z digest=sha256:3d2c68ae6c4c8e383e6eb5491f8b991b3d539f7a9f69dadf2fa5ecff2dbc5eae

Observation b9800eb2-efdb-4547-a677-6792b4e84941 · outbound

This paper cites Stanford alpaca: An instruction-following llama model,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Stanford alpaca: An instruction-following llama model,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:36.997126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:36.997126Z digest=sha256:7cd3727619ef708d5f26c4bf017bfd9a9e98e910fae987cbf69046fc8d55d046

Observation 8ed410b6-ec2c-4dfe-9cb6-af69b0c31ba2 · outbound

This paper cites GPT-4 Technical Report.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling GPT-4 Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.102606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.102606Z digest=sha256:dd0d2920353f71d54b64e36cc2930c1dc0b239f5dad32440ef0d172a17340595

Observation ceffb3a3-2a06-46e5-8057-8f3c954513fc · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Learning transferable visual models from natural language supervision,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.201063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.201063Z digest=sha256:3effad3cfa506543066aeac24340d6ede48eb7c95d6b40e334675446b500cb6f

Observation 0f3c7262-fc54-4fe3-9321-b2dd64f1f0d1 · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:08.137893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:22:37.282011Z digest=sha256:9825af9ec1c1cc2a5633970cee85d8e4a1a2bdacdb291ca1f57bc8e18c6c0d19

Observation d7472608-b0c8-4bb4-97d0-3f532735ae94 · outbound

This paper cites Instructblip: towards general-purpose vision-language mod- els with instruction tuning,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Instructblip: towards general-purpose vision-language mod- els with instruction tuning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:08.025541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:22:37.307798Z digest=sha256:1f1b2825992996f341583d6867b1ff1d06c435cdc8de972c071e2f935ac9a682

Observation be8d938d-911a-477e-be1d-78920bae6ad8 · outbound

This paper cites Visual instruction tuning,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Visual instruction tuning,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:07.904765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:22:37.340244Z digest=sha256:3a5c1f0be56de021bfee01ff4fe72234f1c859ba4856d14d4649f2590190e5c3

Observation 35dd2474-e173-4e13-abf0-5519d5d892b4 · outbound

This paper cites Cogvlm: Visual expert for pretrained language models,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Cogvlm: Visual expert for pretrained language models,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.379971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.379971Z digest=sha256:ca5856137ec2b721acc8b92125aae41b00edbd7873cd6574c11f480110dba1f8

Observation ba83cc8c-8833-4ca0-8872-f73e3468f0f7 · outbound

This paper cites mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.425036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.425036Z digest=sha256:7bec9da333aab33bf9ad4af0b58f63944ce835ec8165c6c6be5817d5dbfa5bef

Observation bbf9b0fa-3c69-4555-ab15-6338cf0f33ae · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Flamingo: a visual language model for few-shot learning,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.458153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.458153Z digest=sha256:4f33d0f97cd3c64683b1f9e4643cb0f4f6f2b044aac32e49b5a97d4d06cd51c5

Observation fa6b6200-8df9-4657-bc92-5e9670131043 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.512451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.512451Z digest=sha256:fd0f693aed050d7d08854352fda30bcdf0fe466d09977d4b3b5fb164fe7553ad

Observation dd754c83-1607-4e06-badf-284d63151ec8 · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:07.740051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:22:37.555854Z digest=sha256:d893466250f66d6053aa333f756dc67e4aa9cae68effab40e82873da20e271a8

Observation f94173a1-c875-42c0-b7a9-a514fe7be0c8 · outbound

This paper cites Qwen2.5-VL Technical Report.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Qwen2.5-VL Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.581702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.581702Z digest=sha256:65a0ef174648b2937e3cded1b4a910b91213de9283a1a28ebe00e0c66ab8deb6

Observation 3b8029ff-bebd-4799-878e-2bc272449bc2 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Multimodal Chain-of-Thought Reasoning in Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.620896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.620896Z digest=sha256:b11433ec19392c599c361dc602f958c9f04c3b227e1b964b84c191d727b1ec2b

Observation 38cd8281-1b22-4352-8eff-46b7aecea09b · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.650879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.650879Z digest=sha256:07cd08df3496444329dbb438e0d75905d5362df8f7ce50f4bc6c396097d88e1e

Observation 6a00b048-1af5-49e0-a168-9ac0d95eebb4 · outbound

This paper cites Visionllm: Large language model is also an open-ended decoder for vision-centric tasks,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Visionllm: Large language model is also an open-ended decoder for vision-centric tasks,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:07.626489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:22:37.685120Z digest=sha256:ee8ca5768038b473ee3d55378a002d570c04f3a10097bf924ec21bfa3ee7029b

Observation 006f5fa3-1fc7-4cbe-9933-4fa979b7adf1 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.709672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.709672Z digest=sha256:721e652f86bb803d8145533588e7645ea2c58ca7c5871416715a10f0a6ee3d88

Observation f6a5387e-36b4-4fbd-b673-82fb7b301473 · outbound

This paper cites Learning source-invariant deep hashing convolutional neural networks for cross-source remote sensing image retrieval,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Learning source-invariant deep hashing convolutional neural networks for cross-source remote sensing image retrieval,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:07.384742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:22:37.735468Z digest=sha256:5f8d2463c26a6c98cd6d9adb51166523b98f0cf530028c39fb399518a24c2d55

Observation 4825f794-4740-45eb-b570-3a66cf21449b · outbound

This paper cites Automatic radiometric normalization for multitemporal remote sensing imagery with iterative slow feature anal- 12 ysis,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Automatic radiometric normalization for multitemporal remote sensing imagery with iterative slow feature anal- 12 ysis,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:07.237268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:22:37.759575Z digest=sha256:821af270175106e853fadfa864d8cb40a792edd8e6c00166b2c6b58aec084661

Observation 8cd2adb0-69d2-4b66-90b4-5bf616d57c2b · outbound

This paper cites A supervised progressive growing generative adversarial network for remote sensing image scene classification,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling A supervised progressive growing generative adversarial network for remote sensing image scene classification,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:07.054740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:22:37.795105Z digest=sha256:a41f72522ea25623d177075997d911db213c1f4c89573006fd96ab3cc324404e

Observation b237e4cc-abc1-4e2d-bcec-3e145eadc888 · outbound

This paper cites Multimodal remote sensing image matching combining learning features and delaunay triangula- tion,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Multimodal remote sensing image matching combining learning features and delaunay triangula- tion,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:06.906216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:22:37.825702Z digest=sha256:b355ebfaa177ec33b885cb57e6cc83a34f2a937ca1122cf264ce66b4980a0e3e

Observation aa33687a-f580-4333-97d4-2a08a59da360 · outbound

This paper cites Urban flood-related remote sensing: research trends, gaps and opportunities,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Urban flood-related remote sensing: research trends, gaps and opportunities,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:06.680198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:22:37.864695Z digest=sha256:688b59d32ddd132b2d49dab24ac836aeeba0e7bbd7b779dee763a7b7847705c4

Observation 107a23c2-1ed6-4912-8df3-21a2d23c0e9e · outbound

This paper cites Remote sensing-based proxies for urban disaster risk management and resilience: A review,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Remote sensing-based proxies for urban disaster risk management and resilience: A review,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:06.535824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:22:37.917813Z digest=sha256:38680c2ff37bea81264ea7cac10b3389e8655a104a528427d4705ab6e4bf0303

Observation 091bfab1-7671-4e9d-a78a-12980f751795 · outbound

This paper cites Remote sensing in multirisk assess- ment: Improving disaster preparedness,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Remote sensing in multirisk assess- ment: Improving disaster preparedness,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:06.468657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:22:37.951695Z digest=sha256:3d7ab28699f0e8705f2be0e7770e9e32019de15054378aea828a48d930b00b2f

Observation 48d1a99b-3fff-4879-86fb-83e40371dc0c · outbound

This paper cites Geochat: Grounded large vision-language model for remote sensing,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Geochat: Grounded large vision-language model for remote sensing,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:06.359877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:22:37.992766Z digest=sha256:89956fb9260e66208971d3d13e2e5768b2814309a3c7a879484c1f2dbc9f177e

Observation bf95f217-69ae-4b4d-8275-716b44b4473f · outbound

This paper cites Lhrs-bot: Empowering remote sensing with vgi-enhanced large multimodal language model,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Lhrs-bot: Empowering remote sensing with vgi-enhanced large multimodal language model,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:06.234735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:22:38.024677Z digest=sha256:4ed2f445c532f4a9e80420255516d73c179749dde7ed36a4acc96ac007fcc130

Observation 8ed055c6-7b5c-4429-ae66-7caf29855f62 · outbound

This paper cites Skyeyegpt: Unifying remote sensing vision-language tasks via instruction tuning with large language model,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Skyeyegpt: Unifying remote sensing vision-language tasks via instruction tuning with large language model,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:06.061858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:22:38.095736Z digest=sha256:6fa5f8f44e6cb32f256562c5ba251677316536efe6c63d82d7f00f521b4b66d5

Observation 1790d3fa-4750-40bd-8f65-f765a54d9134 · outbound

This paper cites H2rsvlm: Towards helpful and honest remote sensing large vision language model,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling H2rsvlm: Towards helpful and honest remote sensing large vision language model,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:05.609068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:22:38.155117Z digest=sha256:2a03bdd05f1d7b6e2330d21de08a8f872a157844552202a54fd22e257c9d0ba0

Observation aec9247a-2ce7-4106-ba4c-d5a89bbc388e · outbound

This paper cites Rs- llava: A large vision-language model for joint captioning and question answering in remote sensing imagery,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Rs- llava: A large vision-language model for joint captioning and question answering in remote sensing imagery,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:38.225898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:38.225898Z digest=sha256:3bf51169ca224a02778ddca9c7c9be3ef02c577e68879d2516011fe6ae10dedf

Observation f16d1259-155d-4449-8054-a9a99b154a5a · outbound

This paper cites SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:38.288515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:38.288515Z digest=sha256:a8b72ec16ce1be43cfeb485c3a1411824e9e874bad82abe6de9e442280fd51a7

Observation 0f5c73c2-eed5-4256-95f8-765f8e89c25f · outbound

This paper cites Rsgpt: A remote sensing vision language model and benchmark,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Rsgpt: A remote sensing vision language model and benchmark,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:38.369262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:38.369262Z digest=sha256:08e3e2def66f6989299ab410cf72d4c77497b082764ee8f1f7469e1686f5bba3

Observation f83e90f2-a829-40bc-b0ea-8b0f4cb61f5f · outbound

This paper cites RSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mixture of Experts.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling RSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mixture of Experts

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:38.461173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:38.461173Z digest=sha256:b0f2af68a8ef3e927aae496e3a107622254d0bc6e6009c4c3d5259fdbf309cff

Observation ec569c52-8643-4cc7-83d8-c0bd36107b6f · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:38.535644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:38.535644Z digest=sha256:2ceb76141902245e1eda5c6b771d3b2a8eaae1c189fa4bc0e597cc2ab38b440a

Observation bb233f38-f6ac-42d0-844a-370aa015a0e8 · outbound

This paper cites The Llama 3 Herd of Models.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling The Llama 3 Herd of Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:38.617226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:38.617226Z digest=sha256:a8213e9b92449fb672192a115bba0b9b52a13a562b658c8387103e61268ee718

Observation 77549bf0-b983-43f4-84e4-b843211d9a8e · outbound

This paper cites Qwen Technical Report.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Qwen Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:38.685861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:38.685861Z digest=sha256:2064e52efeb3b465e00c02498a1e1e03f5d898a1308049fd344249454ccb80c2

Observation b20196fd-9769-4162-830a-65bd6f5a89c8 · outbound

This paper cites Internlm: A multilingual language model with progressively enhanced capabilities,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Internlm: A multilingual language model with progressively enhanced capabilities,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:38.777895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:38.777895Z digest=sha256:765bf07d1d4c0fa4b23a6881afcac8987f5621fb769e9d1cbfe507667ed3841d

Observation 253ce06b-4d3e-4d1b-bb05-902024b5477d · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Llava-next: Improved reasoning, ocr, and world knowledge,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:38.870626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:38.870626Z digest=sha256:7244c296c274c6af503f31996dc5b1f660fdaf775ad7b0831fa58768676b97e0

Observation d9dcd1a2-1e1c-42ab-934b-5e83343b41f0 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:38.978885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:38.978885Z digest=sha256:a81a33ceafe767d5bae6bd58c52aebf6c91195a88126ffde543c20b0a07c5208

Observation e4d4b9f3-1b07-4cb6-b2d9-5521e39fca66 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:39.049798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:39.049798Z digest=sha256:df93eb926a42d0c81eb1cca69b4b642b82e53842caf69d6d8bffc8346601a421

Observation 2a7d4b5d-d066-4e70-bbd3-4208638aae9f · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:39.123538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:39.123538Z digest=sha256:1a5b156825b62ae6856f08e3a880cb7484ac6ecafb08e30691996639bec8103f

Observation a62c908d-af4f-4995-9b14-dc0f73021938 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Gemini: A Family of Highly Capable Multimodal Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:39.212756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:39.212756Z digest=sha256:92ffb32fadc8c080276511f67ba773a2d5042ac811a28a0451084c105c44beee

Observation e759238f-dab3-4c82-8923-2bcb3f9e021a · outbound

This paper cites KOSMOS-2.5: A Multimodal Literate Model.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling KOSMOS-2.5: A Multimodal Literate Model

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:39.299853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:39.299853Z digest=sha256:d1d2bd02836bf69c08d01a2651eab429bc47aa8d7b1ea063a4e339ccd2c6b418

Observation ff274aec-2bf1-4ed8-8007-f2c30dc020d1 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Scaling up visual and vision-language representation learning with noisy text supervision,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:39.393855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:39.393855Z digest=sha256:693115eecc4959f83844a8f5315d9428c0a607d7032b32c954bea1c5998e2765

Observation ef79ffac-a7ca-4ea5-9e95-168cee76a890 · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Sharegpt4v: Improving large multi-modal models with better captions,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:03.827568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:22:39.491263Z digest=sha256:d60336f0b207fef4ee08a2cde351f58435f6a25090f86a9757a72b3f4185140d

Observation e83556a7-25f4-4a34-845c-d81d4b8bfba9 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:39.583072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:39.583072Z digest=sha256:2f508116d81542538c5c014bd33f1dc61b6e3347d765112588b8fff7b21ddabe

Observation 6634d2c4-535b-4d8f-889a-8032b1f00373 · outbound

This paper cites InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:39.652994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:39.652994Z digest=sha256:263ff36671b828bff350c8419a3ad81811cb4f0d01fcbfca7a3c0eb6da9c2216

Observation b281da44-fbdf-4dc3-9560-9aa1687ef19c · outbound

This paper cites Gpt4roi: Instruction tuning large language model on regionof-interest, 2024,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Gpt4roi: Instruction tuning large language model on regionof-interest, 2024,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:03.548855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:22:39.761352Z digest=sha256:65ba960c2a8c9690e5381d457d9d48f9c2f3c12d870c86d6ffed4e382cea1c7c

Observation 0409c39d-37a3-487d-9e11-a04f0cd333e9 · outbound

This paper cites Glamm: Pixel grounding large multimodal model,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Glamm: Pixel grounding large multimodal model,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:39.844011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:39.844011Z digest=sha256:f8dc4be1834250af6a62a319e2e09127df9a6df51f903f001fcd06760ea5ef38

Observation d4964b9e-ef44-4714-bbbb-b500462ca1fb · outbound

This paper cites LHRS-Bot-Nova: Improved Multimodal Large Language Model for Remote Sensing Vision-Language Interpretation.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling LHRS-Bot-Nova: Improved Multimodal Large Language Model for Remote Sensing Vision-Language Interpretation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:39.933256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:39.933256Z digest=sha256:2ce22ec7497a2542d87d5bcfa6e1bf8a660e47817604c169ec9f53f1cc8ee12b

Observation e5b11529-a876-47bc-afe7-58fcfbb95c0d · outbound

This paper cites EarthMarker: A Visual Prompting Multi-modal Large Language Model for Remote Sensing.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling EarthMarker: A Visual Prompting Multi-modal Large Language Model for Remote Sensing

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:40.050681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:40.050681Z digest=sha256:149a0639a1b87999d0259f26187c5a87b47d9831a3d23c3b008856992af5b549

Observation cddd690e-fa14-49b9-a2bc-321b147da3a4 · outbound

This paper cites GeoLLaVA: Efficient Fine-Tuned Vision-Language Models for Temporal Change Detection in Remote Sensing.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling GeoLLaVA: Efficient Fine-Tuned Vision-Language Models for Temporal Change Detection in Remote Sensing

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:40.114061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:40.114061Z digest=sha256:bca84105674533c402720bde4a0d655bd5c26486b509ef38ca02f2d03b471aea

Observation b837ee10-fd68-4969-a7c7-5ae3e609ebc5 · outbound

This paper cites TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:40.225645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:40.225645Z digest=sha256:f4517bae5a2c2c459badfd773cc06b9b920b45428ccc28ec064799fa36777a59

Observation ee836c34-910a-4761-b513-d54b9060d421 · outbound

This paper cites Earthgpt: A universal multi-modal large language model for multi-sensor image comprehension in remote sensing domain,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Earthgpt: A universal multi-modal large language model for multi-sensor image comprehension in remote sensing domain,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:03.371380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:22:40.319247Z digest=sha256:b7b28d04dc540c1ff674b9fd4f737c4bfc629617f033afe4106a1c28f89be0b3

Observation 76d8790e-b373-4c74-a4c0-9e8653b0605b · outbound

This paper cites Ringmogpt: A unified remote sensing foundation model for vision, language, and grounded tasks,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Ringmogpt: A unified remote sensing foundation model for vision, language, and grounded tasks,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:03.204907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:22:40.405051Z digest=sha256:186ac030217fb117c0ab114245105946abffdf3c3b19a7b0993a30179bc45146

Observation 1e2553fc-aa01-4bcf-9126-187abe71100c · outbound

This paper cites Skyscript: A large and semantically diverse vision-language dataset for remote sens- ing,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Skyscript: A large and semantically diverse vision-language dataset for remote sens- ing,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:40.494200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:40.494200Z digest=sha256:2edc55594a64ce44ba28ae655da1520b5e715953134a9fabc562b2d007948826

Observation f30ee937-9a04-41b4-b979-557da59ec5a0 · outbound

This paper cites Deep semantic understanding of high resolution remote sensing image,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Deep semantic understanding of high resolution remote sensing image,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:45.081454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:22:40.564185Z digest=sha256:0032bbb891e903b8a96cbb5560b5a93c8f6e324e61615ab9ec65a11e9bb226d4

Observation 9e93813d-4eaa-4ec6-8557-c15fbd3fc5cf · outbound

This paper cites Remote sensing image scene classifi- cation: Benchmark and state of the art,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Remote sensing image scene classifi- cation: Benchmark and state of the art,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:40.639987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:40.639987Z digest=sha256:ac3acb3c3a4bb3f6cc22cf13e2fb848c354d4fecb5bba28995805a282e0438e0

Observation 1b3dcfdb-32b9-4c65-a9d6-911c92551ff7 · outbound

This paper cites Nwpu- captions dataset and mlca-net for remote sensing image captioning,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Nwpu- captions dataset and mlca-net for remote sensing image captioning,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:44.694995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:22:40.714308Z digest=sha256:e07afd24bf933df407e648fb13bc712dea5866f40686958b43d93010f48c007c

Observation 4911cca2-b81f-45e1-9d61-c6f362d666ac · outbound

This paper cites Exploring a Fine-Grained Multiscale Method for Cross-Modal Remote Sensing Image Retrieval.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Exploring a Fine-Grained Multiscale Method for Cross-Modal Remote Sensing Image Retrieval

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:40.798399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:40.798399Z digest=sha256:8a404b82f2e11680f183c8ffb0649b9d1402e55d6a4cc80bda02a260a316d71e

Observation 51d48c8a-48b8-4d32-99bb-97de27ae9e6e · outbound

This paper cites METER-ML: A Multi-Sensor Earth Observation Benchmark for Automated Methane Source Mapping.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling METER-ML: A Multi-Sensor Earth Observation Benchmark for Automated Methane Source Mapping

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:40.908652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:40.908652Z digest=sha256:b9059378e7ca03f4881117c63230483ec6be200b8e0a511dfb3145dfaaf2be44

Observation bca7d714-1582-4d43-bce5-a7e9262d9514 · outbound

This paper cites Functional map of the world,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Functional map of the world,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:41.031030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:41.031030Z digest=sha256:a195131a322563488e9c0287b5565cae8393023c70944a76d2097b7d168e94f7

Observation a2dcb4e5-4bc9-4eb3-a623-418848a1c8e2 · outbound

This paper cites Exploring models and data for remote sensing image caption generation,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Exploring models and data for remote sensing image caption generation,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:44.460150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:22:41.078684Z digest=sha256:1dff682d517f217a408ae9e0f23abdac39f39250d68cc904fb78fd081da56ac2

Observation ff9def5d-e6e7-4689-9bc6-520ecde441ec · outbound

This paper cites Rsvqa: Visual question answering for remote sensing data,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Rsvqa: Visual question answering for remote sensing data,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:44.146974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:22:41.132016Z digest=sha256:eb154c30e2091440daea84293169fce606bea85c0c098a3481c90fb7ccc0582d

Observation e0692b84-d5bb-411a-8ae5-72431ce9f3c4 · outbound

This paper cites Visual grounding in remote sensing images,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Visual grounding in remote sensing images,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:43.880645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:22:41.183387Z digest=sha256:239cfaa0e30d456da50b32e360b1befea74b591f38b6c1f93d6bf343db6ce083

Observation ad0bd306-6f06-44b9-a4c6-77ded1d15432 · outbound

This paper cites Rsvg: Exploring data and models for visual grounding on remote sensing data,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Rsvg: Exploring data and models for visual grounding on remote sensing data,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:43.585787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:22:41.237485Z digest=sha256:853bcefe549dc1c71470ef6021657fd5fc7c80f13e6501f3f3cee5969f90c1dc

Observation 0ace819e-8c81-45a1-8691-900ac20deb2a · outbound

This paper cites Object detection in optical remote sensing images: A survey and a new benchmark,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Object detection in optical remote sensing images: A survey and a new benchmark,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:41.294447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:41.294447Z digest=sha256:aeef4099880f85ff7c39c32a9686e88091754c2b8800827f0a208440f79fefbd

Observation 74381c94-fca9-478f-8ce9-fefc3ea57b87 · outbound

This paper cites Dota: A large-scale dataset for object detection in aerial images,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Dota: A large-scale dataset for object detection in aerial images,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:41.350812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:41.350812Z digest=sha256:22dc8a1ce8b0e4c625bd088db872df8878c8198a9632f14f953e87c6a86f6ef4

Observation d5a745c7-c55c-4c8a-8348-2d69ffe7ae2b · outbound

This paper cites Aid: A benchmark data set for performance evaluation of aerial scene classification,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Aid: A benchmark data set for performance evaluation of aerial scene classification,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:43.248276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:22:41.398190Z digest=sha256:a34a6f4e5fbc98e845a07f9ad1e09ced754b247b278a187185052b0ecaebdb5b

Observation f0ac9d96-da54-4527-9b83-f21b3531fddd · outbound

This paper cites Satellite image classification via two-layer sparse coding with biased image representation,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Satellite image classification via two-layer sparse coding with biased image representation,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:43.002910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:22:41.467391Z digest=sha256:cda529f21f355a022470110aa71f011bca8ab61bfb537e6864d8323a2b339980

Observation 16b87da5-b324-4c2a-aa0a-a7711128364f · outbound

This paper cites Bag-of-visual- words scene classifier with local and global features for high spatial resolution remote sensing imagery,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Bag-of-visual- words scene classifier with local and global features for high spatial resolution remote sensing imagery,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:42.750777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:22:41.529916Z digest=sha256:524ad3199ceb64fec24623ee4000cc2ecb1d1a7deba684b451e1d97f7555ed4f

Observation 7c171966-67c3-453a-8619-10b7c97a4124 · outbound

This paper cites Improved baselines with visual instruction tuning,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Improved baselines with visual instruction tuning,

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:41.583771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:41.583771Z digest=sha256:2f2fd911544d2d96951711ceb817f2500f992e52a608ed0e3836513bf9811467

Observation ab0645a4-d411-4393-a611-ed09529fa31e · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:41.662531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:41.662531Z digest=sha256:4656574b4a2a90ca5397a5af63a512cd3be8c1b7c382900bf9cf9e4bce08527e

Observation b3dcb610-4927-4266-88be-ce3d69c6f970 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:41.716132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:41.716132Z digest=sha256:1a205c5c4da9966ce9e024956f86ae29ead1cc4939dbfd221ef3d73537c6512b

Observation f5c432bd-ec86-4cd4-8ee2-4da9858061f4 · outbound

This paper cites Qwen2.5: A party of foundation models,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Qwen2.5: A party of foundation models,

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:41.767782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:41.767782Z digest=sha256:9e565fb5e7bfae9c4eed6d0b058e7b285be0b96f4bfe0a783c2e11f6d4446735

Observation 59bd219b-acaf-4158-9483-a56f0ca825b3 · outbound

This paper cites Vhm: Versatile and honest vision language model for remote sensing image analysis,.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Vhm: Versatile and honest vision language model for remote sensing image analysis,

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:42.582486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:22:41.831349Z digest=sha256:d102c421bb18c87ac8731ab04e009ad0e6def07d72584a34172a2e5b035ee15c

Pith citing papers

Observation 522cebc6-71b4-418f-9961-afaf5b036e74 · inbound

Vision-and-Language Navigation for UAVs: Progress, Challenges, and a Research Roadmap cites this paper.

Vision-and-Language Navigation for UAVs: Progress, Challenges, and a Research Roadmap Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling

Reference 148

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:10:29.816125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:48:08.135538Z digest=sha256:fe2ee60d45167931e2e57f2e8c287332c341ec35a3245c5034a0dc4c1dbd2db6