Pith. sign in

Paper Citation Record · LEDGER

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing

As of 17 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 0 inbound Pith citation observations for arXiv:2607.25993.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.25993 v1

Coverage vector

measured 78 of 78 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T00:59:26.218778Z

measured 78 of 78 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

78 of 78 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved78
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c11bfa9e-2aa5-4ea6-975d-4cb1758b7ac5 · outbound

This paper cites Geollava-8k: Scaling remote-sensing multimodal large language models to 8k resolution.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Geollava-8k: Scaling remote-sensing multimodal large language models to 8k resolution

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:16.872554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:16.872554Z digest=sha256:6032faa8a182f8e491c440dbfc07d76ece76e3650e67298d37b9fc0401c57053

Observation 52b5e044-50a9-4d0c-b289-e25af79c9551 · outbound

This paper cites XLRS-Bench: Could Your Multimodal LLMs Understand Extremely Large Ultra-High-Resolution Remote Sensing Imagery?.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing XLRS-Bench: Could Your Multimodal LLMs Understand Extremely Large Ultra-High-Resolution Remote Sensing Imagery?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:16.958440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:16.958440Z digest=sha256:9182047ee48573397812c0e0f42067d00596bf38c958b87a4ca9858d4faee9ee

Observation 814304eb-51a7-4736-82ce-a7c1f09ffb4a · outbound

This paper cites A benchmark for ultra-high-resolution remote sensing mllms.arXiv preprint arXiv:2512.17319, 2025.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing A benchmark for ultra-high-resolution remote sensing mllms.arXiv preprint arXiv:2512.17319, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:17.014362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:17.014362Z digest=sha256:3de114a4acd015c10e24c627bcef388bf43df64be2d6a1f90fadcbe228853280

Observation 1f60195a-2db3-4a4c-bf89-603f0b4f4429 · outbound

This paper cites Geochat: Grounded large vision-language model for remote sensing.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Geochat: Grounded large vision-language model for remote sensing

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:17.090328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:17.090328Z digest=sha256:9df4021c256253a564d287299fd0b0ffc42384fad243a5ae94a2d9a0364c48f1

Observation 2f6ca9cc-3a94-4c7b-8b69-ad244cf260e6 · outbound

This paper cites Earthgpt: A universal multi-modal large language model for multi-sensor image comprehension in remote sensing domain.IEEE Transactions on Geoscience and Remote Sensing, 2024.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Earthgpt: A universal multi-modal large language model for multi-sensor image comprehension in remote sensing domain.IEEE Transactions on Geoscience and Remote Sensing, 2024

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:17.215092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:17.215092Z digest=sha256:2f91aba17ee5b2cbea472a85edca9a5e90dba5ebfb239188167582823983ac1e

Observation 12285465-c255-4adc-b81b-fb0327377206 · outbound

This paper cites Rsgpt: A remote sensing vision language model and benchmark.ISPRS Journal of Photogrammetry and Remote Sensing, 224:272–286, 2025.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Rsgpt: A remote sensing vision language model and benchmark.ISPRS Journal of Photogrammetry and Remote Sensing, 224:272–286, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:17.318819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:17.318819Z digest=sha256:b327734c013765c88f995ac1a460cb906c33586330de6d6fae976b77bc989ff4

Observation 09b689b6-fee9-451c-8730-1b2f1e95c2aa · outbound

This paper cites Skyeyegpt: Unifying remote sensing vision- language tasks via instruction tuning with large language model.ISPRS Journal of Photogram- metry and Remote Sensing, 221:64–77, 2025.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Skyeyegpt: Unifying remote sensing vision- language tasks via instruction tuning with large language model.ISPRS Journal of Photogram- metry and Remote Sensing, 221:64–77, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:17.439421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:17.439421Z digest=sha256:ceb9edd27636858eb13d3b7f0f821eeadc2402bb8aa0d06339b85924945d20a7

Observation e8b1fcf9-a2d4-49f6-847d-d4841751cd23 · outbound

This paper cites When Large Vision-Language Model Meets Large Remote Sensing Imagery: Coarse-to-Fine Text-Guided Token Pruning.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing When Large Vision-Language Model Meets Large Remote Sensing Imagery: Coarse-to-Fine Text-Guided Token Pruning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:17.520344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:17.520344Z digest=sha256:69ed07b2b7209ce73289787124350c40c9330cb0d0ba144a9312badfcff934d7

Observation ece5afc0-163a-4312-8ea0-0ee64030683d · outbound

This paper cites DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:17.566992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:17.566992Z digest=sha256:ee53f5476fa4b9e7370f3aa219e1631620b5ef09b6ccb726026d830c617bbe65

Observation 18bd7123-58d7-4a9a-a111-c29594e4941d · outbound

This paper cites DeepEyesV2: Toward Agentic Multimodal Model.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing DeepEyesV2: Toward Agentic Multimodal Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:17.736485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:17.736485Z digest=sha256:1152a1a21d90190dfffd3187ef09c1d60dbdeed01500b01368f472fbbbd816ab

Observation 266e68f1-a109-4f4e-ba5c-636564131924 · outbound

This paper cites an unresolved cited work.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:17.889133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:17.889133Z digest=sha256:001032cfe4f5d8f227976e14ba572eaee88b3a85507d6193c3d7d46ed4687792

Observation f8d5eb53-742d-41f6-aa39-8408a6330817 · outbound

This paper cites an unresolved cited work.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:18.019185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:18.019185Z digest=sha256:1435ac9dfc937b15f8c2a0a61f1b4d781f383249030893cbf4d9f218046836fc

Observation 08aba2c1-45c0-46a2-a252-526c7ffe7386 · outbound

This paper cites Zoomearth: Active perception for ultra-high-resolution geospatial vision-language tasks.arXiv preprint arXiv:2511.12267, 2025.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Zoomearth: Active perception for ultra-high-resolution geospatial vision-language tasks.arXiv preprint arXiv:2511.12267, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:18.151653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:18.151653Z digest=sha256:a7550fa010cee5a2ed915a8d28d130f1edce9a76a3c0f967cee188dba78b045a

Observation 2d4f191d-b8e0-4086-bc56-bd4518801e57 · outbound

This paper cites Geoeyes: On-demand visual focusing for evidence-grounded understanding of ultra-high-resolution remote sensing imagery.arXiv preprint arXiv:2602.14201, 2026.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Geoeyes: On-demand visual focusing for evidence-grounded understanding of ultra-high-resolution remote sensing imagery.arXiv preprint arXiv:2602.14201, 2026

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:18.214762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:18.214762Z digest=sha256:bedde894e88e4c3e4f480387caeed5c2e3339cbf5b85d77cc9e5f70dc5bc57d4

Observation 045eb384-c08e-49d9-af74-e8ac2de503a9 · outbound

This paper cites Codev: Code with images for faithful visual reasoning via tool-aware policy optimization.arXiv preprint arXiv:2511.19661, 2025.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Codev: Code with images for faithful visual reasoning via tool-aware policy optimization.arXiv preprint arXiv:2511.19661, 2025

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:18.286594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:18.286594Z digest=sha256:478927836e6680a2648609e95964b28b24136e47b3fcfdafc5d5ff6a22382d04

Observation 7fbffac7-fd71-454a-857e-340f563b60cc · outbound

This paper cites Zooming without zooming: Region-to-image distillation for fine-grained multimodal perception.arXiv preprint arXiv:2602.11858, 2026.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Zooming without zooming: Region-to-image distillation for fine-grained multimodal perception.arXiv preprint arXiv:2602.11858, 2026

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:18.418541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:18.418541Z digest=sha256:47ee571fba27af7d665a7028c2672857937a777a71955db863c108a15c7bcd94

Observation 8d8d84ba-c45b-4320-850e-cf1df276870b · outbound

This paper cites What Does Vision Tool-Use Reinforcement Learning Really Learn? Disentangling Tool-Induced and Intrinsic Effects for Crop-and-Zoom.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing What Does Vision Tool-Use Reinforcement Learning Really Learn? Disentangling Tool-Induced and Intrinsic Effects for Crop-and-Zoom

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:18.518291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:18.518291Z digest=sha256:392ad73153b76f137c14ea2ff21f1f054d9409a74056f8604df114c55d07265c

Observation ab95e434-3b5a-4c3d-9e9b-bc7bfe185316 · outbound

This paper cites Reinforced attention learning.arXiv preprint arXiv:2602.04884, 2026.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Reinforced attention learning.arXiv preprint arXiv:2602.04884, 2026

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:18.640925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:18.640925Z digest=sha256:d61418bd8ddf52041a35e7c7de12715e83801e5efbdcc5c8a7c5dd2dae57d91b

Observation 68b5ea20-e143-40b3-a051-c50550194ebd · outbound

This paper cites Smith, and Ranjay Krishna.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Smith, and Ranjay Krishna

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:18.738278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:18.738278Z digest=sha256:95ead14826617b37deb754d14427ef3d3f20ea2f21a03c259607f9f0fc66248d

Observation df08d272-e693-4579-855c-cdac8a691be9 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:18.845444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:18.845444Z digest=sha256:85c143e1c03d64bbe023f0c585196ee0cc54edd92ef5aeeefc7415c2977a79b5

Observation 13047702-0a6e-46ae-ae43-c7c2fe135170 · outbound

This paper cites an unresolved cited work.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:18.999170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:18.999170Z digest=sha256:f8fd44659eff912a841ce69f1693517e0e0b5ee9cbd7ffb3125a5ee25fca8748

Observation aff2f9cf-eeee-4b84-8efd-093bea6c8f14 · outbound

This paper cites an unresolved cited work.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:19.149637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:19.149637Z digest=sha256:cbe09149aa4bab75693d61d7b8b11b194d8382194e2a0a98e621af9e2978e475

Observation 4f68a163-986b-43dd-a9f1-828720e18bef · outbound

This paper cites Lhrs-bot: Em- powering remote sensing with vgi-enhanced large multimodal language model.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Lhrs-bot: Em- powering remote sensing with vgi-enhanced large multimodal language model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:19.257501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:19.257501Z digest=sha256:49af01181be9a4d168ab345908308f1843147084a5c067a0ac8dccb899ce86f1

Observation 848e696c-d13d-4c81-8ba3-f004776e211c · outbound

This paper cites LHRS-Bot-Nova: Improved Multimodal Large Language Model for Remote Sensing Vision-Language Interpretation.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing LHRS-Bot-Nova: Improved Multimodal Large Language Model for Remote Sensing Vision-Language Interpretation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:19.359533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:19.359533Z digest=sha256:296ac4e0773753cbdcfb13ff7bb69c48b403229fd83184da08e89e1776fc25fd

Observation b29b282b-bab2-41eb-835f-e708125f6286 · outbound

This paper cites Earthmind: Leveraging cross-sensor data for advanced earth observation interpretation with a unified multimodal llm.arXiv preprint arXiv:2506.01667, 2025.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Earthmind: Leveraging cross-sensor data for advanced earth observation interpretation with a unified multimodal llm.arXiv preprint arXiv:2506.01667, 2025

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:19.435240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:19.435240Z digest=sha256:25292a643e7d24dfb2d6e2fca8a283d73c7818b95f6547257d2dc1973275fbad

Observation b5817d38-dbd8-4c31-aca0-310942558b45 · outbound

This paper cites Klein, Salman Khan, and Fahad Khan.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Klein, Salman Khan, and Fahad Khan

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:19.486302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:19.486302Z digest=sha256:88e90986deb08554b19964a38533f7e0a2817b4620efa083f999a3d0324ef34d

Observation a97d61d4-c369-45e6-a60b-b6fa5fc67fb6 · outbound

This paper cites an unresolved cited work.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:19.578996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:19.578996Z digest=sha256:0b3d66263f32b1faae2e2af7be4e9bb351e89d3e9c6f58566b031b8730e49bfd

Observation d23355bd-07ba-4550-bdd9-d3925996e483 · outbound

This paper cites an unresolved cited work.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:19.658554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:19.658554Z digest=sha256:b16ea768395564ced2c7ae99ca511d392d45c50f52280565df4dbee28c67c004

Observation 2563d57c-1345-4621-baed-f8ab87791800 · outbound

This paper cites Earthmarker: A visual prompting multi-modal large language model for remote sensing.IEEE Transactions on Geoscience and Remote Sensing, 2024.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Earthmarker: A visual prompting multi-modal large language model for remote sensing.IEEE Transactions on Geoscience and Remote Sensing, 2024

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:19.789342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:19.789342Z digest=sha256:6965a12f7422794a00cf1fb67347ce9ca13052780046cbc9f5240b6ee668154b

Observation 53f4792a-f73e-43c8-a7ad-4263692d2508 · outbound

This paper cites Rsunivlm: A unified vision-language model for remote sensing via granularity-oriented moe.Pattern Recognition, 179:113717, 2026.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Rsunivlm: A unified vision-language model for remote sensing via granularity-oriented moe.Pattern Recognition, 179:113717, 2026

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:19.925428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:19.925428Z digest=sha256:f3dd8210b2bd8db31e1d9f5b9999b15c750a6cb1582a993a4ad6cf6c629784dd

Observation 59ae6266-280f-4309-814a-f39ee181ff7e · outbound

This paper cites Bermano, and Ohad Fried.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Bermano, and Ohad Fried

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:20.039300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:20.039300Z digest=sha256:bdb9c9ede0953b688dd43510e9a9d4f9c5a900bfe97abf46965a32cf62737fa3

Observation bc98720a-f71e-4bdf-8341-3d066909722d · outbound

This paper cites Chatearthnet: A global- scale image-text dataset empowering vision-language geo-foundation models.Earth System Science Data Discussions, pages 1–24, 2024.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Chatearthnet: A global- scale image-text dataset empowering vision-language geo-foundation models.Earth System Science Data Discussions, pages 1–24, 2024

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:20.113408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:20.113408Z digest=sha256:3a1a3d617e61d097623320a1e9f1745e0e4f464f150c8cceea2246efe85c93c5

Observation 26665e13-7435-416b-8642-190480e9833e · outbound

This paper cites Rsvqa: Visual question answering for remote sensing data.IEEE Transactions on Geoscience and Remote Sensing, 58(12):8555– 8566, 2020.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Rsvqa: Visual question answering for remote sensing data.IEEE Transactions on Geoscience and Remote Sensing, 58(12):8555– 8566, 2020

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:20.219206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:20.219206Z digest=sha256:42cada2df40a27e61a36e45c35334ee6a19bc358683708c7930522fe614568e9

Observation a384ffea-9a3a-4e32-80c3-00cc45c5a7b0 · outbound

This paper cites Mutual attention inception network for remote sensing visual question answering.IEEE Transactions on Geoscience and Remote Sensing, 60:1–14, 2021.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Mutual attention inception network for remote sensing visual question answering.IEEE Transactions on Geoscience and Remote Sensing, 60:1–14, 2021

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:20.396216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:20.396216Z digest=sha256:fcaabdaf0a7f20f4e4cb8e753748d87d867ecbce30c6e1a6d6a9c3706c9de365

Observation 38bcf153-7b20-4316-b4a6-0ae3f5e5e6bd · outbound

This paper cites an unresolved cited work.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:20.604144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:20.604144Z digest=sha256:649180a42a22de7e9575179471e114f94abfe24bd4354728c08dee22cb5fc85e

Observation 6e171e44-4f04-46e4-97f5-1d0ae0d5def9 · outbound

This paper cites VHM: Versatile and Honest Vision Language Model for Remote Sensing Image Analysis.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing VHM: Versatile and Honest Vision Language Model for Remote Sensing Image Analysis

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:20.770802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:20.770802Z digest=sha256:05c8b69ff0eb075f2b0b2a475290424600215b22d32969f919f60706ca111684

Observation efde5a10-6ec6-4f5a-a206-f4331c1d89bd · outbound

This paper cites SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:20.911720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:20.911720Z digest=sha256:981fcfa5861a907b90bf2ac0845c4d03e000d41621dbd3ddcefdf50d47160313

Observation e31ab13a-0a6f-49c5-83a1-e314d39e4076 · outbound

This paper cites Vhm: Versatile and honest vision language model for remote sensing image analysis.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Vhm: Versatile and honest vision language model for remote sensing image analysis

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:21.061018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:21.061018Z digest=sha256:8e388adba14b01230dec1d2e12d54bab67e3982cff12926636cb9597c2302ce8

Observation 3dab8c60-d90d-4613-93e0-11591fb83241 · outbound

This paper cites VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:21.171823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:21.171823Z digest=sha256:98262feba73c04881c20f7e9d82c78ce046034279118daef2519ffdad7e21a55

Observation 8aae2d44-7d1e-4172-ad1c-c475cfd70279 · outbound

This paper cites Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:21.369593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:21.369593Z digest=sha256:255ff7a1b7b3fc84b1c6c49d5c56ffa16375ccf2f90c33029b7586396c5b44f3

Observation 12c40016-1136-4111-a9a1-a24b220f201d · outbound

This paper cites VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:21.527737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:21.527737Z digest=sha256:9e9971972a2af746fac6ef401121252a65dac9af953fa6ec5ebad055bfdfd380

Observation b05c17d7-b308-4078-b392-782864e90a7f · outbound

This paper cites Spacetools: Tool-augmented spatial reasoning via double interactive rl, 2025.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Spacetools: Tool-augmented spatial reasoning via double interactive rl, 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:21.686719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:21.686719Z digest=sha256:215f32bc93d06251213914ec71072d85d1b2ba6a04d5dab9cbe80492c2dd56a7

Observation b16b97a9-0b9a-4018-ab06-12f1346796d1 · outbound

This paper cites Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:21.754125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:21.754125Z digest=sha256:e83196d5a37c7d33469753e5464fd20cbc06391fca85df5941dd60f1d752f73b

Observation 81d1eaf5-cb09-45af-abb6-2528bd1a5cad · outbound

This paper cites CropVLM: Learning to Zoom for Fine-Grained Vision-Language Perception.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing CropVLM: Learning to Zoom for Fine-Grained Vision-Language Perception

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:21.920265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:21.920265Z digest=sha256:c962f2ab8a1b3da6831c1aa31b68ed253ac4b7a4e1e29e5168953725e762f68c

Observation b5c9ff74-6e74-48f0-b20a-0614f1eb827d · outbound

This paper cites Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:22.118346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:22.118346Z digest=sha256:1368df73deb6f7df456bc439fe1ce0dc1225062e09a2ef48a4ce7dd99453501d

Observation 5ea80939-15a8-41db-919c-9e1fd68a0d8b · outbound

This paper cites Sensenova- mars: Empowering multimodal agentic reasoning and search via reinforcement learning.arXiv preprint arXiv:2512.24330, 2025.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Sensenova- mars: Empowering multimodal agentic reasoning and search via reinforcement learning.arXiv preprint arXiv:2512.24330, 2025

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:22.279050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:22.279050Z digest=sha256:e4760987d8d3cbd2ba0caf64fe8ce7fd5d7b3b05654c5c6db29a2d1ac8b2c2c5

Observation f3b598d4-7731-463b-9f71-738c39bc372d · outbound

This paper cites Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:22.526800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:22.526800Z digest=sha256:57a20f36345b28029e5e838b793ab662b9a1e3f4abea03199e471537f580d517

Observation 4da40ed0-c0d3-441c-9daa-7a8a0c18bd8f · outbound

This paper cites Reinforcing spatial reasoning in vision-language models with interwoven thinking and visual drawing, 2025.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Reinforcing spatial reasoning in vision-language models with interwoven thinking and visual drawing, 2025

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:22.676414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:22.676414Z digest=sha256:1bad0963e18511ca2f5ab34153b3e244303b713367d7e487681003d9d187ec10

Observation 9813b036-09b0-4d65-ac6a-ba7a1bd105bd · outbound

This paper cites Qwen2.5-VL Technical Report.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Qwen2.5-VL Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:22.748692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:22.748692Z digest=sha256:18ace0ec6f80f8f12e7dd8ba5354b1061b313fbf1f0b393221059dae517a1762

Observation 110b0c09-67c1-4d65-b886-2b4f74bfd8a8 · outbound

This paper cites Introducing gpt-5.4.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Introducing gpt-5.4

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:22.837480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:22.837480Z digest=sha256:3e40762c2d56087d615ab27c1b031072a57f37d70f18b89c01926c4b20b1bf05

Observation 969d2ab0-ff69-4ca3-ac98-c6f804f4bb31 · outbound

This paper cites The claude 4.6 model family.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing The claude 4.6 model family

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:23.081848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:23.081848Z digest=sha256:a73d17b22a0bd5c3686e79dade99e47483b784609668c6a144cb157ece26d817

Observation 2f458609-188c-4961-b143-3132fb560d04 · outbound

This paper cites Qwen3-VL Technical Report.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Qwen3-VL Technical Report

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:23.284240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:23.284240Z digest=sha256:6ce22d1745b94367325aea2842b78be6224e452d51582b12be1ff9ee6f5ce79d

Observation 8fcdab0b-ec14-4b69-a1d5-7ab571857588 · outbound

This paper cites Gpt-5.2: Advancing science and math.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Gpt-5.2: Advancing science and math

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:23.420815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:23.420815Z digest=sha256:97d6cdf0724441559f8ebc7649e3d0673eca6c68631539d7ee16b27abd353acd

Observation ed036522-578f-41d6-b106-3daf4b5e86d6 · outbound

This paper cites CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:23.594318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:23.594318Z digest=sha256:74eb945dde59ffe42948a7253e30632e8f34b217b4bb0d5d88e31d4795573115

Observation 4623532c-32af-45c6-a840-567602617eb1 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:23.699554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:23.699554Z digest=sha256:d8bf559bd3f5266b892f2b878489c8055373f666001e8f8be0cda02d76224bf6

Observation 173d56fb-755b-424c-afd9-e481bf4dd932 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:23.904221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:23.904221Z digest=sha256:f594d5648c9676efc56e3a2db9cd7f02542774449b8defe0a575dfe555c38d1e

Observation c0ea6155-5476-498b-8c74-6e94153e5854 · outbound

This paper cites Claude 3.7 sonnet.https://www.anthropic.com, 2025.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Claude 3.7 sonnet.https://www.anthropic.com, 2025

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:24.119707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:24.119707Z digest=sha256:6e239172c06e232f666814fb77b749ff4a5c23127980ab3b0bb5c134869bea93

Observation ce446c75-0f10-49c5-b742-d489d7fe1c7f · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:24.259442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:24.259442Z digest=sha256:0b8e0bc22197dac6c1bbe4b2b3d5379f926e8f7ebca630bfd58b763b0f9940ea

Observation b27e1f69-2c10-42bc-8d0f-79df976f9b78 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:24.347294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:24.347294Z digest=sha256:8dfc1e330ad851cd257227915bedc7e1ad768026b983c4666d4047738a97ad99

Observation e4edfa49-1e15-4af3-9132-b1e55c692976 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:24.455888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:24.455888Z digest=sha256:83f14456fd50101d2f92f8d1226633716c7a97f4430579e4444f44f9245f972a

Observation c2acb3f3-628c-4a4d-aec6-4e9aa2b459a4 · outbound

This paper cites Internvl 2.5: Scaling up vision-language models with enhanced visual encoding.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Internvl 2.5: Scaling up vision-language models with enhanced visual encoding

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:24.529399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:24.529399Z digest=sha256:35037505be1d786cc8e44a71cf4beeb30836f48f968f1980f201249b06f58d2b

Observation ba836203-b390-4abe-9c4b-65f75ad38898 · outbound

This paper cites Intern-s1-mini.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Intern-s1-mini

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:24.602875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:24.602875Z digest=sha256:cd5a2201bbf5a490a33bbb36fb037c5aa66cdce3d93b51c7332133aa6ee2e4c4

Observation 04314910-d95d-46a0-ba5a-f4510459b4a7 · outbound

This paper cites Scaling vision pre-training to 4k resolution, 2025.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Scaling vision pre-training to 4k resolution, 2025

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:24.712561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:24.712561Z digest=sha256:0e6e0032397ae0f23c348200cd2d7d7feb5cf3f7b7641bd8871a0e1e4bcc5c22

Observation 89cc5eed-8ea0-411f-b726-447b9a64a00f · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:24.852448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:24.852448Z digest=sha256:4b26a3279529497fa11b183950b098d1a0be4b8010e22040457d1128c550750e

Observation 65d12c25-be27-4cbb-a7bf-f55fddb49ef5 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:24.958431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:24.958431Z digest=sha256:04daa0bca422c3e1ff26540049c4bdfacbc3c6387f47fb6096a0bc7af3836d28

Observation 4cdacec1-aaf1-4bdd-ae93-d183aa36a112 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:25.064532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:25.064532Z digest=sha256:80612b5872a03ba76d4124313507c52d98054c8afa0738a9b750861ea347eb38

Observation 131db01f-fad7-4211-b698-e3a62726865b · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:25.168541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:25.168541Z digest=sha256:d9fc275e2d681489357035adb28cd8f83a1912115074ae9f21be8db15ff29200

Observation fdd8507f-2f26-4c3f-84b7-6da8c8bacf4f · outbound

This paper cites Hello gpt-4o.https://openai.com/index/hello-gpt-4o, 2024.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Hello gpt-4o.https://openai.com/index/hello-gpt-4o, 2024

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:25.243527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:25.243527Z digest=sha256:408a161294168651cd4a6bf24c6c2f65f0aae6775e96c7882e21232dcb82cade

Observation d090fcbd-c8b2-458d-9669-43a09a9333b1 · outbound

This paper cites Gpt-4o mini: advancing cost-efficient intelligence.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Gpt-4o mini: advancing cost-efficient intelligence

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:25.345011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:25.345011Z digest=sha256:bbdd6f44e9eb54558c65dc095c8f14f683cc6c27d4af68463ef7fb8dc5556aa1

Observation 75d0d3e2-ae73-4ecf-b933-363859c22b00 · outbound

This paper cites LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:25.422944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:25.422944Z digest=sha256:5793ceb7365327130476a1dd7179ab859316aa4219a2b3ed674da80ff39fb143

Observation fe6a7bc3-b2ba-48e8-89bd-d1f06d7c25d6 · outbound

This paper cites Qwen3 Technical Report.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Qwen3 Technical Report

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:25.492470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:25.492470Z digest=sha256:f8fff974c62264399d3caf28685146292f69a31e2b8c55a418fca9ef9eeb0367

Observation 3ade99b4-5603-491e-b1ab-50aab15b220c · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:25.571086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:25.571086Z digest=sha256:7a95150d19865415b96cbc7842afaa5618855262e419901a27084e752c54025a

Observation cffc18e7-6fd1-4e01-9dd5-0225277e5e2b · outbound

This paper cites Internlm2 technical report.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Internlm2 technical report

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:25.630835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:25.630835Z digest=sha256:f38b360a964fadef9b8d776eac21986dc38f6ab7347c92b350caff13efdfe2b9

Observation f45abd77-c0e3-48b5-aa2a-f4fe01a12570 · outbound

This paper cites Internlm3.https://github.com/InternLM, 2024.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Internlm3.https://github.com/InternLM, 2024

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:25.734409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:25.734409Z digest=sha256:a4806b71ba2da2999a8ce2dd064a8e012f2a72c051a98ca5e451e6c042088c21

Observation 5d6e9cef-d45f-4363-b862-c74f629bddba · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:25.930022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:25.930022Z digest=sha256:a8df8251a8d457509fe6137f0e057dc32ed0240289e730ef0e977088a8e9534f

Observation b360c749-1614-45a1-8676-f81404cd65e4 · outbound

This paper cites VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:25.996486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:25.996486Z digest=sha256:776e8dda20779ab3d923b5f3f5eb7948d601b2153bab853a5ffbdff347e24770

Observation eaae9c14-519a-480c-a673-644432dd9cd2 · outbound

This paper cites Fair1m: A benchmark dataset for fine-grained object recognition in high-resolution remote sensing imagery.ISPRS Journal of Photogrammetry and Remote Sensing, 184:116–130, 2022.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Fair1m: A benchmark dataset for fine-grained object recognition in high-resolution remote sensing imagery.ISPRS Journal of Photogrammetry and Remote Sensing, 184:116–130, 2022

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:26.104529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:26.104529Z digest=sha256:42fe6c8f7278beac00276e982d595b79c26509600e21f3de263c87b4f6b38810

Observation 2567ef06-e15a-40c5-b4c0-98febc8c1b2c · outbound

This paper cites an unresolved cited work.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Unresolved cited work

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:26.218778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:26.218778Z digest=sha256:ccec27a7351d5c597364168033b3d2bc8e138f0ca7e009d007c7f84eea82dd6c

Pith citing papers

No inbound Pith citation observations are available.