Pith. sign in

Paper Citation Record · LEDGER

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints

As of 23 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 4 inbound Pith citation observations for arXiv:2506.14821.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.14821 v3

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:57:50.582268Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T00:59:21.754125Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:54.916129Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 807be10b-9395-4077-8cca-cfe0d2bde021 · outbound

This paper cites Qwen Technical Report.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Qwen Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:50.340731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:57:50.340731Z digest=sha256:a895399d4e00edebb0c2daf883638fc36e543b4aca0a6bd516e9f8952aacd277

Observation e80221b9-7f65-44dc-a55c-248078a8b337 · outbound

This paper cites Qwen2.5-VL Technical Report.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:50.347226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:57:50.347226Z digest=sha256:01ee8a790985113b44dd86d46303f1361efa063a89d917480cc4dc9814f927e0

Observation b0fa55a5-b713-49fc-9f3b-ff77372806c6 · outbound

This paper cites Words or Vision: Do Vision-Language Models Have Blind Faith in Text?.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Words or Vision: Do Vision-Language Models Have Blind Faith in Text?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:50.355178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:57:50.355178Z digest=sha256:a54bdc92f33aa22374bae0592c53c80d5bf8fe0b82fa38b0f4c32d99aae07cde

Observation 7d65f2de-5a24-42a4-a984-dd68cf4ac653 · outbound

This paper cites OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:50.363366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:57:50.363366Z digest=sha256:924aeeac6773f6007fec8a82e6cd10dcf6ee89749dd44c5ed94da7fbe16c2d85

Observation 45647678-922f-44e3-a91a-47a44320d024 · outbound

This paper cites How Well Can Vision Language Models See Image Details?.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints How Well Can Vision Language Models See Image Details?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:50.370077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:57:50.370077Z digest=sha256:f5c91a4ed9c8171f889041796fbfcd4c9b021d1f171906b915fbefe9b9cad06d

Observation 6fea62e8-17af-4438-ae76-049c1c4e08f9 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:50.376121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:57:50.376121Z digest=sha256:661eafbf348c18c4a51d2d8b624960a6c3463deef221ce6c75c2fbfa66dfed19

Observation 115bedb6-730c-4e7a-861e-7d0c591c3f09 · outbound

This paper cites Visual programming: Compositional visual reasoning without training.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Visual programming: Compositional visual reasoning without training

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:57:51.567458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T04:57:50.386824Z digest=sha256:8a1d179fc907ab0f7862c98de5df48928617c26ef1d51ea68aeebc86facb2d71

Observation 2376941c-1789-41cd-959f-712c3846433b · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:50.395418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:57:50.395418Z digest=sha256:2c72b157fe0f5a12a1b235812dab4cf4164de0e06af9568e1716b85720fd7bd7

Observation a6b6d596-1516-4e9c-826d-5f380ebd5e28 · outbound

This paper cites GPT-4o System Card.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints GPT-4o System Card

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:50.403978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:57:50.403978Z digest=sha256:2f4ac4d13fc2c864e852eb6fa7fa31a07cfb68e068a987ee91cd11323eeb62b5

Observation 57dc231c-abb1-4408-ab63-d9983aa0474c · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:50.409615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:57:50.409615Z digest=sha256:3cda5a42f07b1ef47898d297c0c8b3b5b35f32275e6c46230c77c7b5758af52b

Observation fbc78e49-d107-4875-9227-0be3f22a816d · outbound

This paper cites DyFo: A Training-Free Dynamic Focus Visual Search for Enhancing LMMs in Fine-Grained Visual Understanding.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints DyFo: A Training-Free Dynamic Focus Visual Search for Enhancing LMMs in Fine-Grained Visual Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:50.419860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:57:50.419860Z digest=sha256:f21c7f81590d55519fa39aa42c6a7685e48574c029f0712de06c7c35dbe5b095

Observation bb934223-e1ef-4b8d-8011-0ad370fc0956 · outbound

This paper cites API -bank: A comprehensive benchmark for tool-augmented LLM s.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints API -bank: A comprehensive benchmark for tool-augmented LLM s

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:50.426006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:57:50.426006Z digest=sha256:42d099191765e44dbde1509c778729611c255b5ce77c9fcba2efb6f6b10fb48f

Observation f8d24794-e0cf-4293-9398-a5e18f4e90df · outbound

This paper cites InfiMM-HD: A Leap Forward in High-Resolution Multimodal Understanding.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints InfiMM-HD: A Leap Forward in High-Resolution Multimodal Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:50.431220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:57:50.431220Z digest=sha256:4cf00ba72021ce0eb6026b27077e5310e4a54c7ca7445d9dc5719448ea965910

Observation c58f9d50-3dee-4194-a524-81aad9ff36cf · outbound

This paper cites TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:50.436455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:57:50.436455Z digest=sha256:1b809b6ddebf0cf916441b5107c5704f7d4f7a06ef8408e9763a385ef7730a29

Observation 5c50d509-3051-4a26-a553-5419dcc3eb98 · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:50.441648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:57:50.441648Z digest=sha256:c312939b2e7c1955d1df355d669375d07e4e4d1f6434015bfa97e632fb662a48

Observation 95f5041c-ab15-4854-9ef6-ab890f92ed51 · outbound

This paper cites an unresolved cited work.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Unresolved cited work

Reference 16

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T04:57:51.083101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T04:57:50.446995Z digest=sha256:dc82a1718fd47c6c875f2243d88d9fb9f0150dcf27fdfe9abec5dd2d552b0de5

Observation b419a21e-9801-476d-bbb9-f56277c8cb72 · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints ToolRL: Reward is All Tool Learning Needs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:50.452790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:57:50.452790Z digest=sha256:c9911e6dd418adc7a17a0197738f69c7d8d7e47542b0e618e2887c1df957a5d1

Observation 25f6e155-d836-4c44-84df-d5ef548a87fc · outbound

This paper cites Vision language models are blind.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Vision language models are blind

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:57:51.543102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T04:57:50.459936Z digest=sha256:bbbcbfdde59e06a6bce1730c0d364c44cdc40dd50673dc3c41669830187147c0

Observation 2b432fc9-d0e0-4a6b-85fe-20d7c1df7828 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:50.465503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:57:50.465503Z digest=sha256:5218fea2cb0838b54e8986d0057292960a806de4f456a0b4d57dd2eed1ffbfd6

Observation 52602015-4285-4698-93ab-0b88b0fc2e98 · outbound

This paper cites ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:50.473467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:57:50.473467Z digest=sha256:a98499d74d4ea5ca5ebb81fade1cfe0e2cfa872e075c0539914515d5c9399e2e

Observation ee95d8b3-e1b9-411a-aca4-e01c330c29ec · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:50.480102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:57:50.480102Z digest=sha256:888d77c0105c515a7f48ab72aa8e6d6df292b5f8cd38431cb74b4b0e6936f739

Observation 0fe2aaaa-71ca-4903-873d-9525efa9739c · outbound

This paper cites Towards vqa models that can read.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Towards vqa models that can read

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:57:51.521447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T04:57:50.487303Z digest=sha256:f31ddf759bd6753e214dc9f940d31c02a87d58e8708d9b44d6aa2587ed48bafe

Observation 4370f78f-ad7e-460a-afac-bbb72edbc717 · outbound

This paper cites Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:50.493602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:57:50.493602Z digest=sha256:c47468665a76c0ba19365bab5d2f39475248b190b648919c5d389b92bd6ef4e1

Observation bd96d3a5-b0e9-43ea-9bb2-b19e134a2252 · outbound

This paper cites Eyes Wide Shut ? Exploring the Visual Shortcomings of Multimodal LLMs.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Eyes Wide Shut ? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:57:51.495365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T04:57:50.500717Z digest=sha256:bff292b573231000437abe8b44860fa0225ede1f9df82fd210a17450972eb705

Observation 1cf72b14-1ce3-462a-a7c8-279e6a630443 · outbound

This paper cites Mllm-tool: A multimodal large language model for tool agent learning.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Mllm-tool: A multimodal large language model for tool agent learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:50.507372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:57:50.507372Z digest=sha256:53f5ebfe9cce879fc33096d742350996f1570e94ca1483409d26f5cb62d1c89c

Observation eddcca11-3e5d-4acd-9b7a-77af903b8cd7 · outbound

This paper cites VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:50.514130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:57:50.514130Z digest=sha256:f9cf26a05827a824c1e1cca05f40bf388ee9abd5194ecc77c3f16c925d141b3e

Observation 810329ef-15bb-457a-8210-5e9744edee3b · outbound

This paper cites Divide, conquer and combine: A training-free framework for high-resolution image perception in multimodal large language models.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Divide, conquer and combine: A training-free framework for high-resolution image perception in multimodal large language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:50.525845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:57:50.525845Z digest=sha256:3b5bd1f9d62103a7c86bbb7dd055d8d2b052318cfb0a68242e842c5027eb9dcd

Observation 2c994aca-7fb1-4127-81e4-3235be072627 · outbound

This paper cites Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:50.533988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:57:50.533988Z digest=sha256:3f6b8a2b469c4ae09505211abfd65d8d6174df8c3e4e162f3a9bd93e1e6f9eda

Observation c1da3e08-99d8-4710-89cf-989299039783 · outbound

This paper cites V?: Guided visual search as a core mechanism in multimodal llms.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints V?: Guided visual search as a core mechanism in multimodal llms

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:57:51.469386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T04:57:50.540959Z digest=sha256:e5cab4e9f5f6252f993bb6984b9c926499b5c36a1689998164e1638885805179

Observation a1c98a9e-53d7-41d8-ba19-673014b47b5b · outbound

This paper cites R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:50.546246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:57:50.546246Z digest=sha256:9dab55babfc36c478d5b3d87f521ebaf5afe8472905372234565104442f496b0

Observation db15b5e0-d7ea-4612-8a4c-0d59a88f5afa · outbound

This paper cites MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:50.551997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:57:50.551997Z digest=sha256:c5ebb0fffa93d7bb03eb9d1433917fb10bd4ea4a12d7742b30f2f24f53606a3a

Observation e7826dbf-e05f-49db-b7d3-180dc53aebe5 · outbound

This paper cites Predicting goal-directed human attention using inverse reinforcement learning.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Predicting goal-directed human attention using inverse reinforcement learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:57:51.448207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T04:57:50.557650Z digest=sha256:9eb01b7d8fe145a24753dd79eb2618c1c63365281a0f8f886bfa9c9f7e89e07f

Observation b1161c2b-61e7-4430-9521-4776fc732956 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:50.563657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:57:50.563657Z digest=sha256:8c9b60c61f88704d61aba3df7cb633295239da2b8c6748130b91bec0ca91d689

Observation b19bc3e2-5c75-401a-aca3-6dbb135c814f · outbound

This paper cites Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:50.569467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:57:50.569467Z digest=sha256:9c98d716595bb5dd0d58f2bea782af8ab0dd030196ad3cdcec9ddb61be0c89af

Observation 90a9cec5-14ed-48b0-9266-55363b3a1af6 · outbound

This paper cites LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:50.575266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:57:50.575266Z digest=sha256:070bdd5a3f998dbf719a125344fc1bd73ed58a756c2cfe51d702b2b127b643d0

Observation 36bcf9ab-8f31-4149-b913-3657639cb1dd · outbound

This paper cites DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:50.582268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:57:50.582268Z digest=sha256:037e57e8d521afaafc60d3017a6621df27bd13135ae765ca6aa9115f6e9a7bfb

Pith citing papers

Observation d14628f5-581b-44e4-8269-0737da2e3b59 · inbound

Visual Reasoning through Tool-supervised Reinforcement Learning cites this paper.

Visual Reasoning through Tool-supervised Reinforcement Learning Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:03.011383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T03:05:21.688216Z digest=sha256:5a12a76feab11a8953cabaa4a16841a8a90e9ff45cf3880bf46539a59d20c9a0

Observation 286d5a95-1ac0-460a-8654-e38e7b89ef06 · inbound

Robusto-2: Benchmarking Humans & VLMs for Autonomous Driving in Lima & New York City cites this paper.

Robusto-2: Benchmarking Humans & VLMs for Autonomous Driving in Lima & New York City Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:59:32.460615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T17:30:07.926246Z digest=sha256:68530a48725a8ae38a29a8202fb976ae2ebaf6b0876e8f2a886c24b4ac485e0d

Observation caba1028-2cc3-4038-a103-2f69c434e245 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints

Reference 208

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:54.917925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:83683bbdf078d0c5b5a4828a89ee0ff613aea3ef5273412a43bfcf680b4de9c9

Observation b16b97a9-0b9a-4018-ab06-12f1346796d1 · inbound

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing cites this paper.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:21.754125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:21.754125Z digest=sha256:e83196d5a37c7d33469753e5464fd20cbc06391fca85df5941dd60f1d752f73b