Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:57:50.582268Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 4 inbound Pith citation observations for arXiv:2506.14821.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:57:50.582268Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T00:59:21.754125Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T15:09:54.916129Z
36 of 36 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 807be10b-9395-4077-8cca-cfe0d2bde021 · outbound
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Qwen Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e80221b9-7f65-44dc-a55c-248078a8b337 · outbound
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Qwen2.5-VL Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0fa55a5-b713-49fc-9f3b-ff77372806c6 · outbound
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Words or Vision: Do Vision-Language Models Have Blind Faith in Text?
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d65f2de-5a24-42a4-a984-dd68cf4ac653 · outbound
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45647678-922f-44e3-a91a-47a44320d024 · outbound
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints How Well Can Vision Language Models See Image Details?
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fea62e8-17af-4438-ae76-049c1c4e08f9 · outbound
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 115bedb6-730c-4e7a-861e-7d0c591c3f09 · outbound
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Visual programming: Compositional visual reasoning without training
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2376941c-1789-41cd-959f-712c3846433b · outbound
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6b6d596-1516-4e9c-826d-5f380ebd5e28 · outbound
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints GPT-4o System Card
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57dc231c-abb1-4408-ab63-d9983aa0474c · outbound
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbc78e49-d107-4875-9227-0be3f22a816d · outbound
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints DyFo: A Training-Free Dynamic Focus Visual Search for Enhancing LMMs in Fine-Grained Visual Understanding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb934223-e1ef-4b8d-8011-0ad370fc0956 · outbound
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints API -bank: A comprehensive benchmark for tool-augmented LLM s
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8d24794-e0cf-4293-9398-a5e18f4e90df · outbound
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints InfiMM-HD: A Leap Forward in High-Resolution Multimodal Understanding
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c58f9d50-3dee-4194-a524-81aad9ff36cf · outbound
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c50d509-3051-4a26-a553-5419dcc3eb98 · outbound
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95f5041c-ab15-4854-9ef6-ab890f92ed51 · outbound
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b419a21e-9801-476d-bbb9-f56277c8cb72 · outbound
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints ToolRL: Reward is All Tool Learning Needs
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25f6e155-d836-4c44-84df-d5ef548a87fc · outbound
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Vision language models are blind
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2b432fc9-d0e0-4a6b-85fe-20d7c1df7828 · outbound
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52602015-4285-4698-93ab-0b88b0fc2e98 · outbound
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee95d8b3-e1b9-411a-aca4-e01c330c29ec · outbound
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fe2aaaa-71ca-4903-873d-9525efa9739c · outbound
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Towards vqa models that can read
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4370f78f-ad7e-460a-afac-bbb72edbc717 · outbound
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd96d3a5-b0e9-43ea-9bb2-b19e134a2252 · outbound
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Eyes Wide Shut ? Exploring the Visual Shortcomings of Multimodal LLMs
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1cf72b14-1ce3-462a-a7c8-279e6a630443 · outbound
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Mllm-tool: A multimodal large language model for tool agent learning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eddcca11-3e5d-4acd-9b7a-77af903b8cd7 · outbound
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 810329ef-15bb-457a-8210-5e9744edee3b · outbound
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Divide, conquer and combine: A training-free framework for high-resolution image perception in multimodal large language models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c994aca-7fb1-4127-81e4-3235be072627 · outbound
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1da3e08-99d8-4710-89cf-989299039783 · outbound
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints V?: Guided visual search as a core mechanism in multimodal llms
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a1c98a9e-53d7-41d8-ba19-673014b47b5b · outbound
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db15b5e0-d7ea-4612-8a4c-0d59a88f5afa · outbound
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7826dbf-e05f-49db-b7d3-180dc53aebe5 · outbound
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Predicting goal-directed human attention using inverse reinforcement learning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b1161c2b-61e7-4430-9521-4776fc732956 · outbound
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b19bc3e2-5c75-401a-aca3-6dbb135c814f · outbound
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90a9cec5-14ed-48b0-9266-55363b3a1af6 · outbound
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36bcf9ab-8f31-4149-b913-3657639cb1dd · outbound
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d14628f5-581b-44e4-8269-0737da2e3b59 · inbound
Visual Reasoning through Tool-supervised Reinforcement Learning Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 286d5a95-1ac0-460a-8654-e38e7b89ef06 · inbound
Robusto-2: Benchmarking Humans & VLMs for Autonomous Driving in Lima & New York City Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation caba1028-2cc3-4038-a103-2f69c434e245 · inbound
From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints
Reference 208
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b16b97a9-0b9a-4018-ab06-12f1346796d1 · inbound
Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.