Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-19T06:10:57.219445Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 6 inbound Pith citation observations for arXiv:2507.05920.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-19T06:10:57.219445Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T20:19:14.175809Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-04T16:29:57.242422Z
47 of 47 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8316a74b-7040-4fcf-9dd5-fb93fb15ac57 · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Qwen2.5-VL Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a4bad2e2-5f09-4ddd-bdfd-72206fb5ec49 · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c3340de6-38de-4680-a081-5f57c9a14964 · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1f07e7e0-eee0-4671-8179-07877bbbceeb · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cc991096-7495-4f71-8763-49a95021207b · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 26111992-6bc2-4e55-8c56-83fdbe0f1905 · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Patch n’pack: Navit, a vision transformer for any aspect ratio and resolution
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7f2e98f7-69b6-4ef1-a6ce-cad52c7a2f7e · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 932f6aae-56d4-4782-a543-4a0fdd953aae · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5a506313-273e-48bf-ad42-05120d387ce9 · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d7eafe4e-337a-4b45-9220-d046863627a6 · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Llava-uhd: an lmm perceiving any aspect ratio and high- resolution images
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f33e2de3-2122-4ca7-9519-4496730a929c · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 410158e2-06a0-4f04-a543-e319dee066ca · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning GPT-4o System Card
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f81de1b1-c5fb-4e6a-8362-81e7e1b1bbd0 · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning OpenAI o1 System Card
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation de95bd19-7d34-430d-881c-c35c877467a1 · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning The hungarian method for the assignment problem
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 796e8c22-5586-499a-9325-692b7ef248b0 · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Efficient memory management for large language model serving with pagedattention
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 44f74fa9-d835-4992-8914-020ba9faaf52 · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning LLaVA-OneVision: Easy Visual Task Transfer
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 80b8bca0-a559-4b5b-82a1-f870a84d7c70 · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 31c2b4e6-2685-4110-bcb6-df4e1146e7c6 · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 12d30eed-5ba8-4cab-aead-1be6a388f186 · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Improved baselines with visual instruction tuning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation eb6a274d-38e4-4fd2-a979-ceed535e3c69 · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Lost in the Middle: How Language Models Use Long Contexts
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f5616df8-bba1-4eb2-a19e-34372d632036 · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cd23a595-bf5b-4d0f-8550-15b82f3f5b94 · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 891797d4-fed9-483a-87f6-fb8f24729b7f · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Ola: Pushing the Frontiers of Omni-Modal Language Model
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b10b30ed-54ee-4cb2-a7da-5d398020fa5f · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Decoupled Weight Decay Regularization
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 89bfe261-63ba-4bc5-ac2f-b5aed3907ad6 · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e339be48-cd2c-4a14-af0e-e90af4fbda39 · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Openai o3 and o4-mini system card
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9bfa58a6-b28e-4251-a279-0ff11e7a0dea · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c115eee6-bf3d-4e05-b8f2-01510781adba · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Learning to count everything
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2c9a6eb4-748c-41c1-848f-c1349661304c · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c7d4c3a1-5c5c-4b2c-8f3e-4999b9a823b3 · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Visual cot: Advancing multi-modal language models with a comprehen- sive dataset and benchmark for chain-of-thought reasoning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fd731518-f764-4cfc-a7c4-1b8f30ff4254 · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 923604e5-c07d-4a78-a39f-dffbfaaf30f4 · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7ddb696f-2b3d-41c6-ab51-48cf9469cb21 · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Scaling Vision Pre-Training to 4K Resolution
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 241fea74-b9e2-4e6f-92e7-1a40a79df36c · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1be0386f-ef27-4a47-a10d-675bc08961c9 · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Visual Agents as Fast and Slow Thinkers
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2e5864fe-54cc-4b85-b577-67c86d72d40b · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Kimi-VL Technical Report
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 58985c59-4a99-4994-b49d-8a8fb9d9a85e · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bacf6ec2-a998-426a-a4b6-1c514d5fa113 · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6a0e124b-8ccd-43a2-ac2b-00f51ff3bebd · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning V?: Guided visual search as a core mechanism in multimodal llms
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a705ae18-2b50-4fb0-a0b4-7c4504fa2e1f · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6c680704-a2cd-44a5-976e-9af560ee71dc · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Qwen2.5 Technical Report
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7d9f764b-c8d3-4086-8666-6f7e6d107054 · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Octopus: Embodied vision-language programmer from environmental feedback
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a3b664d2-7d71-49d1-b1c1-52360459a5fb · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Egolife: Towards egocentric life assistant
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0ebd4066-3e4d-4946-ae53-10c1276a1491 · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6ae173af-24c7-4ac9-8d56-a0b265d43824 · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 858439a1-fa26-45ca-af18-97a616067348 · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation beef2e43-a6f6-45ef-8106-e0dc21feb3a6 · outbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b1326d94-34ca-4045-b2c9-b94ba72f6096 · inbound
Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 75664f17-c17d-45fa-b449-11f4e9e77b56 · inbound
HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5beb4e0-9243-4940-a28a-9be871b815b3 · inbound
Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 94aa7353-924e-401a-9d55-d2c7562c497b · inbound
MARINER: A 3E-Driven Benchmark for Fine-Grained Perception and Complex Reasoning in Open-Water Environments High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3b6b1080-1e07-4cae-bac0-b75c37e863a1 · inbound
Latent Visual States for Efficient Multimodal Reasoning High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b645c24f-c040-4dba-99c3-b7ab268b3b23 · inbound
BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning
Reference 280
Source-reported events for the cited work
Unavailable: canonical work link unavailable.