Pith. sign in

Paper Citation Record · LEDGER

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning

As of 11 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 6 inbound Pith citation observations for arXiv:2507.05920.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.05920 v2

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-19T06:10:57.219445Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T20:19:14.175809Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T16:29:57.242422Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact31
  • verified fuzzy11
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8316a74b-7040-4fcf-9dd5-fb93fb15ac57 · outbound

This paper cites Qwen2.5-VL Technical Report.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Qwen2.5-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:12:07.121746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:138e3427dcc1fae383c86a22417b5d5bc8a7dbd06811e19681010bdf48f357f2

Observation a4bad2e2-5f09-4ddd-bdfd-72206fb5ec49 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:12:07.109330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:5507c5bd30a6bb687a5c9d4f289e624280f3920815ccb67131107903bf5b69a3

Observation c3340de6-38de-4680-a081-5f57c9a14964 · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:12:08.249189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:090c0902cefd87841d5058dbd52cf4b43154f433dba398f845aad52f64b85d03

Observation 1f07e7e0-eee0-4671-8179-07877bbbceeb · outbound

This paper cites an unresolved cited work.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-19T06:12:08.262107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:8c9b6e4f5c52afb63757a931347815ec38b59b06afbebba8ec9a6473040c4bb3

Observation cc991096-7495-4f71-8763-49a95021207b · outbound

This paper cites an unresolved cited work.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-05-19T06:12:08.265032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:361ee9a68948d05fc61552f0591ca627186a9c008a7dcafea4ef3186e28c0d40

Observation 26111992-6bc2-4e55-8c56-83fdbe0f1905 · outbound

This paper cites Patch n’pack: Navit, a vision transformer for any aspect ratio and resolution.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Patch n’pack: Navit, a vision transformer for any aspect ratio and resolution

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:12:08.232117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:7926975299447deab91dd909651a8d8ad13fc66d045b731f39354093cd1719e5

Observation 7f2e98f7-69b6-4ef1-a6ce-cad52c7a2f7e · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:12:07.024774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:175cca971ba713e73db1e13d582c2dd2d4b9dac14efab00790496e27bd991baf

Observation 932f6aae-56d4-4782-a543-4a0fdd953aae · outbound

This paper cites Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:12:07.105593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:f27e726a73505e4f8787a54872579bbdde9d40b12185cb276d99070556c144b4

Observation 5a506313-273e-48bf-ad42-05120d387ce9 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:12:07.049792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:47bc81a3b0e3dec1a49895bbca7c3e2897a5db68832e70f937c08374bd03da03

Observation d7eafe4e-337a-4b45-9220-d046863627a6 · outbound

This paper cites Llava-uhd: an lmm perceiving any aspect ratio and high- resolution images.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Llava-uhd: an lmm perceiving any aspect ratio and high- resolution images

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:12:08.239479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:63cd861647e577a1fadc93657278cc5503725612d4e127e2cdbe4b8358bc7fd1

Observation f33e2de3-2122-4ca7-9519-4496730a929c · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:12:07.073787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:7961e500f8c8c1c5440a6a528c564fe76c899f9a963901719454a28e59608525

Observation 410158e2-06a0-4f04-a543-e319dee066ca · outbound

This paper cites GPT-4o System Card.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning GPT-4o System Card

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:12:07.065804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:e1f4bcb0d9de0d44ed22cb72d8dad3645009c96cfff9ef585855540a9eb80b99

Observation f81de1b1-c5fb-4e6a-8362-81e7e1b1bbd0 · outbound

This paper cites OpenAI o1 System Card.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning OpenAI o1 System Card

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:12:07.093722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:d92021b0aa5b6558dc94b64a70118222bbed0aabb48eade4f3b1d4ff273b404c

Observation de95bd19-7d34-430d-881c-c35c877467a1 · outbound

This paper cites The hungarian method for the assignment problem.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning The hungarian method for the assignment problem

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:12:08.242912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:ece4d5eec11c77508bde9a71451ade0f604c94e4cdbb31be44919c6fd48081c5

Observation 796e8c22-5586-499a-9325-692b7ef248b0 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Efficient memory management for large language model serving with pagedattention

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:12:08.228285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:89b07cbe5eb98b04f0b809c22acc90edfcfeb46d57b1f9429c5013084bd7bc62

Observation 44f74fa9-d835-4992-8914-020ba9faaf52 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning LLaVA-OneVision: Easy Visual Task Transfer

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:12:07.097719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:b2777ed5586cd817cb8753f99b0ded405e587eeea2f10f80ffe6e9e07af969b1

Observation 80b8bca0-a559-4b5b-82a1-f870a84d7c70 · outbound

This paper cites Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:12:07.077961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:49da274a78ff1c638ffb2132a63f7dc1584e238f47d1e83e867825781c72b026

Observation 31c2b4e6-2685-4110-bcb6-df4e1146e7c6 · outbound

This paper cites Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:12:07.045708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:e35990aca433c3926a1dd5caac23a367104e3aa4722472fd260c0f4ef545d817

Observation 12d30eed-5ba8-4cab-aead-1be6a388f186 · outbound

This paper cites Improved baselines with visual instruction tuning.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Improved baselines with visual instruction tuning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:12:08.224154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:a3608d9fffc9311eb305f6c5f80a6a1caeb963ac0dd6dc571430d3d3f1e86999

Observation eb6a274d-38e4-4fd2-a979-ceed535e3c69 · outbound

This paper cites Lost in the Middle: How Language Models Use Long Contexts.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Lost in the Middle: How Language Models Use Long Contexts

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:12:07.069754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:8fd80db75a5a74e124b687dcc2762cfeb5257ae8423c11dadd5092b050255a59

Observation f5616df8-bba1-4eb2-a19e-34372d632036 · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:12:07.016238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:523dffcdf99ea1bd9890f727057c0086f33bd974b21bf5e64a53b57a28cc3f55

Observation cd23a595-bf5b-4d0f-8550-15b82f3f5b94 · outbound

This paper cites Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:12:07.081795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:472daa890248ad6a57438148e62ef6ce9d55f07efce608a1136d9ef43cc3ba42

Observation 891797d4-fed9-483a-87f6-fb8f24729b7f · outbound

This paper cites Ola: Pushing the Frontiers of Omni-Modal Language Model.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T06:12:07.011738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:6f5f41e6d50a443e800da3b08798767904c27acfc227f71a9d3127c8d42171e0

Observation b10b30ed-54ee-4cb2-a7da-5d398020fa5f · outbound

This paper cites Decoupled Weight Decay Regularization.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Decoupled Weight Decay Regularization

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:12:07.053607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:617dd9d73ee2f0eb596309d5b8cec58213c88a08dbbdea8bcbb2ccf184b2da81

Observation 89bfe261-63ba-4bc5-ac2f-b5aed3907ad6 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:12:07.062177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:4df3cc6439935b139bc791a2865293c988ad8e0589b8b77b00d9ff807726947a

Observation e339be48-cd2c-4a14-af0e-e90af4fbda39 · outbound

This paper cites Openai o3 and o4-mini system card.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Openai o3 and o4-mini system card

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:12:08.268271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:d9033aa764af54017c39a769956a63dfa192904ce4f8417a18002d455ea028bf

Observation 9bfa58a6-b28e-4251-a279-0ff11e7a0dea · outbound

This paper cites Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:12:07.003212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:2dee283dc05d6d97ca9afad815d37cf9b4ca54e12fbf57c8d225cb16a0de11f8

Observation c115eee6-bf3d-4e05-b8f2-01510781adba · outbound

This paper cites Learning to count everything.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Learning to count everything

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:12:08.255788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:39f3ba6096a9963fd8fc1e80cd79c17235deff6f7b6037a2102d9c48664f3a15

Observation 2c9a6eb4-748c-41c1-848f-c1349661304c · outbound

This paper cites Proximal Policy Optimization Algorithms.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:12:07.028914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:c38ff8ab39929c96e783cf147d7d64d399b6061e5e5471685398ad65ef76b88b

Observation c7d4c3a1-5c5c-4b2c-8f3e-4999b9a823b3 · outbound

This paper cites Visual cot: Advancing multi-modal language models with a comprehen- sive dataset and benchmark for chain-of-thought reasoning.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Visual cot: Advancing multi-modal language models with a comprehen- sive dataset and benchmark for chain-of-thought reasoning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:12:08.252592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:d4711e9a1755fc7a5cf046e3690ee409a342945fa3b2f058fde8beb19c5f4615

Observation fd731518-f764-4cfc-a7c4-1b8f30ff4254 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:12:07.033133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:f8074f1725f8228577c4e9eb08e2f68767a0c2c9d5b6b7b260b0f8fad8c2142f

Observation 923604e5-c07d-4a78-a39f-dffbfaaf30f4 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:12:07.037396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:8eb475857063c5d9f5a78e7f8d6fbbf5c5ba18dd1f15d72b100bb196ce728d11

Observation 7ddb696f-2b3d-41c6-ab51-48cf9469cb21 · outbound

This paper cites Scaling Vision Pre-Training to 4K Resolution.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Scaling Vision Pre-Training to 4K Resolution

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:12:06.991297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:776288a81e6401aa8c9d6586934b8dc10f181f40d59030625677a3bade867523

Observation 241fea74-b9e2-4e6f-92e7-1a40a79df36c · outbound

This paper cites an unresolved cited work.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-05-19T06:12:08.245694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:a852e97c81f73cf0724bbdfc985b51b06fbe6b7551b75a4570a7f2baa17f9c1b

Observation 1be0386f-ef27-4a47-a10d-675bc08961c9 · outbound

This paper cites Visual Agents as Fast and Slow Thinkers.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Visual Agents as Fast and Slow Thinkers

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:12:06.995531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:9696fe492b7286a40add90965f3bc73d9eaa6b8ea193842cabed061d2504105f

Observation 2e5864fe-54cc-4b85-b577-67c86d72d40b · outbound

This paper cites Kimi-VL Technical Report.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Kimi-VL Technical Report

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:12:07.007047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:321188ce6fded132dbf5ed2f3a3970a96d222f0140c18df3ad18bffb2e2a3e6f

Observation 58985c59-4a99-4994-b49d-8a8fb9d9a85e · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:12:07.117138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:5c08036a06807bbfad3578534bd36e197ce35ba1db9e40ea0a3ec2794c2a0c09

Observation bacf6ec2-a998-426a-a4b6-1c514d5fa113 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:12:07.090045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:0a4f3cb330cacf0d5d75bc28fd25dd17e0ddab478175e864bbe6c70ed72f3a30

Observation 6a0e124b-8ccd-43a2-ac2b-00f51ff3bebd · outbound

This paper cites V?: Guided visual search as a core mechanism in multimodal llms.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning V?: Guided visual search as a core mechanism in multimodal llms

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:12:08.258959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:6967b46d11b19f14e92c611a6fc7849455ec29bbcefcef2875cdac7b1984cd45

Observation a705ae18-2b50-4fb0-a0b4-7c4504fa2e1f · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:12:07.041323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:66479efcf3f2e12d6902e2fff840536cade945ddd070993dff1d2f1aed5edead

Observation 6c680704-a2cd-44a5-976e-9af560ee71dc · outbound

This paper cites Qwen2.5 Technical Report.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Qwen2.5 Technical Report

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:12:07.020138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:f3d27d0466a7dac9742990f7d3aea6b279727e9f692899828d76859cd0da179e

Observation 7d9f764b-c8d3-4086-8666-6f7e6d107054 · outbound

This paper cites Octopus: Embodied vision-language programmer from environmental feedback.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Octopus: Embodied vision-language programmer from environmental feedback

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:12:08.235939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:1f729c407a90ffe1fc20056e3ed6bcd8e3592732aa51127c245ed01e7202f8c7

Observation a3b664d2-7d71-49d1-b1c1-52360459a5fb · outbound

This paper cites Egolife: Towards egocentric life assistant.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Egolife: Towards egocentric life assistant

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:12:06.999080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:a5fd5a5c9b80b361d32b2dc7c014310f2c053bb6e017a832ca5af68e0475f018

Observation 0ebd4066-3e4d-4946-ae53-10c1276a1491 · outbound

This paper cites R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:12:06.987250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:811bb1e565ace65dbfb5811fe9a8c7155f0d6c9f0fcf77cf3bd385081c881311

Observation 6ae173af-24c7-4ac9-8d56-a0b265d43824 · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:12:07.057635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:f79aa5f9249e7ebfafd4e0bc0ced25c557693e9305c9d23daf93175bb82cee8c

Observation 858439a1-fa26-45ca-af18-97a616067348 · outbound

This paper cites Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:12:07.101562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:c16841002ec9f1c3c4279cd228d3adbf5964e4b1ceff68e65e8ffb178625b125

Observation beef2e43-a6f6-45ef-8106-e0dc21feb3a6 · outbound

This paper cites MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?

Reference 49

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T06:12:07.086160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:b282ac31a2e0bfcaa5e1e4afabd0b5df9b02b81e66b3eb4d31bc5cbfd79a9e0e

Pith citing papers

Observation b1326d94-34ca-4045-b2c9-b94ba72f6096 · inbound

Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search cites this paper.

Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:17:55.662734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T01:17:55.500268Z digest=sha256:f96aeae6901cf92560282d6571987e2f5278dd5dc4f5883a01d6adec944c28bb

Observation 75664f17-c17d-45fa-b449-11f4e9e77b56 · inbound

HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework cites this paper.

HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T20:19:14.175809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:19:14.175809Z digest=sha256:b9b2cf36f857f8432e28ad38bc4c088dd1fdf9bae9217ae614a7f1a2f396096b

Observation f5beb4e0-9243-4940-a28a-9be871b815b3 · inbound

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization cites this paper.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:45:50.112672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:62c81f6de3a0909a3058e9a6c45ec0c5f149a184a6437171788bdc54b0ceca11

Observation 94aa7353-924e-401a-9d55-d2c7562c497b · inbound

MARINER: A 3E-Driven Benchmark for Fine-Grained Perception and Complex Reasoning in Open-Water Environments cites this paper.

MARINER: A 3E-Driven Benchmark for Fine-Grained Perception and Complex Reasoning in Open-Water Environments High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:11:00.807634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T17:46:15.107175Z digest=sha256:134cdf2559d1928c8af702c4f0384891552ee007a2c2413d760829e30682db11

Observation 3b6b1080-1e07-4cae-bac0-b75c37e863a1 · inbound

Latent Visual States for Efficient Multimodal Reasoning cites this paper.

Latent Visual States for Efficient Multimodal Reasoning High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-04T16:29:57.243827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T00:38:11.619574Z digest=sha256:f5d0c3895177998af829253becac2f902240d9caf40289ceb9e58ea7f996448a

Observation b645c24f-c040-4dba-99c3-b7ab268b3b23 · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning

Reference 280

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:4d2461548ac194e5b7d641500acf87da31e4dce179cf7983e88d3d7303a40e89