Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 1 inbound Pith citation observation for arXiv:2605.20177.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-31T14:46:30.334667Z
A source-named dated measurement, never combined with another source.
Source: cited_works
30 of 30 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 58b2735c-2917-443e-abb3-75e0bff75d96 · outbound
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models Nore- geo: Non-reasoning geometry benchmark.arXiv preprint arXiv:2601.10254
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4b90cb14-efeb-4715-a8d3-433eb05e6166 · outbound
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models Qwen3-VL Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ec0b5e64-b557-4071-bc35-f14648cee6bc · outbound
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8445d69d-6a82-432a-bee2-4873c9088011 · outbound
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models Video-R1: Reinforcing Video Reasoning in MLLMs
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5715ec37-505e-47c5-b169-dff5fe7e45ab · outbound
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e07336f4-347c-40a0-9901-4679646b9a37 · outbound
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models CogVLM2: Visual Language Models for Image and Video Understanding
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 51ea26a9-8f75-4e2d-bd1e-165529af6f56 · outbound
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models Do Vision-Language Models Really Understand Visual Language?
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3e7971cf-bbaf-4d2c-8221-235a3d504af6 · outbound
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ec38e397-937d-4a55-b115-949d5acba2ab · outbound
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models arXiv preprint arXiv:2508.02669 , year=
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d7aa1115-cc2c-4ed1-9598-d22cb6d7878d · outbound
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4a1f0191-67c1-401c-a645-e41a857a2b47 · outbound
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9fa14ff1-e144-4db8-936a-0e635473e881 · outbound
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models Mmr1: Enhancing multimodal reasoning with variance-aware sampling and open resources
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c5eaa9fa-2347-4767-9243-568217237c48 · outbound
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8ee5a468-1450-4a4b-baaf-9f19392b0e77 · outbound
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8161aa63-f862-49da-ae67-ed2dca173ef3 · outbound
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ca1ce379-d2b8-419e-bb4b-08a71fff2d98 · outbound
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fd501f9b-ba9e-4480-bdc0-61dba6a73518 · outbound
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models Improved Baselines with Visual Instruction Tuning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 36e93f43-8a0a-4285-acb9-a2c247a2eb60 · outbound
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1020f756-3f7b-4e10-911d-ebb048065042 · outbound
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 285a4026-1e00-4f33-a829-679f7d2deeb3 · outbound
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ae7bedce-d04a-4558-ab12-656c1ddcfc70 · outbound
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ab60bffd-42be-4c82-a373-9b4b783f5001 · outbound
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models Descriptive caption enhancement with visual specialists for multimodal perception.arXiv preprint arXiv:2412.14233
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 04a38104-5c6b-48cf-8edf-7f6ac1384ef9 · outbound
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 11049ae2-ecda-41a2-9b5f-b6b5004bc4ba · outbound
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6983eb48-fcc7-4ed4-9e77-1de6f3922bfd · outbound
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8160789e-f389-48e2-ab36-3b6de0efeffc · outbound
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d7a7893f-c13f-4f80-ba1d-8d33346d535b · outbound
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models Improve Vision Language Model Chain-of-thought Reasoning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b85715cc-147b-4f31-a7f8-6813f5d76688 · outbound
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models Dynamath: A dynamic visual benchmark for evaluating mathematical reasoning robustness of vision language models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 44244442-f5a8-44d9-a6a3-b3e34868c57c · outbound
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models Table 6.Key hyperparameters used in our Stage-3 training
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dba7d2c7-074f-4993-aa57-2f2a59abf1ee · outbound
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7d07c8f7-0df1-49ad-83bf-b49a503c4082 · inbound
RP-OPSD: Resolution-Privileged On-Policy Self-Distillation for Multimodal Large Language Models From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.