Pith. sign in

Paper Citation Record · LEDGER

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models

As of 6 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 1 inbound Pith citation observation for arXiv:2605.20177.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.20177 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T14:46:30.334667Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact21
  • verified fuzzy2
  • unresolved1
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch5

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 58b2735c-2917-443e-abb3-75e0bff75d96 · outbound

This paper cites Nore- geo: Non-reasoning geometry benchmark.arXiv preprint arXiv:2601.10254.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models Nore- geo: Non-reasoning geometry benchmark.arXiv preprint arXiv:2601.10254

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:13:21.522642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T05:13:03.237427Z digest=sha256:bc85b13341a619c94616ec39f2873a0af90521268dbdb2b791a84fe8acfd7b3f

Observation 4b90cb14-efeb-4715-a8d3-433eb05e6166 · outbound

This paper cites Qwen3-VL Technical Report.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models Qwen3-VL Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:13:21.535514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T05:13:03.237427Z digest=sha256:502636a162b16d70758fceecfad695529d498125be2e3e21203cdbc444ab4176

Observation ec0b5e64-b557-4071-bc35-f14648cee6bc · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T05:13:21.508096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T05:13:03.237427Z digest=sha256:2dc99e997bdc349e412938b78a05f21aded6e6d19b0632c533b4ce78a244b9f3

Observation 8445d69d-6a82-432a-bee2-4873c9088011 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:13:21.545362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T05:13:03.237427Z digest=sha256:94f4e0feb70e7510c34abc250c4717fbf25e6e14ca56e990e21d70c932b912d3

Observation 5715ec37-505e-47c5-b169-dff5fe7e45ab · outbound

This paper cites G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:13:21.504876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T05:13:03.237427Z digest=sha256:12ad82681813ac08c7dafaa1baf14112eeacd50da96eb9ae8e950e97028faed4

Observation e07336f4-347c-40a0-9901-4679646b9a37 · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models CogVLM2: Visual Language Models for Image and Video Understanding

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T05:13:21.481162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:ddc7bc7d2d9b6909e5872f3e48ddda3c33b874926c7cb0373f5c879f0ed2bb14

Observation 51ea26a9-8f75-4e2d-bd1e-165529af6f56 · outbound

This paper cites Do Vision-Language Models Really Understand Visual Language?.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models Do Vision-Language Models Really Understand Visual Language?

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:13:21.514602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T05:13:03.237427Z digest=sha256:159164f6c03db4568428f3da4ac7dd2efffb68c7fca49418f57fdaf5357e3863

Observation 3e7971cf-bbaf-4d2c-8221-235a3d504af6 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:13:21.561611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T05:13:03.237427Z digest=sha256:72a539dd34b3d24222d277a2b04b5af9c56b3833b9a0885dcb607d5772376179

Observation ec38e397-937d-4a55-b115-949d5acba2ab · outbound

This paper cites arXiv preprint arXiv:2508.02669 , year=.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models arXiv preprint arXiv:2508.02669 , year=

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:13:21.525956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T05:13:03.237427Z digest=sha256:6d85e752d59235898662129d97596345cef0691274a107f2f806b0167bb5aa4d

Observation d7aa1115-cc2c-4ed1-9598-d22cb6d7878d · outbound

This paper cites GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T05:13:21.546757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T05:13:03.237427Z digest=sha256:eb1fa64b6f91108b3d338f3fb8af1b474aec040f1db85d11108528fb1cca72fb

Observation 4a1f0191-67c1-401c-a645-e41a857a2b47 · outbound

This paper cites VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:13:21.501063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T05:13:03.237427Z digest=sha256:a65440d6f4352166228ade0039853daa993cf59ddef3aa65d2b5550d94c69615

Observation 9fa14ff1-e144-4db8-936a-0e635473e881 · outbound

This paper cites Mmr1: Enhancing multimodal reasoning with variance-aware sampling and open resources.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models Mmr1: Enhancing multimodal reasoning with variance-aware sampling and open resources

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:13:21.560557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T05:13:03.237427Z digest=sha256:657755dfadc4baf12675a9a45712ffab3abf1f40ab4cc4867aa27571746cd7da

Observation c5eaa9fa-2347-4767-9243-568217237c48 · outbound

This paper cites Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:13:21.507568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T05:13:03.237427Z digest=sha256:cee00d6e4b405ecb1236e9c3dc14889676d7cb45705125600b06855f2ce28e6d

Observation 8ee5a468-1450-4a4b-baaf-9f19392b0e77 · outbound

This paper cites VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:17:25.615471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T05:13:03.237427Z digest=sha256:ed3e9f4aeec715619e685328222b7103fbd4721f1efae7907d5b7720a2b11866

Observation 8161aa63-f862-49da-ae67-ed2dca173ef3 · outbound

This paper cites CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:13:21.543836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T05:13:03.237427Z digest=sha256:cb99304e32e89f5e615321a59f47c9ece6703060b3deebb01bbfa2b20849b775

Observation ca1ce379-d2b8-419e-bb4b-08a71fff2d98 · outbound

This paper cites More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T05:13:21.551928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T05:13:03.237427Z digest=sha256:ee1ec9b64ba0344b594c1b62767ab89fb2ea20677f6d71ca5c09a4862e0e55ce

Observation fd501f9b-ba9e-4480-bdc0-61dba6a73518 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models Improved Baselines with Visual Instruction Tuning

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:13:21.557487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T05:13:03.237427Z digest=sha256:8225866afbb9b8d66404508b909d976a994206f1840b2768d1dafce717006213

Observation 36e93f43-8a0a-4285-acb9-a2c247a2eb60 · outbound

This paper cites OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:13:21.570421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T05:13:03.237427Z digest=sha256:173c0f6a7a78c84890dfa665e7b9aa6078cf1abc4971f9acf67ebf0467ddac19

Observation 1020f756-3f7b-4e10-911d-ebb048065042 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T05:13:21.573194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T05:13:03.237427Z digest=sha256:411d57c75a8031ab3c2b711915553744d851aecce8cd7d76aeb2bf3e0c43a77d

Observation 285a4026-1e00-4f33-a829-679f7d2deeb3 · outbound

This paper cites LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:13:21.567258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T05:13:03.237427Z digest=sha256:21af4ca4759d46c5c08ec262fbf236c0592b24a92f4a56abf933d31b169096b7

Observation ae7bedce-d04a-4558-ab12-656c1ddcfc70 · outbound

This paper cites A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:13:21.565954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T05:13:03.237427Z digest=sha256:6065ee1441fa1ecb50e8e4e56359f538896326e4c491db92585cdecf8feee333

Observation ab60bffd-42be-4c82-a373-9b4b783f5001 · outbound

This paper cites Descriptive caption enhancement with visual specialists for multimodal perception.arXiv preprint arXiv:2412.14233.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models Descriptive caption enhancement with visual specialists for multimodal perception.arXiv preprint arXiv:2412.14233

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:13:21.519154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T05:13:03.237427Z digest=sha256:aca8b9e2412831bb72b99b9a5f98e42ce1d26773ab3e3d1b078797bfb71007f2

Observation 04a38104-5c6b-48cf-8edf-7f6ac1384ef9 · outbound

This paper cites LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:13:21.529260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T05:13:03.237427Z digest=sha256:8c173450b500a7633ec6765af2aea237366bb0da3c44ada98b594a7544f620a3

Observation 11049ae2-ecda-41a2-9b5f-b6b5004bc4ba · outbound

This paper cites an unresolved cited work.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-05-20T05:13:21.578814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T05:13:03.237427Z digest=sha256:4b1c7d074b705c3efb10e4f8867f3e8378b09641290d3a50cd9291555d5977f1

Observation 6983eb48-fcc7-4ed4-9e77-1de6f3922bfd · outbound

This paper cites WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:13:21.563530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T05:13:03.237427Z digest=sha256:92fac76a9ed0efda811874ce9f010b9adacc8f4d223f925535b155ef1eca8c29

Observation 8160789e-f389-48e2-ab36-3b6de0efeffc · outbound

This paper cites GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:13:21.576609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T05:13:03.237427Z digest=sha256:9cf1b3c00802145c365e6e6f51075b2ba06838c6190353f293d8d33c14bd35b9

Observation d7a7893f-c13f-4f80-ba1d-8d33346d535b · outbound

This paper cites Improve Vision Language Model Chain-of-thought Reasoning.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models Improve Vision Language Model Chain-of-thought Reasoning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:13:21.549579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T05:13:03.237427Z digest=sha256:89165eb577405a55b19e4a1e3670e23d12f1f121420b49f38bf66b593eedac71

Observation b85715cc-147b-4f31-a7f8-6813f5d76688 · outbound

This paper cites Dynamath: A dynamic visual benchmark for evaluating mathematical reasoning robustness of vision language models.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models Dynamath: A dynamic visual benchmark for evaluating mathematical reasoning robustness of vision language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:13:21.583182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T05:13:03.237427Z digest=sha256:5da16da5ee7e9c4d835663dfac8cef11eee6a5f1bb800a3f299a6870c328a753

Observation 44244442-f5a8-44d9-a6a3-b3e34868c57c · outbound

This paper cites Table 6.Key hyperparameters used in our Stage-3 training.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models Table 6.Key hyperparameters used in our Stage-3 training

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:13:21.581098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T05:13:03.237427Z digest=sha256:798f177cd84835792a82d09ba13984fe2723adece4ffd1437adfe6d9d7c9d961

Observation dba7d2c7-074f-4993-aa57-2f2a59abf1ee · outbound

This paper cites an unresolved cited work.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models Unresolved cited work

Reference 30

Resolution
malformed identifier
arxiv_id, observed 2026-05-20T05:13:21.555039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T05:13:03.237427Z digest=sha256:314cfc160295f02cc2c8ffd14440e79a7dd42e3113a3f3572433864a4dae93de

Pith citing papers

Observation 7d07c8f7-0df1-49ad-83bf-b49a503c4082 · inbound

RP-OPSD: Resolution-Privileged On-Policy Self-Distillation for Multimodal Large Language Models cites this paper.

RP-OPSD: Resolution-Privileged On-Policy Self-Distillation for Multimodal Large Language Models From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-31T14:46:30.334667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T14:46:30.334667Z digest=sha256:050afe33bfa7ec63225c977e14569318e8cd807a196b2d443cd6334355787962