Pith. sign in

Paper Citation Record · LEDGER

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning

As of 8 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 4 inbound Pith citation observations for arXiv:2507.13362.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.13362 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:54:48.162766Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-25T05:51:18.233043Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T05:55:25.118073Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact1
  • verified fuzzy6
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 026491a9-7bf9-4eb7-93f4-0d6ff127268f · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.058714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.058714Z digest=sha256:edc049e765410b5f704e051f2ba1b75f5d501331ee4c62b0d3e6766f7abf47ea

Observation 660692e1-7f6f-4f16-a9a7-57b0b0c008ba · outbound

This paper cites Visual Spatial Reasoning.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Visual Spatial Reasoning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.062289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.062289Z digest=sha256:0dc5cf00889aac3f0121ba82b91bd375eb755de70ef6d11daf8f6d1477e64e85

Observation 7237e9ee-d699-4820-a06a-8563f308b19f · outbound

This paper cites Sat: Dynamic spatial aptitude training for multimodal language models,.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Sat: Dynamic spatial aptitude training for multimodal language models,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.065213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.065213Z digest=sha256:413222a5f1ee5df7a1497ecd1900f1801c4ac43a4acab95078af6c523d100167

Observation 2cede708-7895-4775-8ce0-f9e6f9ce8811 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.068292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.068292Z digest=sha256:f84bfb305ac354aa675585360ebe6dc3aef983d9f0e7bbc57194a9bbc68cb9fb

Observation 8b8cc7d2-446f-4491-bc6b-7dfa4da6239e · outbound

This paper cites Clevr: A diagnostic dataset for compositional language and elementary visual reasoning,.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Clevr: A diagnostic dataset for compositional language and elementary visual reasoning,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:48.613029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:54:48.071617Z digest=sha256:edc49eac6c0880ec0666a98ea80f82053e347d55b73f5b6b6e26058dcac81022

Observation 344944d5-44eb-4806-a839-9e0064ee6aed · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.074302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.074302Z digest=sha256:3290cd68ec3b5d8195bd472883418eb1a626fcecf8ee7460d46f3fdfdf42fd0d

Observation 31ce173d-4c71-403f-8b5b-69c00dc94c58 · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.077400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.077400Z digest=sha256:c63ee57847192f0b5e334d0ff33d9ca61a94ec713216cbb4b28c75047231d756

Observation f0f284c0-8e9b-41f3-93a9-c12961eda8de · outbound

This paper cites R1-v: Reinforcing super generalization ability in vision-language models with less than $3,.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning R1-v: Reinforcing super generalization ability in vision-language models with less than $3,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:48.604078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:54:48.079868Z digest=sha256:4c8ac28cb1ee54dc5859d29fb7a56811b32d4747953a6f23c40b01f14108142b

Observation b5752c06-a821-4f2b-832b-c9732a089370 · outbound

This paper cites Rlvr in vision language models: Findings, questions and directions,.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Rlvr in vision language models: Findings, questions and directions,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:48.595750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:54:48.082191Z digest=sha256:6bb0f512db60f7cb011fb20dbf2fde31cee0209f55002e0a2cd5700e24ad052b

Observation 93e83563-3570-4ed1-b0e7-64ee2a7f3db7 · outbound

This paper cites R1-zero’s.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning R1-zero’s

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:48.587733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:54:48.084730Z digest=sha256:bbde7ff1a7ba0ba90ec6e8ce3443dcbf2425dde2ada58ae7650df5c0f178e4e5

Observation c9714968-8465-4d2e-960e-7030c0c7d297 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.090566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.090566Z digest=sha256:d262e908d2315bf00e68e1ae2459e885d449e2d48894e3ef759e8352915b6c34

Observation 1aeb14d0-f97e-4d0c-b62a-c21146162283 · outbound

This paper cites Think or not think: A study of explicit thinking in rule-based visual reinforcement fine-tuning,.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Think or not think: A study of explicit thinking in rule-based visual reinforcement fine-tuning,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.093206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.093206Z digest=sha256:ebbbce3ed287e344f96738cf362a2866f10817fb5a66bee1c38bad4cdaff1daf

Observation 459ba278-8570-4a00-a6cf-778ca1f6a78e · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.095753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.095753Z digest=sha256:9758587a3a5ec161849d5b04aa041cfdfee1c737c4492c5cc0d5ae0c7ceda701

Observation 403e28e2-fd40-4f71-9f1f-50b839a7a5a1 · outbound

This paper cites Spatial-r1: Enhancing mllms in video spatial reasoning,.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Spatial-r1: Enhancing mllms in video spatial reasoning,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:48.579259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:54:48.098417Z digest=sha256:f782ea08efe786a79d2155684385279836ffa24e54fc04cd112246dd9ce80ca1

Observation d1911141-ba6d-45a0-9592-05121f0c56cf · outbound

This paper cites Improved Visual-Spatial Reasoning via R1-Zero-Like Training.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.103936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.103936Z digest=sha256:d6737522f9f762171d76962d441a2d3fc0b93e9e18d7a39a8a3ec3bbd5b954cf

Observation 08e55a4f-c9fb-4548-899a-300f02ff56b2 · outbound

This paper cites SpaceR: Reinforcing MLLMs in Video Spatial Reasoning.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning SpaceR: Reinforcing MLLMs in Video Spatial Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.101125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.101125Z digest=sha256:891dc9502ac12a3593177c1d08137868718da86c68adfe6854a0f53afc718fad

Observation ef361559-baec-48aa-b250-31ffba03bdd6 · outbound

This paper cites Compositional Chain-of-Thought Prompting for Large Multimodal Models.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Compositional Chain-of-Thought Prompting for Large Multimodal Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.110228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.110228Z digest=sha256:31c48bba42190335f7edb1fface6686d78eff86ede03eaafa701332ea5be7d1d

Observation 8bcc650d-7dcb-4299-bb70-9ff14e903ad3 · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.107118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.107118Z digest=sha256:be9ad65c2f6ee096d3aec267486b261e3fbd352676e46166a23842ca0180075e

Observation 5666202e-5beb-4160-8ff1-4c59f3288bbb · outbound

This paper cites Mind's Eye of LLMs: Visualization-of-Thought Elicits Spatial Reasoning in Large Language Models.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Mind's Eye of LLMs: Visualization-of-Thought Elicits Spatial Reasoning in Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.116134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.116134Z digest=sha256:4bf1cf932f2c402b2ca1daa4fca321cefc38b65c8baa2109d8edbfadfa1bbe47

Observation 474139ab-99b5-46cc-801e-99b140d37b0b · outbound

This paper cites SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.113277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.113277Z digest=sha256:ebab0239b58ad2075e288fed6a4c10308c15671da8a9544663a8475cf8c5786a

Observation b4a23978-495e-42ab-a85a-5013e648f521 · outbound

This paper cites Clevr cogent valb,.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Clevr cogent valb,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:48.571061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:54:48.122022Z digest=sha256:f5839bbe6d5432883cea668fd6714f761673f639d97bf4128b2fae9844390a4e

Observation a241cba3-ce08-431e-90ed-5525cdbb95cc · outbound

This paper cites Chain-of-Symbol Prompting Elicits Planning in Large Langauge Models.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Chain-of-Symbol Prompting Elicits Planning in Large Langauge Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.118975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.118975Z digest=sha256:4042ea51b876f8be4d862670700b7d405949855256fa3ab4bc1bccb044a62e92

Observation 7f53ab88-b6a3-4746-9d32-22d246d838f3 · outbound

This paper cites Depth Anything V2.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Depth Anything V2

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.127358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.127358Z digest=sha256:edaa3cfcadc37e5d02ea06ff7fbd17acd69c7ea64c5ba36b5c72201d9b6eb86a

Observation e63589c4-b6e1-4889-b392-73cd413e2956 · outbound

This paper cites Super-CLEVR: A Virtual Benchmark to Diagnose Domain Robustness in Visual Reasoning.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Super-CLEVR: A Virtual Benchmark to Diagnose Domain Robustness in Visual Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.124661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.124661Z digest=sha256:48fad91f340998d02fd325d977716cf41f912713ec87d569b644f6cb1c6625a7

Observation 6fe160ae-13f8-4435-baf3-d27e065c7463 · outbound

This paper cites ClipSAM: CLIP and SAM Collaboration for Zero-Shot Anomaly Segmentation.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning ClipSAM: CLIP and SAM Collaboration for Zero-Shot Anomaly Segmentation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.132799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.132799Z digest=sha256:8fe4a8060ac2f3a82cb1020feafed735fbaaf9fac887aac89da44616c0ca90cf

Observation 426118d5-58bd-4f7f-a067-ed12597e2dc4 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.130017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.130017Z digest=sha256:3db1f5cb57fe4d1bce24e9ae974914dd72a59d32786d54f373fb3b8b5263dde5

Observation a27f093c-9609-49fb-8719-68d3b7ce78f9 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.139134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.139134Z digest=sha256:09535d6f92c63d24db6de2c841f50fa3b697951dd2232a01956f494fe42c3c35

Observation 5b1b369b-9af6-4c77-bc28-85882706c23c · outbound

This paper cites PaliGemma 2: A Family of Versatile VLMs for Transfer.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning PaliGemma 2: A Family of Versatile VLMs for Transfer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.136009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.136009Z digest=sha256:c651e267138ce4fabe1037953d37340f26249c35119d6b9c9f61c2b6e99b234f

Observation 50ae8837-e520-4134-81e2-36491a27c56e · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.144642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.144642Z digest=sha256:eae4b0c2978fb72772056b1c7f5f418fa1d21a48e43168c144df8bdb9486153d

Observation fb0e9815-b80d-47f6-b11a-6cf7316cedc1 · outbound

This paper cites Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.142033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.142033Z digest=sha256:9573f532ea8ff9ca1bd4c7a9bb8360035df2d9cb8cd3594ddf17d108bc587956

Observation 248d13b4-d54a-449a-91f8-49431fa02145 · outbound

This paper cites Training Large Language Models to Reason in a Continuous Latent Space.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Training Large Language Models to Reason in a Continuous Latent Space

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.150801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.150801Z digest=sha256:9b2c2949fe3d4d820251a8aa658844c30c381311c00ff4f567fde469915ffe46

Observation 25ecb92a-5ce5-4dd3-9db3-cc0770f60934 · outbound

This paper cites EmbodiedVSR: Dynamic Scene Graph-Guided Chain-of-Thought Reasoning for Visual Spatial Tasks.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning EmbodiedVSR: Dynamic Scene Graph-Guided Chain-of-Thought Reasoning for Visual Spatial Tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.147566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.147566Z digest=sha256:7c52bab128dae3e0954530d566a8bb257bfb79cf050b124557d1cf9123b13b91

Observation 46a2075a-2a04-4750-af4a-544bd439c12c · outbound

This paper cites SpatialBot: Precise Spatial Understanding with Vision Language Models.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.156627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.156627Z digest=sha256:ff100173833dcdff69dee98fb09a35d7180cc589209c396c6196d0ca279d28da

Observation 9faaa57f-7cfc-4f53-a10a-cf3e13efeb9a · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.153646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.153646Z digest=sha256:652cea5c905fbea532dc73d8d3428efa0876e16023ad54669e4f3f7fba5adff2

Observation e7f85b5b-2a86-40c6-b069-8646390f3e0c · outbound

This paper cites SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.162766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.162766Z digest=sha256:92419f56c4a0c43b22cd52b8207efadc978d38d4252e0d725d4d2996dba640d5

Observation bd595542-d108-48c3-8ae1-71279158c07e · outbound

This paper cites SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:54:48.204284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:54:48.159895Z digest=sha256:ca6c24fd798612ec667e95051b0bc13103ff4ff0151c322c9fd64ef79e0649c3

Observation d72bffca-7bf1-4666-bb7f-eb04a5452b88 · outbound

This paper cites R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.087387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.087387Z digest=sha256:960c8575a32e13b3d47f3b6ffdf77c56b9c5af933b9ab69efe0494284e19e4f2

Pith citing papers

Observation 205bf49e-2912-4f14-97fd-2c0ad26eddea · inbound

Does RLVR Extend Reasoning Boundaries? Investigating Capability Expansion in Vision-Language Models cites this paper.

Does RLVR Extend Reasoning Boundaries? Investigating Capability Expansion in Vision-Language Models Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:00:34.244407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T00:59:29.584299Z digest=sha256:68ff10846b12c25306391edd23afdecaac234d8e7b45173fdb5e49b049284044

Observation 8864a7cd-4258-41a7-8c93-e9db60e40cd3 · inbound

SPATIOROUTE: Dynamic Prompt Routing for Zero-Shot Spatial Reasoning cites this paper.

SPATIOROUTE: Dynamic Prompt Routing for Zero-Shot Spatial Reasoning Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:08:13.509580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T11:06:29.022994Z digest=sha256:04cae3411398a79c2985f772942d3f03904334f7c5683e51adb587735af67df9

Observation fafb401a-850a-41f4-a92c-12801b397889 · inbound

Eyes on VLM: Benchmarking Gaze Following and Social Gaze Prediction in Vision Language Models cites this paper.

Eyes on VLM: Benchmarking Gaze Following and Social Gaze Prediction in Vision Language Models Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:58:04.874738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T05:57:45.964145Z digest=sha256:ea9eb5273c490c16c9a1535317e9687c61f9d07439966dccc4e5d72911a4ffdd

Observation 12a1fb11-346b-4f01-b440-4f0280f224b8 · inbound

Eyes on VLM: Benchmarking Gaze Following and Social Gaze Prediction in Vision Language Models cites this paper.

Eyes on VLM: Benchmarking Gaze Following and Social Gaze Prediction in Vision Language Models Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:55:25.121803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T05:51:18.233043Z digest=sha256:27bee3104203d19e9c303cc9329031b77a635fed296c8dd17f56da0731d11e5c