Pith. sign in

Paper Citation Record · LEDGER

Improved Visual-Spatial Reasoning via R1-Zero-Like Training

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2504.00883.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.00883 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:31:18.424642Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T16:18:37.298889Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cfa03b87-9845-4c98-a800-14a7f564b40f · inbound

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding cites this paper.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:40:06.550790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:4c53e8cac5624e93f374ba4ef35cc2bce244a12a882a27645ed17d1354003ae7

Observation f7b0dbec-41e6-4426-b8ee-c8dd450953f1 · inbound

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models cites this paper.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 108

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:18.424642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:18.424642Z digest=sha256:d800902cffd26656cf04fe7825e499c5b429daf5c8b297ec0538d883144a0f19

Observation ff15f0c3-de02-41ae-b59b-5e0d8646e711 · inbound

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? cites this paper.

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:40:56.029917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T05:40:55.944288Z digest=sha256:ff7e36c2603f6ccbf459335ec5e0a8267b55b307f28781642f6d5fb2d2ca145c

Observation 292307aa-50aa-4b8e-ad88-d44a0372db00 · inbound

Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames cites this paper.

Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:17.956196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:17.956196Z digest=sha256:1415e5adec45b3f3e53a3c93aef728fbdb6fe111cb23fe1f1c4e49745fd9aecc

Observation d04d06b1-776c-47a2-be9f-0e489c873a9d · inbound

Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning cites this paper.

Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:14.722576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:14.722576Z digest=sha256:e01dbdee03f9fc5b8a16cfddc46c8c0f464f5af649d5c35f4944cf136f26b124

Observation 900592f3-cb75-485e-8f6d-57cc9329113b · inbound

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning cites this paper.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.114508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.114508Z digest=sha256:8ea592ef6841910e6e07574606679e1a80005e3eb2911d5a0bbf8b6383caf848

Observation 1d5a5c88-f935-41fe-ae91-41cd19f4e1c3 · inbound

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning cites this paper.

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:52:08.116778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T06:50:02.607136Z digest=sha256:de6742c5c16340fae16f19b88b0e4fff57d9e18b2d5f1d4de332386ee8a6f276

Observation d1911141-ba6d-45a0-9592-05121f0c56cf · inbound

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning cites this paper.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.103936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.103936Z digest=sha256:d6737522f9f762171d76962d441a2d3fc0b93e9e18d7a39a8a3ec3bbd5b954cf

Observation c3a40141-58cc-4308-b02d-db22c41b6f97 · inbound

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture cites this paper.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:22.714354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:22.714354Z digest=sha256:1348f556a781990262287add5c50f801e7ae0ba6bca17141d4a2bcace6c6131d

Observation 2a7ed2cc-0db1-417c-8b68-aae1f1a6ccf7 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:33.995648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:33.995648Z digest=sha256:9023e5af7546934abca481136d9ab734c1c63131ad34b3f193133444b59967da

Observation 85ebea19-bc89-4b8f-a73c-c97ba66a00b4 · inbound

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding cites this paper.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:05.079698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:05.079698Z digest=sha256:9e254926f238ed2280da0c4d53b91faf55f789a8eec8234540b0528a3850759c

Observation b4d6e639-687f-48bd-b459-b9d41f707997 · inbound

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs cites this paper.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:04.926420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:04.926420Z digest=sha256:78007d1e521acc6ae55e56961ca5ee3625bc4940b2506724e20276ce89e34759

Observation 87bc1d6f-4537-455f-bca5-9b40c5226ec0 · inbound

From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language Models cites this paper.

From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language Models Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T03:16:49.999803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:16:49.999803Z digest=sha256:4abb6b2627042d2bc9135c7caff777742c6870d088211f06e17c68f8898a8f21

Observation 873458e3-a56d-4bd4-a4d4-ed0ab4552919 · inbound

Deeper Thought, Weaker Aim: Understanding and Mitigating Perceptual Impairment during Reasoning in Multimodal Large Language Models cites this paper.

Deeper Thought, Weaker Aim: Understanding and Mitigating Perceptual Impairment during Reasoning in Multimodal Large Language Models Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-21T12:00:04.278492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T11:56:29.743281Z digest=sha256:f46fab97b0ed6c0196afdc463eee92ee977c6c06447f5f373b6653cab14e1eda

Observation 28551100-a395-4bf3-8a26-bb319e01047b · inbound

Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding cites this paper.

Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T17:54:05.713076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:54:05.713076Z digest=sha256:fb7686e6e301b0e509df70351a4ffe58ecb1782ffea66853bc0c74d13cf24bbd

Observation a1dff7ec-a172-4dc1-b672-2ab691354125 · inbound

EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs cites this paper.

EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:43:22.842223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T22:41:09.840792Z digest=sha256:30f7902db2a318887b4692895f77bb6ad99dc9e572656dab775f3328a43704c6

Observation d8ac2c06-1c86-44a5-9652-b424b447625c · inbound

EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs cites this paper.

EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-13T14:39:55.177552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T14:39:55.177552Z digest=sha256:43e2ae4b19a8bea2d5ca3fec79b97b152c46413cc8a464caa71c5b569a62bbb2

Observation 3548575d-a02f-41b1-b7c1-3d3883a46ef6 · inbound

Enhancing MLLM Spatial Understanding via Active 3D Scene Exploration for Multi-Perspective Reasoning cites this paper.

Enhancing MLLM Spatial Understanding via Active 3D Scene Exploration for Multi-Perspective Reasoning Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:15:47.770329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T19:16:46.753641Z digest=sha256:2d762e02ae333b1b4b5c847d4c8c34a14f575b4ce7b79b8d2c283e3cdbd70747

Observation 68324a95-075f-4510-afde-f682474e04b7 · inbound

Robotic Manipulation is Vision-to-Geometry Mapping ($f(v) \rightarrow G$): Vision-Geometry Backbones over Language and Video Models cites this paper.

Robotic Manipulation is Vision-to-Geometry Mapping ($f(v) \rightarrow G$): Vision-Geometry Backbones over Language and Video Models Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:46:01.362635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:20:47.418874Z digest=sha256:7ea8758857cb468797e387da4a693876bc13ece9ab23077bff98defb2ba9167c

Observation 75226c71-ce97-427c-87db-a0944021055f · inbound

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models cites this paper.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:36:08.748348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:5f559fe002a783b10904d007c495876c563fb596e713edb8a7212cc2b0eec50e

Observation 1ec4db68-218b-4ec9-968b-c9b9a5e89eb6 · inbound

CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models cites this paper.

CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T05:28:04.514201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T05:27:30.938311Z digest=sha256:5bb7dfbc980bd66c014585911c35bb0e49f35acd8b00351935173a5efae21a78

Observation ff9b5c72-0230-4538-ab8f-ae71625a189f · inbound

Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration? cites this paper.

Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration? Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:12:24.798806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T17:08:59.111306Z digest=sha256:798f441ce8d167136ba5ac6beb14248f4808d08b269b65794e8ae08bc6e247d3

Observation b5b5edfe-f574-464e-9f75-e5a3662b22e0 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 219

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.583697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:12218d1fdf4142cf71d478445e9d8ee23e9e8997ca02a68e25c395c6c6ea8c57

Observation 83a601e4-40bf-4523-8d9a-22308548fa91 · inbound

Reason, Then Re-reason: Cross-view Revisiting Improves Spatial Reasoning cites this paper.

Reason, Then Re-reason: Cross-view Revisiting Improves Spatial Reasoning Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:47:50.499205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T10:41:41.697216Z digest=sha256:c0b7be4c44e6033403ced86fec061f7053956138da4ca32cf045d7379ef05df8

Observation 11d33447-b8ce-4e57-bc5e-7f3061ada16f · inbound

SpaceEra++: A Unified Framework Towards 3D Spatial Reasoning in Video cites this paper.

SpaceEra++: A Unified Framework Towards 3D Spatial Reasoning in Video Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:18:37.300491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-03T16:16:41.412451Z digest=sha256:0ba134c4d9fa58b715ab480f30111c0fe84b787b7419a1c992db26646e56b938

Observation 8c4a1fa3-70d6-499e-b42b-d96f56b54eb2 · inbound

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding cites this paper.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.979540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.979540Z digest=sha256:edf1988747e8f3e4b431356b576cc4d82f3de0eb703d132e42d1ff2e7af6a72a