Pith. sign in

Paper Citation Record · LEDGER

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2505.15804.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15804 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-13T16:55:20.099628Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.322058Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4cb3a065-2600-492c-a285-090f70fc77d3 · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Reference 297

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:36:24.521889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:a941df35fff0746d0aa107d96a5eaf4c3876bfc5d77c0c20dae19a125188040a

Observation cad82a6b-df4e-4c32-8a85-f002afae1123 · inbound

Video-R1: Reinforcing Video Reasoning in MLLMs cites this paper.

Video-R1: Reinforcing Video Reasoning in MLLMs STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:43:00.404330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T09:43:00.208065Z digest=sha256:64b9835f3bd91530c5dbfe0c1a124d947082d065dfb7e70929d1e90b532e7bb7

Observation 23c52071-9374-4517-9e4b-b5b62d315cf2 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Reference 233

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.322627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:5f8ea095a07bd30f808a7219268e684092a2efcf097d40890fd708ced30bf23a

Observation c0c118ee-ba2c-4943-91d0-197b451d7758 · inbound

VIDEOP2R: Video Understanding from Perception to Reasoning cites this paper.

VIDEOP2R: Video Understanding from Perception to Reasoning STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:25:22.690136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T22:24:41.760120Z digest=sha256:95e90241448a5c7c367065a5d8cc4323850170d169cf39807d1f168abfa13694

Observation c5ef8303-4565-4db1-a785-e21994e09dc6 · inbound

OneThinker: All-in-one Reasoning Model for Image and Video cites this paper.

OneThinker: All-in-one Reasoning Model for Image and Video STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:11:26.499729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T02:09:39.820651Z digest=sha256:942829807428210ce0919b6891a5f520224233af77d27051035793899fe80590

Observation 0b1a2ca8-b126-4ea0-a620-aba6a117bbea · inbound

CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning cites this paper.

CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:02:42.413324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T10:02:20.477517Z digest=sha256:ad734e696ce6d64a58f411e9cd9dc71e30891de06346d5a166df0d40ceff0403

Observation f5fe546e-47f0-4f33-b2df-f9ef43d8d2e8 · inbound

Fully Spiking Neural Networks with Target Awareness for Energy-Efficient UAV Tracking cites this paper.

Fully Spiking Neural Networks with Target Awareness for Energy-Efficient UAV Tracking STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-13T16:55:20.099628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:55:20.099628Z digest=sha256:7d724c8ab623f6bc008d07a0a871d4053d510d7ea83f0d5580bf651cb6f1001d

Observation e5b02838-3488-422e-b562-c080880d485d · inbound

Learning to Focus and Precise Cropping: A Reinforcement Learning Framework with Information Gaps and Grounding Loss for MLLMs cites this paper.

Learning to Focus and Precise Cropping: A Reinforcement Learning Framework with Information Gaps and Grounding Loss for MLLMs STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:38:00.985595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T21:35:12.859669Z digest=sha256:ffe3d7284392eaed24db2a4fc4b991007a872167f0036cd903729481e333e7c4

Observation 9ae53715-7a4f-4529-9a94-fbda2d4735b4 · inbound

Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning cites this paper.

Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:30:52.525638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:46:16.975267Z digest=sha256:38c10ff961f910eba4f97bb32f8bb2b7279983577d7609656ea57ba8571a2aeb

Observation dd904fd6-2380-4591-afc2-19ec70eb6245 · inbound

Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning cites this paper.

Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:35:52.573891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:40:41.642852Z digest=sha256:b7a5ee07367c1db096f4e43030c67c22eaa74a7c32fe9cc14e0c5de5948f7f73

Observation a1bcdc9a-4721-4290-a3e5-7842709a2325 · inbound

Beyond Semantic Search: Towards Referential Anchoring in Composed Image Retrieval cites this paper.

Beyond Semantic Search: Towards Referential Anchoring in Composed Image Retrieval STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:35:50.543718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:42:23.224746Z digest=sha256:ba9ee6c10e136c10c91e3a807c945a44fb570ea166ea94597a32d5d8b47ae5e5

Observation 76164de6-9435-4e6a-a1ef-2695f34cf3de · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Reference 192

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.323884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:e6007b3f5e0b5c00d2080e626080f443d0561ebb0a8afb5900c077f9b3d7c7b8