Pith. sign in

Paper Citation Record · LEDGER

GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2505.17022.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17022 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T18:48:03.631140Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T16:39:58.241853Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1a05092a-b4a2-4775-b79c-f58b6f9084fc · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning

Reference 258

Resolution
verified exact
local_arxiv, observed 2026-05-18T19:21:48.230655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:7cf71221b85b4ba4e27fcd81e94dee2d6b20cee883787a54a7de1587f7ca94ef

Observation 937a204b-cf09-4159-8f3d-86a1fac5637c · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning

Reference 122

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:02:25.313296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:d9806ad638156c916fe0cf01a1a4c1bf050ecf21f5a8eac9bff76fddcad225ec

Observation 5aa4c4a5-d4bd-4dba-a56a-700e5605e989 · inbound

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark cites this paper.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.631140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.631140Z digest=sha256:44493710243f933e499d1e959185137c1322603dbbb599e29bb9bd6fca450f0c

Observation caa6bf72-8cd2-47eb-9a59-4c222b40aa02 · inbound

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models cites this paper.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:24.184568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:24.184568Z digest=sha256:081c1d3a186452939044ccf54edbdced375d1ac3bd2ec67c31300ecb2f4702fc

Observation 00b19032-37ed-41a1-916d-f871d6cedbff · inbound

PSR: Scaling Multi-Subject Personalized Image Generation with Pairwise Subject-Consistency Rewards cites this paper.

PSR: Scaling Multi-Subject Personalized Image Generation with Pairwise Subject-Consistency Rewards GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:51:29.372261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T03:49:05.489626Z digest=sha256:22b781c4ddf40cb51ea485db36d63d82cb59e2c32ed43d0993f8682ec2eab60e

Observation db44c475-71a3-401b-b341-fb815a499ceb · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T23:27:11.006580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:27:11.006580Z digest=sha256:2583e24cb03a988e36bf1e64e78046dde2ae716459cd606e24d6a052e62de726

Observation 516faf33-70ee-4909-9683-b156ffc9db19 · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:53.733424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:53.733424Z digest=sha256:885f7e6140afafeacc9973f457db5ed043f6b4bfac298f95e58828e22633fce4

Observation 6dac6061-0d13-46e9-829c-c94867eeecdb · inbound

Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward cites this paper.

Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:20:47.844523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:59:19.379119Z digest=sha256:8b908ca15d4e788d6e87b5894c57dbbb25546fd37d94a8a18d4a84874b825dc2

Observation e500cf13-84a6-4c94-95e4-5c45e0173aef · inbound

Meta-CoT: Enhancing Granularity and Generalization in Image Editing cites this paper.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:46:13.279370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:c49ebc4533a15c72fe3af241fa152e22766a60d444ab51bc51f39a0ac89976d9

Observation 1ee22c8d-8ecd-48d3-a3ee-db5e6837eedf · inbound

Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models cites this paper.

Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T07:22:28.450807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T07:21:40.803291Z digest=sha256:4652c14288fc2cb178183afa402927a10eb7fb16be8032788faba5bf11fcc9a9

Observation edd9b6e7-bb74-4838-8f0c-001b6a967dac · inbound

Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models cites this paper.

Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T16:52:40.185513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T16:48:20.132240Z digest=sha256:0625d744d86bd3977de8ae45ea342f37e6b5f681313183e233afa7ebe240a12b

Observation 1dea298d-778d-4428-9fad-dbbc972204a5 · inbound

Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models cites this paper.

Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T08:04:02.662823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-21T08:02:00.980836Z digest=sha256:41463702e2ff67228ab52515efcc90049f77f8ca1bcd9f60fdecf4547d163d8d

Observation f32ca364-43fe-47c5-a928-79787bb578e5 · inbound

B-GRTO: Bootstrapped Group Relative Tool Optimization for Referring Segmentation cites this paper.

B-GRTO: Bootstrapped Group Relative Tool Optimization for Referring Segmentation GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-25T04:30:20.645237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T04:27:15.911342Z digest=sha256:485a447b01a04d0c96133f287e2062e5ea1a4cec2fdd3afef4be785f1229f335

Observation fd3c0cd2-1340-47a9-b03a-cc0eb5118814 · inbound

B-GRTO: Bootstrapped Group Relative Tool Optimization for Referring Segmentation cites this paper.

B-GRTO: Bootstrapped Group Relative Tool Optimization for Referring Segmentation GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-06-30T16:04:52.997914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T16:00:40.812912Z digest=sha256:b1e24872635f2a09c6adf66577175db3c9ab166456fb65b93c4f45818f8fcfd1

Observation 5e4a12c1-dc9f-4802-a659-453f0c34ffe7 · inbound

Dive into the Scene: Breaking the Perceptual Bottleneck in Vision-Language Decision Making via Focus Plan Generation cites this paper.

Dive into the Scene: Breaking the Perceptual Bottleneck in Vision-Language Decision Making via Focus Plan Generation GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T02:56:29.293936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T10:28:41.952330Z digest=sha256:6a544528fc3d7823cb0be0925a1f2de36bbc497bea17a47dfaada0fce72817af

Observation 47ed9e8d-6ae9-486a-b062-8ba11c7620fc · inbound

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation cites this paper.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T03:06:29.575568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:5298addc3510fe91cd48374e5d596192d17b213d66959adb3ee5003f65a24c2d

Observation d572bf79-1550-4926-aa5b-4abc6e3c5429 · inbound

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation cites this paper.

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T16:39:58.243105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T00:19:49.071495Z digest=sha256:0083122c0e75a97607385ef1cf9ed08adf702bf4c2287ca12d99140d8fc5225f