Pith. sign in

Paper Citation Record · LEDGER

GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2503.10639.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.10639 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:18:32.501235Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T21:28:58.240235Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 43cde051-d7e4-4926-8375-5a61830f08ec · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 255

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T17:18:53.577665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:93e354d94ebf2bf41392385f65f6dd5404894073ed76e0838feea6217534e344

Observation 05e16346-56ee-4424-8657-735ad92e1cc8 · inbound

Step1X-Edit: A Practical Framework for General Image Editing cites this paper.

Step1X-Edit: A Practical Framework for General Image Editing GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:36:41.949615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T14:36:41.467429Z digest=sha256:521cb22f690a57f82b8aeed0772df012f86af8ff298df6a1a424882301b98e6a

Observation db554c65-c3ad-45c8-9371-8fe68adb408c · inbound

ImgEdit: A Unified Image Editing Dataset and Benchmark cites this paper.

ImgEdit: A Unified Image Editing Dataset and Benchmark GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:17:45.508850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:17:45.123690Z digest=sha256:e729f1345c9172e4d9b56b579b35e6803ab47f635bdef1beddbe2edbaf0cec40

Observation f9e05ca0-be08-4147-969c-b18a3d36e5a5 · inbound

Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation cites this paper.

Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:18:32.501235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:18:32.501235Z digest=sha256:d7d577e9c3da55f813b002fd369ab4a0ee2f09150cac43972bf6fab38b837900

Observation e85fb9d6-0103-4439-b9cc-e33646806b97 · inbound

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics cites this paper.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.165875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.165875Z digest=sha256:590465648b5845ed63a84f648851ceda730f2459d3f1369fb96f22de8373aacc

Observation 74ba5345-9cb4-4895-ac99-6692f515617b · inbound

ComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional Dependencies cites this paper.

ComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional Dependencies GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:19.605957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:19.605957Z digest=sha256:5a150da58d031e9c44d6518c81d9875f85fc67cfae011656015a1165b993be3b

Observation 0feb2e14-1f15-4d9d-b0ec-05c9ae66f644 · inbound

Via Score to Performance: Efficient Human-Controllable Long Song Generation with Bar-Level Symbolic Notation cites this paper.

Via Score to Performance: Efficient Human-Controllable Long Song Generation with Bar-Level Symbolic Notation GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T05:40:56.881671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:40:56.881671Z digest=sha256:0f39b8b94c9275d8e256a1c21ec6f0e9417ac04beab8921f6904ae709c36150e

Observation 5be1d564-b5ec-4ea8-b747-5d355dacab17 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 256

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.132193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:45546d708cb37aef23002174f24a0be95393380ed6b929a3efdbbbd4428a4b6b

Observation 86739b7b-e6d6-4ebd-a99e-dd28b5e1bf17 · inbound

MEPG:Multi-Expert Planning and Generation for Compositionally-Rich Image Generation cites this paper.

MEPG:Multi-Expert Planning and Generation for Compositionally-Rich Image Generation GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T10:24:55.232863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:24:55.232863Z digest=sha256:111ec19e90d3d2b1bdf74cdf4ff18614176f3bd6d7249ba2e5f27cb1ec377751

Observation 9f56752a-f678-4670-b871-473500b61ff1 · inbound

Interleaving Reasoning for Better Text-to-Image Generation cites this paper.

Interleaving Reasoning for Better Text-to-Image Generation GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.877594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.877594Z digest=sha256:c537e9176500c250db7bafc8ddc93357938d5ee8673c5098e52519647a89b9e3

Observation 1ad52245-d29a-460c-bd5d-77ca3b5b34fb · inbound

Reconstruction Alignment Improves Unified Multimodal Models cites this paper.

Reconstruction Alignment Improves Unified Multimodal Models GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T22:36:07.936009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:36:07.936009Z digest=sha256:6017015eb9ab43fddeef4882c86c12996e7e06f45cc6ed2017f6869e5966cabd

Observation 15feb0bb-9423-4260-a030-ccd50ba0b956 · inbound

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark cites this paper.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.645811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.645811Z digest=sha256:d5e5a77cfe25996786abad761700d91c5eacce57a25cd968504e68e05d882bd5

Observation c317703b-34b2-4365-81ba-33f4da95b3ef · inbound

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models cites this paper.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:24.239016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:24.239016Z digest=sha256:e410683e614e9c6e357b79b3529fdf1b5b7461b3b63508b5dce7407a17a03b28

Observation b5450353-038d-4dac-a264-4d7fa7247780 · inbound

IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation cites this paper.

IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T17:06:54.722954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:06:54.722954Z digest=sha256:5e97aef4e0f4e0e43ad39219e2c7ba39eb3755f5e27a357231a05e08c3edc595

Observation f8b3094d-71f9-4f13-968f-a51d8674b083 · inbound

Do-Undo Bench: Reversibility for Action Understanding in Image Generation cites this paper.

Do-Undo Bench: Reversibility for Action Understanding in Image Generation GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:11:18.473960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T22:10:29.304091Z digest=sha256:8b64a6d2e858c7a8d561d185efe7e002fcbc58b1586c6b5c11aac9474446fc4a

Observation b4f0b072-f1dc-495e-9848-97a5fb37d484 · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-13T23:27:11.006580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:27:11.006580Z digest=sha256:6bbb4412dd2b172d0d53f46e19937fab986247a38d52dbba585f3f2d569b4e76

Observation 3766ccda-d2e8-4b89-a5ea-a612eab982a5 · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:53.930732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:53.930732Z digest=sha256:8f4b2fc36eb6a096890c186f1202ce06f8b8d20cfb73921736cec6c9747722c6

Observation 14adf7f1-dbb3-474f-be97-4dc5bf312958 · inbound

AIM-Bench: Benchmarking and Improving Affective Image Manipulation via Fine-Grained Hierarchical Control cites this paper.

AIM-Bench: Benchmarking and Improving Affective Image Manipulation via Fine-Grained Hierarchical Control GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:11:04.280660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:35:49.575589Z digest=sha256:9f9a9a9341cf4a8549f85a1fb21f82418ff7199c577e1a49d7ca6f39dd2a3b3f

Observation 62b9b7dc-027b-415e-9552-c6643be237f4 · inbound

Edit Where You Mean: Region-Aware Adapter Injection for Mask-Free Local Image Editing cites this paper.

Edit Where You Mean: Region-Aware Adapter Injection for Mask-Free Local Image Editing GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:06:13.722077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T06:47:22.451132Z digest=sha256:858d2d3dee668b29ea14f6c5e04f1bfddc28ec07f0a57e61ff48d1f00df156fd

Observation c251ec09-3171-4a66-8787-bfef63b28c40 · inbound

Meta-CoT: Enhancing Granularity and Generalization in Image Editing cites this paper.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:19.274947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:8f3fcf2bc85cd5192e9c8cf97f0d520dbc0634232ae85e4ec5f02456864e08df

Observation 2cbfaf2e-1b78-4469-9421-422f690c10d1 · inbound

RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation cites this paper.

RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:31:00.029682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:16:25.031349Z digest=sha256:22bb8aae81c8fa3550695ef7b699d35450c3b5f7ae0d5ebca2c34b6169bcdf89

Observation af9e71d7-eece-42b0-b3e3-f446fe819732 · inbound

Masked Generative Transformer Is What You Need for Image Editing cites this paper.

Masked Generative Transformer Is What You Need for Image Editing GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:06:26.027511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:35:52.355925Z digest=sha256:0212194f2fca6dfc6a03f3b5c3497624f5d600c24d45614fdcdb08666630377e

Observation 2eb83584-ced1-4f53-b90d-a9e56d3d8c6c · inbound

UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning cites this paper.

UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:47:04.960167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:29:23.774196Z digest=sha256:1d1706fb63188d3254315e1f4b6a975fb08e1688e897c43260ecc07d8a3b939c

Observation 65c4a275-426e-46cd-b7ad-8b45e09534c1 · inbound

Editor's Choice: Evaluating Abstract Intent in Image Editing through Atomic Entity Analysis cites this paper.

Editor's Choice: Evaluating Abstract Intent in Image Editing through Atomic Entity Analysis GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:25:46.043811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T21:41:10.852265Z digest=sha256:0bfd7295f1b2d49547ba16d6c9647c2955bbae88126c3490624c0286088700eb

Observation a57079f0-abd8-431f-ac79-e66a5a1ac02e · inbound

Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning cites this paper.

Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:18:05.182007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T06:16:47.650748Z digest=sha256:80717116dd023894669557afeaf0b347c7811a6e297a83fd1e7f465d00d47102

Observation 14463bd6-eece-4c48-8105-59f7a7b1fb00 · inbound

Evaluating Reasoning Fidelity in Visual Text Generation cites this paper.

Evaluating Reasoning Fidelity in Visual Text Generation GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:16:44.516479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T07:03:38.967856Z digest=sha256:de27e392809073bb0d9ebf784582d5513425da5b1f8c4ef2c97b43adcf360e1e

Observation 814bee08-d142-41f8-bbd4-71e59b4c17f1 · inbound

MetaPoint: Unlocking Precise Spatial Control in Agentic Visual Generation cites this paper.

MetaPoint: Unlocking Precise Spatial Control in Agentic Visual Generation GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T08:26:47.752091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T06:07:06.056441Z digest=sha256:f381e2b116669e4f9c71382edd4d81b0c92dccdc7692006c5bfb14a36d349cc8

Observation 364406dd-a65b-4b65-8622-0a880d401306 · inbound

GMO-E$^2$DIT: Grounded Multi-Operation Editing for E-Commerce Images cites this paper.

GMO-E$^2$DIT: Grounded Multi-Operation Editing for E-Commerce Images GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:56:58.969221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T13:56:43.671622Z digest=sha256:f1588e58102378d0d043dede66d80e3fbe64785d2445ac374a37a73bd8af369c

Observation e6a8179a-8d93-479a-90dc-1b99efc7a669 · inbound

GMO-E$^2$DIT: Grounded Multi-Operation Editing for E-Commerce Images cites this paper.

GMO-E$^2$DIT: Grounded Multi-Operation Editing for E-Commerce Images GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:28:58.241767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T21:23:41.271521Z digest=sha256:811bb2fdba146bc579d0d0c13bd1b456a586353fb2e90baf1db82f7685feefaf

Observation cec7d7d6-c797-4468-a227-9ac89c904e7a · inbound

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning cites this paper.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:cc4e72bb846c55a83e693201db1da009979fb5a4644ea14b850c0f4524640ec5