Pith. sign in

Paper Citation Record · LEDGER

Interleaving Reasoning for Better Text-to-Image Generation

As of 22 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 21 inbound Pith citation observations for arXiv:2509.06945.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.06945 v2

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T22:55:44.984216Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T04:44:28.524092Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:39:58.253188Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5124f506-35d6-405b-8e95-25bfe6394496 · outbound

This paper cites Qwen2.5-VL Technical Report.

Interleaving Reasoning for Better Text-to-Image Generation Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.859430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.859430Z digest=sha256:fe4bec42f561c3532489a9c5fad28d2a7abe9f2b1e0f5f798cf144b570083727

Observation 10576f2d-416c-403f-a149-1fc83f13bcee · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

Interleaving Reasoning for Better Text-to-Image Generation Emerging Properties in Unified Multimodal Pretraining

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.873401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.873401Z digest=sha256:f5281fad87fe0170c40fcdb8df24e8c118e8d27afa67cd1c36deb20c552b3e9d

Observation 9f56752a-f678-4670-b871-473500b61ff1 · outbound

This paper cites GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing.

Interleaving Reasoning for Better Text-to-Image Generation GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.877594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.877594Z digest=sha256:3d6f0f048dab722f39f9d0feb864f3f7e89f597724860ee18ba36ad409c242dc

Observation 6f0d6a40-9d29-4a19-8936-10f296ef8d0b · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

Interleaving Reasoning for Better Text-to-Image Generation SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.881929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.881929Z digest=sha256:d8405f31f84694075f872b740c6ca264b628c278bfa88434bdda94a045fddae3

Observation 44623535-7df1-419b-9f58-44f48e8dbe26 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Interleaving Reasoning for Better Text-to-Image Generation Classifier-Free Diffusion Guidance

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.890907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.890907Z digest=sha256:f94d3351656e8126cb6d76b32c71216f0d4a4d2e0dc4dadbe673e88eaada8775

Observation 5b9bcf84-7d04-4b07-b310-f5bb9393f428 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

Interleaving Reasoning for Better Text-to-Image Generation Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.894915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.894915Z digest=sha256:5f400fc9a7269a97094c07be76c52d9b4d9d0b88d9229efc485d50bd6cda426e

Observation 400b2981-4301-4a12-b296-a09abfdf8ee5 · outbound

This paper cites OpenAI o1 System Card.

Interleaving Reasoning for Better Text-to-Image Generation OpenAI o1 System Card

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.898925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.898925Z digest=sha256:f0b857e92dbe68ce275c4c0d2401f51746a5eede0db215f375cb8c89c3a8e711

Observation 90676a64-8725-46a9-ad1c-173bd31af4e9 · outbound

This paper cites T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT.

Interleaving Reasoning for Better Text-to-Image Generation T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.903011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.903011Z digest=sha256:335c3159570f79b08e040b210230202a7b77405cd439daaeefa28841c4257b76

Observation e913c196-2937-4c88-9adc-fc01578ed526 · outbound

This paper cites Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation.

Interleaving Reasoning for Better Text-to-Image Generation Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.907085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.907085Z digest=sha256:d5307b875e648cb1f7c2a5c5d74eba3794a5da671a4840fc5998a01b10923f2a

Observation c94664b9-3afc-4601-b1c1-1ee5a6dfc69f · outbound

This paper cites Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation.

Interleaving Reasoning for Better Text-to-Image Generation Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.911235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.911235Z digest=sha256:b9a2a01295796c35690974e72dd9a81ffec6bbfd3a52cd43f770858e02987e82

Observation 1ecae873-fb11-4561-9f18-6b0b8472fce7 · outbound

This paper cites UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation.

Interleaving Reasoning for Better Text-to-Image Generation UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.915056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.915056Z digest=sha256:ec9f53b49fb601ec7ed5754b14ce69aeb32ff218ff6e7e8e8b7d6f3c0379c357

Observation e958345e-03c1-4134-a54f-10d17e56ae3b · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

Interleaving Reasoning for Better Text-to-Image Generation World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.919199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.919199Z digest=sha256:b8307aa690a83bb9d077bcc9596c825ee2ea26e99b977c05d4d3f5b8dabdcf6a

Observation 143acc15-20e2-4963-bea2-b262de52529a · outbound

This paper cites Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321,.

Interleaving Reasoning for Better Text-to-Image Generation Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.923109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.923109Z digest=sha256:9ed6e5a9d254c3d98a1529a5aa203f0c052d94446bbb41e4076489af606252f9

Observation 1de9c5af-6f41-4e77-854c-14ae561f2ab8 · outbound

This paper cites JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation.

Interleaving Reasoning for Better Text-to-Image Generation JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.926786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.926786Z digest=sha256:d06f8328ad9c4bc1862bbd5536a7f1af88574b9b31e6f999d7ea17ef179e1016

Observation f8fef04f-988b-4924-b6c2-dcdc3e59eb9e · outbound

This paper cites WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation.

Interleaving Reasoning for Better Text-to-Image Generation WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.930599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.930599Z digest=sha256:35404370dc88f4c093c22fa82f5502d6389a9eef4226d5cad755e278c803a495

Observation 8417e21d-2b57-426d-bb17-b6ae00a836fd · outbound

This paper cites Transfer between Modalities with MetaQueries.

Interleaving Reasoning for Better Text-to-Image Generation Transfer between Modalities with MetaQueries

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.934415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.934415Z digest=sha256:9aa2c30c891fc722f5744bf5ae6217a1c99e55d3b51fe5d36c708bfbf87e5947

Observation aea23b05-9223-43c1-b718-afbfdddcb612 · outbound

This paper cites Lumina-Image 2.0: A Unified and Efficient Image Generative Framework.

Interleaving Reasoning for Better Text-to-Image Generation Lumina-Image 2.0: A Unified and Efficient Image Generative Framework

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.938727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.938727Z digest=sha256:a2733f1308be60c339eb1a01320fc7d1149a49a9d7ddede459b3d7a204602612

Observation b6a80fd0-984a-4d3a-b633-fccfc4c216a0 · outbound

This paper cites TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation.

Interleaving Reasoning for Better Text-to-Image Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.942757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.942757Z digest=sha256:a60cd2a2334508682e038fb502132231d531650bb0568cf3fb104b62175891d9

Observation 7c5b707e-5997-40d3-8563-10b5c8b89bce · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Interleaving Reasoning for Better Text-to-Image Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.946720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.946720Z digest=sha256:0149687fad9b04d29d9a337f83f940106b899171cb929396a80856d346d1bad7

Observation 6149699a-4a80-4225-a911-7bf60ff7540a · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Interleaving Reasoning for Better Text-to-Image Generation Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.954534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.954534Z digest=sha256:18c6498946195afea57366337f6fa7bcc9e1c61421b15fec2b0c7fe6f240c784

Observation 7f3efb74-57f0-4b11-ae1b-346a7a6f4f81 · outbound

This paper cites MetaMorph: Multimodal Understanding and Generation via Instruction Tuning.

Interleaving Reasoning for Better Text-to-Image Generation MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.958495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.958495Z digest=sha256:ffc475d7cb5cc94f768c1005216a2b301bea02a0e80d7ed1f524e674522b1a8d

Observation f7a83eec-5066-409a-87c3-31ec6c7fb86f · outbound

This paper cites ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance.

Interleaving Reasoning for Better Text-to-Image Generation ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.963071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.963071Z digest=sha256:75bf5c1cdb94d9d75abc1d1e001ed7bebd10a7f3de6f1a748358a4d7000e1a8f

Observation 84579299-7ced-436f-8ae8-b816cbe7fedd · outbound

This paper cites OmniGen2: Towards Instruction-Aligned Multimodal Generation.

Interleaving Reasoning for Better Text-to-Image Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.967227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.967227Z digest=sha256:74cfe192f10b7dc872ee7246759fe81e6c04575fabe967831255956b90b2b5ca

Observation 558594d0-d1f6-48f9-a683-2152cbed355b · outbound

This paper cites Show-o2: Improved Native Unified Multimodal Models.

Interleaving Reasoning for Better Text-to-Image Generation Show-o2: Improved Native Unified Multimodal Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.971860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.971860Z digest=sha256:7c84414ffcec356fb787278f9702e0b2815e2624c57a0c1b3849236f575425fc

Observation 573de087-0f8f-42a2-9344-7eb92521f4f1 · outbound

This paper cites Qwen3 Technical Report.

Interleaving Reasoning for Better Text-to-Image Generation Qwen3 Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.976360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.976360Z digest=sha256:52edc7d725f66a2f5c3c474285d8035177e16c9666504a43a4e26ddeee0447e3

Observation b87d6789-8fed-404b-bf6e-8d880ea1d689 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Interleaving Reasoning for Better Text-to-Image Generation Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.980217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.980217Z digest=sha256:d49a74a2a86557ebcc678ba040d081bc0811254c7476a3c46e1543ce77ae31b1

Observation cbe2b674-3fa5-42d5-8723-e92f31c46a4b · outbound

This paper cites From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning.

Interleaving Reasoning for Better Text-to-Image Generation From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.984216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.984216Z digest=sha256:44f94e43cd869cae37ef1a0fc4e40115a2c2b37db0ba14d43f013e2c1d4ca1b9

Observation 93875204-2588-4310-b563-7bab02756f4c · outbound

This paper cites LMFusion: Adapting Pretrained Language Models for Multimodal Generation.

Interleaving Reasoning for Better Text-to-Image Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.950531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.950531Z digest=sha256:540dd26570e659a6ba8bcd36de9da3906e3771a9b147b0f23dc6fdd960272aa4

Observation 31b420a8-9e27-48f7-be8e-059af6970582 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Interleaving Reasoning for Better Text-to-Image Generation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.886846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.886846Z digest=sha256:f454ec13c2cd4bc8957b4f33974b38baa68876f8655922887efbc46c073abc29

Observation 8235d2a8-716d-4e57-a318-1de3f1d2d672 · outbound

This paper cites Advancing multimodal reasoning: From optimized cold start to staged reinforcement learning.arXiv preprint arXiv:2506.04207, 2025b.

Interleaving Reasoning for Better Text-to-Image Generation Advancing multimodal reasoning: From optimized cold start to staged reinforcement learning.arXiv preprint arXiv:2506.04207, 2025b

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.868981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.868981Z digest=sha256:b6e8b58f0b3657bebe561065291a8b0a2b92997734b2b05a5a1d723debcce9ea

Observation c1c9beee-3b0d-41a7-af2f-2b5562734f09 · outbound

This paper cites OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation.

Interleaving Reasoning for Better Text-to-Image Generation OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.864389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.864389Z digest=sha256:9edd57395b9f0b85c6cb14a79231e70ea14f9d519dd0b8033c5d6560034a7a6d

Pith citing papers

Observation 7a4ba74c-4fae-4ac1-a9d8-8357d14b0a7a · inbound

IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation cites this paper.

IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation Interleaving Reasoning for Better Text-to-Image Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T17:06:55.166111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:06:55.166111Z digest=sha256:d88a15757980de79ff181bc4ffa717d6d85d63e8455de99446a248d3f675243a

Observation 685c24f6-8ff2-4100-87cb-c52c064ff8e9 · inbound

How RL Unlocks the Aha Moment in Geometric Interleaved Reasoning cites this paper.

How RL Unlocks the Aha Moment in Geometric Interleaved Reasoning Interleaving Reasoning for Better Text-to-Image Generation

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:20:14.441950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T18:16:25.372627Z digest=sha256:e727a4dde81f7641e0a9135f6742152f858ae8b3e9d77341a685f276131a3069

Observation b316f3e0-46ec-46b5-8fb3-fd08f3daa2ec · inbound

Learning Preference-Based Objectives from Clinical Narratives for Dynamic Sepsis Treatment cites this paper.

Learning Preference-Based Objectives from Clinical Narratives for Dynamic Sepsis Treatment Interleaving Reasoning for Better Text-to-Image Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-12T22:22:06.385856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:22:06.385856Z digest=sha256:3172070ff8db504de798185f3660dc25f04d8b5f050b9c722c435faba0f079cc

Observation c5f51071-fd0f-44b1-a625-9e7400416a2c · inbound

TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training cites this paper.

TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training Interleaving Reasoning for Better Text-to-Image Generation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:21:03.058943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T15:32:01.551829Z digest=sha256:99e3137f03b8dc5ef0dd9a844f0a020ad6b524653ae9b4c13ac628ad5869a160

Observation 3f383230-214a-4a49-86a1-679e9d4442f1 · inbound

TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training cites this paper.

TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training Interleaving Reasoning for Better Text-to-Image Generation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:59:55.292592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T08:59:10.877437Z digest=sha256:0cc7eee6c266b874ce78f6fd1242436920d078cc48f4f95611f188e5942798a0

Observation 78c1b218-e674-4b05-b64c-8811eeaa877c · inbound

Meta-CoT: Enhancing Granularity and Generalization in Image Editing cites this paper.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Interleaving Reasoning for Better Text-to-Image Generation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:19.288085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:dc8c775673f7380a7ae29dfb64c9f53abea2b5c29133d2bda315e822d1f88197

Observation b890564b-ba3d-47d4-89d5-74e044db0147 · inbound

Refinement via Regeneration: Enlarging Modification Space Boosts Image Refinement in Unified Multimodal Models cites this paper.

Refinement via Regeneration: Enlarging Modification Space Boosts Image Refinement in Unified Multimodal Models Interleaving Reasoning for Better Text-to-Image Generation

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:31:14.310799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-07T16:55:19.763050Z digest=sha256:56b4e1bfd3a07c8ac95dacac5bec777dcc103b7547d799fd923097c0d658a2ad

Observation d1ab0b94-b4e9-4270-915d-cad94ddabee2 · inbound

SCOPE: Structured Decomposition and Conditional Skill Orchestration for Complex Image Generation cites this paper.

SCOPE: Structured Decomposition and Conditional Skill Orchestration for Complex Image Generation Interleaving Reasoning for Better Text-to-Image Generation

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:30:57.539775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-11T02:27:36.104714Z digest=sha256:10bd175f0edb292d6ff459ecf8d5b9992845ae161cb5f0b2a90e176c5ed75e4a

Observation cd069034-a83f-49d8-b350-c40ee2a3e0bb · inbound

Flow-OPD: On-Policy Distillation for Flow Matching Models cites this paper.

Flow-OPD: On-Policy Distillation for Flow Matching Models Interleaving Reasoning for Better Text-to-Image Generation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:00:54.250991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T02:04:16.335479Z digest=sha256:86a6dd10d58a1f36036fae44ef0b71da8839c2627466638964449819dff2210c

Observation e91cf0d4-220b-4294-bbb2-e2c4dc308fb7 · inbound

Flow-OPD: On-Policy Distillation for Flow Matching Models cites this paper.

Flow-OPD: On-Policy Distillation for Flow Matching Models Interleaving Reasoning for Better Text-to-Image Generation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:27:02.795232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T01:17:08.955601Z digest=sha256:c1aa1ea17df6deab3a2c92b2c2d211a5983701195803376e7f1c3ee00282d3bf

Observation b968cffc-d6c6-4451-a5bb-9ad3dae66c79 · inbound

Flow-OPD: On-Policy Distillation for Flow Matching Models cites this paper.

Flow-OPD: On-Policy Distillation for Flow Matching Models Interleaving Reasoning for Better Text-to-Image Generation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:55:05.137604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T05:52:26.736304Z digest=sha256:61c1e4ad3aac537d5f3b045dfcb6b0b7f0b2cfc15f8ad8ec93470112e240f167

Observation d3b2a36d-e9d5-4c90-bf62-5f9a86decb77 · inbound

Flow-OPD: On-Policy Distillation for Flow Matching Models cites this paper.

Flow-OPD: On-Policy Distillation for Flow Matching Models Interleaving Reasoning for Better Text-to-Image Generation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:43:51.740816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T22:39:32.496303Z digest=sha256:968bb5da2b0a72c46f37af4b067e16e22aaf944fab5db7587d9ec2e330485355

Observation b10ec737-893f-4828-8850-e260f3b6d7e4 · inbound

Flow-OPD: On-Policy Distillation for Flow Matching Models cites this paper.

Flow-OPD: On-Policy Distillation for Flow Matching Models Interleaving Reasoning for Better Text-to-Image Generation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:05:07.211199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:02:29.120150Z digest=sha256:4c5a9f58b23429d259d19966e04003ce6b7125e444f2b48c3c9456dd3c72df35

Observation 2480ca98-713e-4083-877f-7bcb775e4451 · inbound

UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning cites this paper.

UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning Interleaving Reasoning for Better Text-to-Image Generation

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:47:04.938139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T01:29:23.774196Z digest=sha256:6d809b6ef5edd60e59bbce69df6c65acc032661daf02c9f3118369de8c1ac127

Observation 4d94f5ca-30ea-45a2-93af-68710ce439fb · inbound

Unlocking Complex Visual Generation via Closed-Loop Verified Reasoning cites this paper.

Unlocking Complex Visual Generation via Closed-Loop Verified Reasoning Interleaving Reasoning for Better Text-to-Image Generation

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:37:39.755072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-19T16:35:43.166697Z digest=sha256:e9e93b6e37d1c8769e291ffe6594c75f63dd8e009b6c0f8b40695fb93c0ebe47

Observation 28de1f0c-778d-4778-bf02-b61f1fa7d7d1 · inbound

LatentUMM: Dual Latent Alignment for Unified Multimodal Models cites this paper.

LatentUMM: Dual Latent Alignment for Unified Multimodal Models Interleaving Reasoning for Better Text-to-Image Generation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:43:17.447238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:39:28.058049Z digest=sha256:cdee207824f418592c089cc2cbfa15dc35d083dd533821911fa5e11d0973a86c

Observation 86c8d491-79c1-4d6b-89f9-afd5f2f656cf · inbound

OmniVerifier-M1: Multimodal Meta-Verifier with Explicit Structured Recalibration cites this paper.

OmniVerifier-M1: Multimodal Meta-Verifier with Explicit Structured Recalibration Interleaving Reasoning for Better Text-to-Image Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:03:25.978252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T13:01:08.529979Z digest=sha256:9b5dce72d53319fffc62b0690591c56ba466b8d763973e85ed52320f9449a3b0

Observation 2e4f6a9a-d3a6-417a-897c-15e969a7484e · inbound

Reason, Then Re-reason: Cross-view Revisiting Improves Spatial Reasoning cites this paper.

Reason, Then Re-reason: Cross-view Revisiting Improves Spatial Reasoning Interleaving Reasoning for Better Text-to-Image Generation

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:47:50.518611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T10:41:41.697216Z digest=sha256:1bcf078c2991ce589736a8bfcf5c52009ba1a5af36155130b7feac8acb640bdc

Observation 66a7c912-fa2b-4992-b99e-417e7cce5358 · inbound

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation cites this paper.

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation Interleaving Reasoning for Better Text-to-Image Generation

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:39:58.254509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-26T00:19:49.071495Z digest=sha256:ac1a430c30a1b39af5cd584ec8a8158319b255704d19600859ac1b6dfabf0699

Observation eee4709d-d147-47d7-81fa-b1206a14b403 · inbound

SimpleSearch-VL: A Simple Recipe for Multimodal Agentic Deep Search cites this paper.

SimpleSearch-VL: A Simple Recipe for Multimodal Agentic Deep Search Interleaving Reasoning for Better Text-to-Image Generation

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:55:41.095644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-01T06:02:48.532478Z digest=sha256:2072c66a2ab36f0d3337ad5c1594171249d638b6a1282f830e6bd0d83627341d

Observation 097f2bcd-f015-46ea-b082-5db26a881451 · inbound

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent cites this paper.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Interleaving Reasoning for Better Text-to-Image Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:28.524092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:28.524092Z digest=sha256:3bf3cfddcb6758064ad49ecf18c4727ff1a037c055d1cfa8355c7c2d967176fe