Pith. sign in

Paper Citation Record · LEDGER

Interleaving Reasoning for Better Text-to-Image Generation

As of 6 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 21 inbound Pith citation observations for arXiv:2509.06945.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.06945 v2

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T22:55:44.984216Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T04:44:28.524092Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:39:58.253188Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5124f506-35d6-405b-8e95-25bfe6394496 · outbound

This paper cites Qwen2.5-VL Technical Report.

Interleaving Reasoning for Better Text-to-Image Generation Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.859430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.859430Z digest=sha256:258eb38cb1432d55130264183eb0a092197152655c2c0328cd6dfed8f6814d77

Observation 10576f2d-416c-403f-a149-1fc83f13bcee · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

Interleaving Reasoning for Better Text-to-Image Generation Emerging Properties in Unified Multimodal Pretraining

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.873401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.873401Z digest=sha256:50b66c2f109f91826b802086f93b59e87dca5760dac8efc181714496b162ca62

Observation 9f56752a-f678-4670-b871-473500b61ff1 · outbound

This paper cites GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing.

Interleaving Reasoning for Better Text-to-Image Generation GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.877594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.877594Z digest=sha256:c537e9176500c250db7bafc8ddc93357938d5ee8673c5098e52519647a89b9e3

Observation 6f0d6a40-9d29-4a19-8936-10f296ef8d0b · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

Interleaving Reasoning for Better Text-to-Image Generation SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.881929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.881929Z digest=sha256:381cf3248dfec8f6dd90682ba20e9d26e3c83912f56c965ac285c6fff0dc3b49

Observation 44623535-7df1-419b-9f58-44f48e8dbe26 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Interleaving Reasoning for Better Text-to-Image Generation Classifier-Free Diffusion Guidance

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.890907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.890907Z digest=sha256:1971309283b85d078c1e4456d763603ddd1e09f472e5e8a905162fdf65d70762

Observation 5b9bcf84-7d04-4b07-b310-f5bb9393f428 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

Interleaving Reasoning for Better Text-to-Image Generation Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.894915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.894915Z digest=sha256:85e0dfb965dbbe3d44b6150e838172477e40c7c91ccfb95c15f7e3c3fab98160

Observation 400b2981-4301-4a12-b296-a09abfdf8ee5 · outbound

This paper cites OpenAI o1 System Card.

Interleaving Reasoning for Better Text-to-Image Generation OpenAI o1 System Card

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.898925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.898925Z digest=sha256:3b2073d7c448b51c6c966ad2e3ca7bd80964c1d95b0dd1b03070b471c8e31529

Observation 90676a64-8725-46a9-ad1c-173bd31af4e9 · outbound

This paper cites T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT.

Interleaving Reasoning for Better Text-to-Image Generation T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.903011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.903011Z digest=sha256:24c416ef35887570f5e759fb40cd86ef86d2edc5bcdf7977ece9c764267f27bb

Observation e913c196-2937-4c88-9adc-fc01578ed526 · outbound

This paper cites Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation.

Interleaving Reasoning for Better Text-to-Image Generation Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.907085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.907085Z digest=sha256:35a9153e207933882e027e7411d282264b6d546c9aab7e87682e311552c28dcf

Observation c94664b9-3afc-4601-b1c1-1ee5a6dfc69f · outbound

This paper cites Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation.

Interleaving Reasoning for Better Text-to-Image Generation Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.911235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.911235Z digest=sha256:2eba21073c901a4ed90da6ba1be6204708504d3fd40694bedcdadfea13d6edad

Observation 1ecae873-fb11-4561-9f18-6b0b8472fce7 · outbound

This paper cites UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation.

Interleaving Reasoning for Better Text-to-Image Generation UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.915056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.915056Z digest=sha256:ca9dbbc1ae1375fc2653e3fa91c05bd57a53f6389ce2a8d3f368bc0052026c09

Observation e958345e-03c1-4134-a54f-10d17e56ae3b · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

Interleaving Reasoning for Better Text-to-Image Generation World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.919199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.919199Z digest=sha256:3051bcf8c8e1144946279d5b9c21b088d2b8846e783443efc19554d551549bf6

Observation 143acc15-20e2-4963-bea2-b262de52529a · outbound

This paper cites Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321,.

Interleaving Reasoning for Better Text-to-Image Generation Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.923109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.923109Z digest=sha256:9d9451ab532f46d9fac8642f2908ca1f7852532f56cd043ed6e1b5f11ccb1500

Observation 1de9c5af-6f41-4e77-854c-14ae561f2ab8 · outbound

This paper cites JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation.

Interleaving Reasoning for Better Text-to-Image Generation JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.926786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.926786Z digest=sha256:35c9386f443dcb25feda8e154a850b63d68032a4092a36b71406b06115cafc02

Observation f8fef04f-988b-4924-b6c2-dcdc3e59eb9e · outbound

This paper cites WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation.

Interleaving Reasoning for Better Text-to-Image Generation WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.930599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.930599Z digest=sha256:f856c8e2f713b438ab137a25addadf50f14144c905c92a9ed9bc5d4f1da82f4f

Observation 8417e21d-2b57-426d-bb17-b6ae00a836fd · outbound

This paper cites Transfer between Modalities with MetaQueries.

Interleaving Reasoning for Better Text-to-Image Generation Transfer between Modalities with MetaQueries

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.934415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.934415Z digest=sha256:f1d19c63a25f69dea2466a81d9aa5d65a1a55983f8459a618b947536f7c45394

Observation aea23b05-9223-43c1-b718-afbfdddcb612 · outbound

This paper cites Lumina-Image 2.0: A Unified and Efficient Image Generative Framework.

Interleaving Reasoning for Better Text-to-Image Generation Lumina-Image 2.0: A Unified and Efficient Image Generative Framework

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.938727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.938727Z digest=sha256:8d0643266790c5634a425c25dc0f87116f061a5142a4d2ae5fa117aba50178c0

Observation b6a80fd0-984a-4d3a-b633-fccfc4c216a0 · outbound

This paper cites TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation.

Interleaving Reasoning for Better Text-to-Image Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.942757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.942757Z digest=sha256:c862c7fffc6d7856d3dfb997ac67968c55eb11b32df3e15aebb073a574c0b25c

Observation 7c5b707e-5997-40d3-8563-10b5c8b89bce · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Interleaving Reasoning for Better Text-to-Image Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.946720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.946720Z digest=sha256:49110c1408de51450f009bcfeacf67339fcbc496dc7cff0850bbd02807a00b50

Observation 6149699a-4a80-4225-a911-7bf60ff7540a · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Interleaving Reasoning for Better Text-to-Image Generation Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.954534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.954534Z digest=sha256:ff69e858786211d54f639a2d618a8b49358db03b5f55273ff13bd0658a07ddf8

Observation 7f3efb74-57f0-4b11-ae1b-346a7a6f4f81 · outbound

This paper cites MetaMorph: Multimodal Understanding and Generation via Instruction Tuning.

Interleaving Reasoning for Better Text-to-Image Generation MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.958495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.958495Z digest=sha256:f39f6c960c9d237b1c5020199d6acb1cfa2ee35d889d6771b0cd1f971dec9d07

Observation f7a83eec-5066-409a-87c3-31ec6c7fb86f · outbound

This paper cites ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance.

Interleaving Reasoning for Better Text-to-Image Generation ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.963071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.963071Z digest=sha256:daa5bf9372d2d3823a7dcaa8915a1a6da209275cda85f024bfd2a029f3c79e62

Observation 84579299-7ced-436f-8ae8-b816cbe7fedd · outbound

This paper cites OmniGen2: Towards Instruction-Aligned Multimodal Generation.

Interleaving Reasoning for Better Text-to-Image Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.967227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.967227Z digest=sha256:3110e0b30bcb04948673f04e3c41b09e210e6ebd5333d3a9441575a08d35a0c5

Observation 558594d0-d1f6-48f9-a683-2152cbed355b · outbound

This paper cites Show-o2: Improved Native Unified Multimodal Models.

Interleaving Reasoning for Better Text-to-Image Generation Show-o2: Improved Native Unified Multimodal Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.971860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.971860Z digest=sha256:033e971e41e0cb50c3dbe3c1e2c4161e9b4e6f309103804f076f660b6adfabde

Observation 573de087-0f8f-42a2-9344-7eb92521f4f1 · outbound

This paper cites Qwen3 Technical Report.

Interleaving Reasoning for Better Text-to-Image Generation Qwen3 Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.976360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.976360Z digest=sha256:ed0d20278e843c3601bf41b53c830df0386556bb4b9f64b72b4c1adf067ad421

Observation b87d6789-8fed-404b-bf6e-8d880ea1d689 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Interleaving Reasoning for Better Text-to-Image Generation Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.980217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.980217Z digest=sha256:0db161ab953cacbdfe237ab4d4506efa5483975d3b1f0aab07fb4af603fb7484

Observation cbe2b674-3fa5-42d5-8723-e92f31c46a4b · outbound

This paper cites From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning.

Interleaving Reasoning for Better Text-to-Image Generation From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.984216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.984216Z digest=sha256:e265ed22c7d7ed0a476e5b0480f68171beef24c2d11e2804797ecd623e72ce3e

Observation 93875204-2588-4310-b563-7bab02756f4c · outbound

This paper cites LMFusion: Adapting Pretrained Language Models for Multimodal Generation.

Interleaving Reasoning for Better Text-to-Image Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.950531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.950531Z digest=sha256:dd798772dd51d3d557ed34e95f539cc97c7d94aa60c28c47fce930a729d22816

Observation 31b420a8-9e27-48f7-be8e-059af6970582 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Interleaving Reasoning for Better Text-to-Image Generation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.886846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.886846Z digest=sha256:c63dd553458d3490b7cbe579a4c391e44faa1ecde592e7624500d6068583559b

Observation 8235d2a8-716d-4e57-a318-1de3f1d2d672 · outbound

This paper cites Advancing multimodal reasoning: From optimized cold start to staged reinforcement learning.arXiv preprint arXiv:2506.04207, 2025b.

Interleaving Reasoning for Better Text-to-Image Generation Advancing multimodal reasoning: From optimized cold start to staged reinforcement learning.arXiv preprint arXiv:2506.04207, 2025b

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.868981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.868981Z digest=sha256:d41ef2e9508607b02a355f4f41a73a1bebd812c7a631fa6309bc4b43dd520940

Observation c1c9beee-3b0d-41a7-af2f-2b5562734f09 · outbound

This paper cites OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation.

Interleaving Reasoning for Better Text-to-Image Generation OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.864389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.864389Z digest=sha256:bc2bb3cb0b7cb5574aa8f461702d9f92540c6459deca4b47d217b5677257fe62

Pith citing papers

Observation 7a4ba74c-4fae-4ac1-a9d8-8357d14b0a7a · inbound

IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation cites this paper.

IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation Interleaving Reasoning for Better Text-to-Image Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T17:06:55.166111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:06:55.166111Z digest=sha256:8e2802af84ebadf3e3f1e0942aec6d872666d89fbefeea6b6c6bf5f5baa011cd

Observation 685c24f6-8ff2-4100-87cb-c52c064ff8e9 · inbound

How RL Unlocks the Aha Moment in Geometric Interleaved Reasoning cites this paper.

How RL Unlocks the Aha Moment in Geometric Interleaved Reasoning Interleaving Reasoning for Better Text-to-Image Generation

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:20:14.441950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T18:16:25.372627Z digest=sha256:7d528494467a79fbb96dc80727735aa3c6248e44822a2eed3cf39bd8d5e73648

Observation b316f3e0-46ec-46b5-8fb3-fd08f3daa2ec · inbound

Learning Preference-Based Objectives from Clinical Narratives for Dynamic Sepsis Treatment cites this paper.

Learning Preference-Based Objectives from Clinical Narratives for Dynamic Sepsis Treatment Interleaving Reasoning for Better Text-to-Image Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-12T22:22:06.385856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:22:06.385856Z digest=sha256:6197597181c5abf258e496e458ebf47b4682594b2e05e0432da0758501b50ade

Observation c5f51071-fd0f-44b1-a625-9e7400416a2c · inbound

TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training cites this paper.

TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training Interleaving Reasoning for Better Text-to-Image Generation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:21:03.058943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:32:01.551829Z digest=sha256:16a83867c015138ec1c82a94fdc97fc4ee4570bb8b995d49825af655e125b43e

Observation 3f383230-214a-4a49-86a1-679e9d4442f1 · inbound

TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training cites this paper.

TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training Interleaving Reasoning for Better Text-to-Image Generation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:59:55.292592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T08:59:10.877437Z digest=sha256:7ebceffb87540c4051d4009fc1cffbe1ee99ffab1abf744487f7d0f28bba82c6

Observation 78c1b218-e674-4b05-b64c-8811eeaa877c · inbound

Meta-CoT: Enhancing Granularity and Generalization in Image Editing cites this paper.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Interleaving Reasoning for Better Text-to-Image Generation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:19.288085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:6dfa364bde573e7c354baf0839c55a440f6f3022146cb077d0361e31036fa160

Observation b890564b-ba3d-47d4-89d5-74e044db0147 · inbound

Refinement via Regeneration: Enlarging Modification Space Boosts Image Refinement in Unified Multimodal Models cites this paper.

Refinement via Regeneration: Enlarging Modification Space Boosts Image Refinement in Unified Multimodal Models Interleaving Reasoning for Better Text-to-Image Generation

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:31:14.310799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T16:55:19.763050Z digest=sha256:57fa7a8bb1dc08439dbffdc02b62e949d988913d18849f63006dc6560ae6c44a

Observation d1ab0b94-b4e9-4270-915d-cad94ddabee2 · inbound

SCOPE: Structured Decomposition and Conditional Skill Orchestration for Complex Image Generation cites this paper.

SCOPE: Structured Decomposition and Conditional Skill Orchestration for Complex Image Generation Interleaving Reasoning for Better Text-to-Image Generation

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:30:57.539775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-11T02:27:36.104714Z digest=sha256:faca0d1a8e9e5e4bd59ece3c9ef386eac0252a3c803d2ce1b03fae162f99ce9e

Observation cd069034-a83f-49d8-b350-c40ee2a3e0bb · inbound

Flow-OPD: On-Policy Distillation for Flow Matching Models cites this paper.

Flow-OPD: On-Policy Distillation for Flow Matching Models Interleaving Reasoning for Better Text-to-Image Generation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:00:54.250991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T02:04:16.335479Z digest=sha256:a27d3408b20219de50ff023f56c9226589d4f16b8d595e33e3b420f9f02ab18f

Observation e91cf0d4-220b-4294-bbb2-e2c4dc308fb7 · inbound

Flow-OPD: On-Policy Distillation for Flow Matching Models cites this paper.

Flow-OPD: On-Policy Distillation for Flow Matching Models Interleaving Reasoning for Better Text-to-Image Generation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:27:02.795232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T01:17:08.955601Z digest=sha256:8ac568f4ecedf40ad6bd3762a9504adfb83f17f8da02ccb0d84f0c97224a77dd

Observation b968cffc-d6c6-4451-a5bb-9ad3dae66c79 · inbound

Flow-OPD: On-Policy Distillation for Flow Matching Models cites this paper.

Flow-OPD: On-Policy Distillation for Flow Matching Models Interleaving Reasoning for Better Text-to-Image Generation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:55:05.137604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T05:52:26.736304Z digest=sha256:ac0376afadc3d2e7f1cefebeecd105b0eb60e0be041e07f615b2792c0673285c

Observation d3b2a36d-e9d5-4c90-bf62-5f9a86decb77 · inbound

Flow-OPD: On-Policy Distillation for Flow Matching Models cites this paper.

Flow-OPD: On-Policy Distillation for Flow Matching Models Interleaving Reasoning for Better Text-to-Image Generation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:43:51.740816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T22:39:32.496303Z digest=sha256:533911577912791eee8e8c496574b086fd583d54eb430f882f9b44f7d4bf3434

Observation b10ec737-893f-4828-8850-e260f3b6d7e4 · inbound

Flow-OPD: On-Policy Distillation for Flow Matching Models cites this paper.

Flow-OPD: On-Policy Distillation for Flow Matching Models Interleaving Reasoning for Better Text-to-Image Generation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:05:07.211199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T23:02:29.120150Z digest=sha256:94eea119e4fce11f1fa168955c2fab939a3f90906969b5b059b82b7afd0bc4b1

Observation 2480ca98-713e-4083-877f-7bcb775e4451 · inbound

UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning cites this paper.

UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning Interleaving Reasoning for Better Text-to-Image Generation

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:47:04.938139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T01:29:23.774196Z digest=sha256:f0ff7af4dd01508a6d26436811b399080b17544da682c98ba79af4bb73781f9a

Observation 4d94f5ca-30ea-45a2-93af-68710ce439fb · inbound

Unlocking Complex Visual Generation via Closed-Loop Verified Reasoning cites this paper.

Unlocking Complex Visual Generation via Closed-Loop Verified Reasoning Interleaving Reasoning for Better Text-to-Image Generation

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:37:39.755072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T16:35:43.166697Z digest=sha256:dc4bcaea52c06bc21d9ee1004af7073568f8967b5b2942dbfd9b9c2d4c55deef

Observation 28de1f0c-778d-4778-bf02-b61f1fa7d7d1 · inbound

LatentUMM: Dual Latent Alignment for Unified Multimodal Models cites this paper.

LatentUMM: Dual Latent Alignment for Unified Multimodal Models Interleaving Reasoning for Better Text-to-Image Generation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:43:17.447238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T12:39:28.058049Z digest=sha256:683e8d9357c9a300e958cbc13153c3b7bcef88723fd510f0688850d315bfda00

Observation 86c8d491-79c1-4d6b-89f9-afd5f2f656cf · inbound

OmniVerifier-M1: Multimodal Meta-Verifier with Explicit Structured Recalibration cites this paper.

OmniVerifier-M1: Multimodal Meta-Verifier with Explicit Structured Recalibration Interleaving Reasoning for Better Text-to-Image Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:03:25.978252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T13:01:08.529979Z digest=sha256:097283a9c7631f819fe278bf179eda858a45d699632e120bc5b1762d1f20c6be

Observation 2e4f6a9a-d3a6-417a-897c-15e969a7484e · inbound

Reason, Then Re-reason: Cross-view Revisiting Improves Spatial Reasoning cites this paper.

Reason, Then Re-reason: Cross-view Revisiting Improves Spatial Reasoning Interleaving Reasoning for Better Text-to-Image Generation

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:47:50.518611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T10:41:41.697216Z digest=sha256:d947ac23df84df63d07bf5590ccbc5c21f97b34c7f85a1ff90c2525733aa8d7a

Observation 66a7c912-fa2b-4992-b99e-417e7cce5358 · inbound

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation cites this paper.

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation Interleaving Reasoning for Better Text-to-Image Generation

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:39:58.254509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T00:19:49.071495Z digest=sha256:1da0a23d44ea90d2feb4d76ae6bf18381d479d7b2939237317df0122d9d6c1dc

Observation eee4709d-d147-47d7-81fa-b1206a14b403 · inbound

SimpleSearch-VL: A Simple Recipe for Multimodal Agentic Deep Search cites this paper.

SimpleSearch-VL: A Simple Recipe for Multimodal Agentic Deep Search Interleaving Reasoning for Better Text-to-Image Generation

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:55:41.095644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T06:02:48.532478Z digest=sha256:591590d90d6da31acab3eef522d6d8987008761c96981d74ff95251424ee6e61

Observation 097f2bcd-f015-46ea-b082-5db26a881451 · inbound

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent cites this paper.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Interleaving Reasoning for Better Text-to-Image Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:28.524092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:28.524092Z digest=sha256:8e51ec0337a5d5b7ffb128d962b4a6b1fc915d68f1c274b5606381cc2a98d0aa