Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2401.11708.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:38:56.249658Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-29T13:23:28.257830Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 5f5cdb6a-875b-468f-8208-44001f9fbd15 · inbound
ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d2c356f3-c195-426d-8230-ba8200344824 · inbound
Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d01e336e-499e-4ebe-9ab8-1faeee6f2db4 · inbound
Image Regeneration: Evaluating Text-to-Image Model via Generating Identical Image with Multimodal Large Language Models Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7a4ae6b-1162-475b-97a1-3215af1bbe90 · inbound
MoTrans: Customized Motion Transfer with Text-driven Video Diffusion Models Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d18b211b-9d12-46c5-9a1e-15169c08dc39 · inbound
GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d42e04f-a0b0-41d2-adb3-fabe87a9bd6c · inbound
DyST-XL: Dynamic Layout Planning and Content Control for Compositional Text-to-Video Generation Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2dc2bf1-886b-4bd2-a3e3-16813d97853e · inbound
Step1X-Edit: A Practical Framework for General Image Editing Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b3b398bc-8821-45dd-bc29-fee3c6833b06 · inbound
Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71ba07ae-22b9-49a5-809f-62d0ccf92dca · inbound
Rhetorical Text-to-Image Generation via Two-layer Diffusion Policy Optimization Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc27c992-506a-44c4-8192-8bdb4f0a89db · inbound
Rethinking Cross-Modal Interaction in Multimodal Diffusion Transformers Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff379478-dc49-45a9-9c2d-0c34e5ca44a1 · inbound
Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 276680a2-f601-4dcf-bb86-2f6e84230380 · inbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation fc73f304-751b-4432-b8c0-7f4ee8713757 · inbound
Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs
Reference 232
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 529408f0-8e0a-4e04-84c2-4757a6f8e4cd · inbound
TILT: Improving Compositional Generation in Diffusion Models with a Model-Intrinsic Reward Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.