Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T12:39:59.330475Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 3 inbound Pith citation observations for arXiv:2601.02211.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T12:39:59.330475Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T03:50:45.652959Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T16:07:08.965746Z
53 of 53 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1acbafad-d7d0-44d0-86d2-01cf0551a5af · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Countgd: Multi-modal open-world counting
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf5cfc98-badc-4486-a651-145405c13c89 · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Stable flow: Vital layers for training-free image editing
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31b10555-9829-4064-b7a9-667ffd37b0dc · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Qwen2.5-VL Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df594be9-da2c-47fd-a711-fe65842fd27d · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Imagen 3.arXiv preprint arXiv:2408.07009, 2024
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98087cb2-d25d-4ef9-ad56-5da8a66a4665 · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers All are worth words: A vit backbone for diffusion models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bfe2fc2-c3c2-43ad-b2b6-7f7f16666086 · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Masactrl: Tuning-free mu- 9 tual self-attention control for consistent image synthesis and editing
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6580c660-7bbc-4b7a-889b-42aecfbbba6c · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Pixart-α: Fast training of diffusion trans- former for photorealistic text-to-image synthesis
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44e29ded-6dc0-494a-8a6b-16eb8fc7c0b7 · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Diffusion models beat gans on image synthesis
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33bd5ffc-84ae-444b-ba8a-da1eb66851e3 · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Scaling recti- fied flow transformers for high-resolution image synthesis
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f04187f-5873-46e9-bf57-3e09f4fd306d · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Seedream 3.0 Technical Report
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e71270c5-7abd-41ac-bf7b-0de65280f75a · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Geneval: An object-focused framework for evaluating text- to-image alignment
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01fc6551-66d3-4ebe-9067-d5a27f20c564 · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51d3f7af-ab58-4b1d-98af-23948b2d3f2a · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Prompt-to-Prompt Image Editing with Cross Attention Control
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 468c6688-74ad-47c3-af8b-39fbc9b9ad5d · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Classifier-Free Diffusion Guidance
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a499af7b-d295-4d50-81ef-5a13d4ca6916 · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Denoising dif- fusion probabilistic models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a0dc0d1-5a93-484d-8991-6d7cc8bba12d · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers T2i-compbench++: An enhanced and comprehensive benchmark for compositional text-to-image generation.IEEE Transactions on Pattern Analysis and Ma- chine Intelligence, 2025
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bbaecdb-570c-4a02-add0-10e33e9d3224 · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caba5c44-c215-4525-95ee-5dcf5f2a0fab · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f75b3b1-1986-4cec-9437-b15d04ed03f3 · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Flow matching for generative modeling
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffe1fea6-b322-4bed-ac8f-36b493bc5c8a · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Towards understanding cross and self-attention in stable diffusion for text-guided image editing
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68dbbf78-7163-49de-b3a1-be990f2c9c61 · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Timestep embedding tells: It’s time to cache for video diffusion model
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae64b6fb-a21a-42ba-8929-f53e8ee588b3 · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Flow straight and fast: Learning to generate and transfer data with rectified flow
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67b54e65-822f-41bd-b9e3-5fe551ddcaf1 · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Dpm-solver: A fast ode solver for dif- fusion probabilistic model sampling in around 10 steps
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22acf449-4dcb-4b89-883f-88e59253f854 · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fda6821e-32f3-49b4-b843-8c2684adee44 · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Re- thinking cross-modal interaction in multimodal diffusion transformers
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a9024af-1680-47c0-a149-3eed5e376324 · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Latte: Latent Diffusion Transformer for Video Generation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 483dffff-0608-43d2-a226-17c1ff7f515d · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Video generation models as world simulators
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dec7a44-6ec2-4a1d-b77c-73114fafccb8 · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers DINOv2: Learning robust visual features without supervision.Transactions on Machine Learning Research,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bea0785-31fe-437c-a12a-8b60f53c56aa · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Scalable diffusion models with transformers
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5dc1917-4917-4ec7-b264-26e79e43e1b1 · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Learn- ing transferable visual models from natural language super- vision
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35ecbd6d-7a7b-4e7e-b616-2615a0c751f6 · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers High-resolution image synthesis with latent diffusion models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5db906c9-2034-453d-a023-bff27804ba38 · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Laion aesthetics
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35cbfcbb-ebb8-4de8-838b-d35e04709cd8 · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Exploring multimodal diffusion transform- ers for enhanced prompt-based image editing
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e910398a-4711-4bfa-8d54-ab7d05014a3f · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Freeu: Free lunch in diffusion u-net
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3d00d7b-43cd-4cb2-a8e6-dc8689b4765d · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Denois- ing diffusion implicit models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 783579df-e02e-4bd6-b34f-61d11f1f1327 · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Consistency models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a739707a-5e72-4285-ba41-2ab6605fd67d · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Stable diffusion 3.5
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c0f15bf-3ead-45b1-93ab-7f44e558ac02 · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers OminiControl: Minimal and Universal Control for Diffusion Transformer
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ca6bce0-b200-4195-bbd4-9f705fe5b897 · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Hunyuanvideo: A sys- tematic framework for large video generative models, 2024
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b53c935-42b2-4c9f-b0a7-62b56133ecce · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Hunyuanimage-2.1: An efficient diffusion model for high-resolution (2k) text-to- image generation, 2025
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a9a88b3-9352-4759-abb6-930e227ffce2 · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Attention is all you need
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59737f79-45ed-430c-9b0d-ad8aff6542b3 · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Wan: Open and Advanced Large-Scale Video Generative Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1572279b-ff26-4a4b-bb28-7bf03d1c2f24 · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Freeflux: Understanding and exploiting layer-specific roles in rope-based mmdit for versatile image editing
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6415ef91-6f46-4a58-8c9d-2f011aab16a1 · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers A uni- fied framework for u-net design and analysis
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 327a4d58-e3ec-4920-a17e-96b3642f4a71 · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Qwen-Image Technical Report
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ec50b5c-b186-465f-b332-d45cd87c0e70 · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 611b0c85-4da9-41a9-a245-4c4b37db281b · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Instanceassemble: Layout-aware image generation via in- stance assembling attention
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b538b1eb-af94-4eb5-8a32-ede0895fb125 · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Imagereward: learning and evaluating human preferences for text-to-image generation
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24595342-8f84-46d5-ace6-aeb7b3aa6def · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Cogvideox: Text-to-video dif- fusion models with an expert transformer
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 994c3947-e3d5-48cd-8969-5104ca38c0ab · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Towards understanding the working mechanism of text-to-image dif- fusion model
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ee9f82a-8b81-4d8f-bb65-6fc7d147b935 · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Adding conditional control to text-to-image diffusion models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed1bbda6-819d-449a-bd41-43dab2455bd9 · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Exploring Position Encoding in Diffusion U-Net for Training-free High-resolution Image Generation
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fad3956d-e6d8-40d6-ba84-cac84187b02d · outbound
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Unresolved cited work
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a617b15-aa7d-4b04-892d-0c90d7ab9c50 · inbound
Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers for Text-to-Image Generation TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cb1ba7c0-5ec4-4cdc-8fbc-057a32fd6f27 · inbound
Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers for Text-to-Image Generation TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ed7b4e8-95a8-4511-bfe7-ef482499fa00 · inbound
Breaking the Lock-in: Diversifying Text-to-Image Generation via Representation Modulation TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.