Pith. sign in

Paper Citation Record · LEDGER

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation

As of 14 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 10 inbound Pith citation observations for arXiv:2412.03859.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.03859 v3

Coverage vector

measured 81 of 81 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T22:05:33.961490Z

measured 91 of 91 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:40:43.078627Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T08:26:47.756395Z

Reference resolution

81 of 81 outbound references displayed

  • verified exact0
  • verified fuzzy45
  • unresolved35
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dde7cb6b-6ca6-496c-b485-422903a65b18 · outbound

This paper cites GPT-4 Technical Report.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T22:05:33.607999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:05:33.607999Z digest=sha256:ac49877e7d903b6be245ed10fe24ceecf66bbc17a8193fd9f7b7f2df43aef19e

Observation 8c7eb923-77e6-4f70-8bd8-3d7057203b1b · outbound

This paper cites All are worth words: A vit backbone for diffusion models.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation All are worth words: A vit backbone for diffusion models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T22:05:33.613655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:05:33.613655Z digest=sha256:9fa1d68346083d04c93589ae9eb671d0435099a6345845fae3d1f87bdd70dddf

Observation 677ce3e1-9b45-42f4-b45c-793471d677c1 · outbound

This paper cites Improving image generation with better captions.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Improving image generation with better captions

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:35.378858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.618856Z digest=sha256:6ef412a98cc423785a6e3560f0de7dfb7d2876022aef803eb8c4f2aee3bc1a3e

Observation ed7e54ed-6d3f-4c73-b7e9-9a19ec7bdd5d · outbound

This paper cites Attend-and-excite: Attention-based semantic guid- ance for text-to-image diffusion models.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Attend-and-excite: Attention-based semantic guid- ance for text-to-image diffusion models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:35.365429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.623795Z digest=sha256:529b6fb8b4d6aa755066e44a87e9ac9bd02c6451fbb6fe5f1662a1bedbd9a0b1

Observation 9f84a8f3-c602-4eec-a369-d115eaa22c42 · outbound

This paper cites Textdiffuser-2: Unleashing the power of language models for text rendering.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Textdiffuser-2: Unleashing the power of language models for text rendering

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:35.350978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.628329Z digest=sha256:5e39f89966b11998c37779774a77e5c99b80cd2e801c24af24b5742cffc2c5a9

Observation e67f4de6-2d24-4899-b046-64bdfff50aad · outbound

This paper cites Pixart- α: Fast training of diffusion transformer for photorealistic text-to-image synthesis.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Pixart- α: Fast training of diffusion transformer for photorealistic text-to-image synthesis

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:35.334739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.632652Z digest=sha256:78b6ced3d12e8e6c42dc1c24bfc8dfa3cc4aaaacddfa36a9786f462281cabfc8

Observation f56d5274-4670-4d52-a659-4a0762c5ca6b · outbound

This paper cites Reason out Your Layout: Evoking the Layout Master from Large Language Models for Text-to-Image Synthesis.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Reason out Your Layout: Evoking the Layout Master from Large Language Models for Text-to-Image Synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T22:05:33.637744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:05:33.637744Z digest=sha256:30ed6d690c150807699574666614d531737217f2f15177c300c42f97197ea45e

Observation d242a493-9f20-4e08-b10c-97060422e4eb · outbound

This paper cites Visual program- ming for step-by-step text-to-image generation and evaluation.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Visual program- ming for step-by-step text-to-image generation and evaluation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:35.318705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.642220Z digest=sha256:a93aa2fa808eb8665db261339aa39bcc187069d44063175d5f81849aa471d979

Observation 82a46ef0-b064-4648-82ab-02345ab05bbe · outbound

This paper cites Be yourself: Bounded attention for multi-subject text-to-image generation.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Be yourself: Bounded attention for multi-subject text-to-image generation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:35.298889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.646602Z digest=sha256:12a649b7e7b76f65ad0471c454f6b39d3538b538f2d211c769f6de89018448f8

Observation ee663d83-3818-44a4-9e26-0a1118a9c889 · outbound

This paper cites The Llama 3 Herd of Models.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T22:05:33.650509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:05:33.650509Z digest=sha256:7f27d0539f0eaa5666b656ba5c198bf964cab2fdbe0d2bd13f447936893a731f

Observation 0978cecc-9323-4cc9-80e4-c821c3e254a0 · outbound

This paper cites Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T22:05:33.655435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:05:33.655435Z digest=sha256:86cbbad46a0abb23ed3b1bd87884d36c32f76e000230bfb22898df38391937e3

Observation 917bc8d4-2e9e-44d7-9931-24031007c25f · outbound

This paper cites Layoutgpt: Compositional visual plan- ning and generation with large language models.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Layoutgpt: Compositional visual plan- ning and generation with large language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:35.281786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.660019Z digest=sha256:f5897fb375d1db9f86caab3d69f59cecd66e4da45b959985e8b6b9750b7ab9fe

Observation 96d17313-8599-424a-a9c3-4a47ae00a08d · outbound

This paper cites Ranni: Taming text-to-image diffusion for accurate instruction following.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Ranni: Taming text-to-image diffusion for accurate instruction following

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:35.262897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.664188Z digest=sha256:b4a088f409d0d1019dc8895ca9d7c8e5defb539c4aac9a1b2347caa3b10d6735

Observation 6a0b6586-9b77-499a-b583-d36255a26bf2 · outbound

This paper cites Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T22:05:33.668005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:05:33.668005Z digest=sha256:bf25b4a5a1f606df029286194ab8c2fd252c8fe19060519770ce7dcc90f55d9d

Observation 1cfe4114-b8ec-4d6f-a53e-936f75e4d4eb · outbound

This paper cites Check locate rectify: A training- free layout calibration system for text-to-image generation.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Check locate rectify: A training- free layout calibration system for text-to-image generation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:35.244578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.672009Z digest=sha256:5f2f47cbe05ad96f869c5df5d6f3279c9088ff6ff32641258997e04e7dbf710d

Observation d950db3b-7894-4755-b2d3-2502ed0ef0ba · outbound

This paper cites Gemma-7b-it.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Gemma-7b-it

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:35.219104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.676336Z digest=sha256:1de748d2d86ea21f1fc7c1c71e98006e86c1fd279451cc0b6693d02188703205

Observation 5115498c-3644-43a3-af71-7bd0b4a8d75e · outbound

This paper cites Qwen-vl-chat-finetuned-dense-captioner.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Qwen-vl-chat-finetuned-dense-captioner

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:35.192500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.680326Z digest=sha256:bd1472ffdba397e92d143e79eff2f9cda9eb31b5274f0ce9371f7d8f06f2f3eb

Observation a96ccd68-fbfa-4327-bb67-19751e36e785 · outbound

This paper cites Qwen2.5-7b-it.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Qwen2.5-7b-it

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:35.175481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.684170Z digest=sha256:5b78235fba8cc7faacf13b43dd62514ba7f085cd46dd97627d9a415a97fa54ab

Observation dee13355-d281-402e-a115-1ff847f4bcb7 · outbound

This paper cites Layoutflow: Flow matching for layout generation.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Layoutflow: Flow matching for layout generation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:35.151023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.688344Z digest=sha256:5e1ec5fa86af64585056f2adea6afdfeedf11a2f3c4bc0dfb1838b4f7de160fc

Observation 51082934-53bc-4534-a1d4-fafe8f5ea16f · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Gans trained by a two time-scale update rule converge to a local nash equilibrium

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T22:05:33.692142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:05:33.692142Z digest=sha256:b842a22b8f92a91c4c1d7f2ba917fd526f4ede498eaff58438fe6db433302d0f

Observation 0c98c773-dca2-4065-8859-041c8e51a584 · outbound

This paper cites Denoising diffu- sion probabilistic models.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Denoising diffu- sion probabilistic models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T22:05:33.696264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:05:33.696264Z digest=sha256:a7b6962936c9dbc3b21615ded65823e6592b353ba4c51089ff7278604d5e69f3

Observation 4523de25-1d9c-4421-acde-23eea7302b50 · outbound

This paper cites Interactdiffusion: Interaction control in text-to-image diffusion models.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Interactdiffusion: Interaction control in text-to-image diffusion models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:35.115405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.699710Z digest=sha256:5547ecc425f0bb4e7b68d4cea35014163e998c5d37c134fc2ed6773e008f51f6

Observation 5aa908e3-9e8e-479b-b439-37faadb1b35e · outbound

This paper cites Retrieval-augmented layout trans- former for content-aware layout generation.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Retrieval-augmented layout trans- former for content-aware layout generation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:35.094819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.703860Z digest=sha256:716a2bdf22c795fcaac6275887746316c8004e25b242fd09ea9fb965fae2b303

Observation 77cfa17d-416a-4a7d-96bb-26656625b847 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation LoRA: Low-Rank Adaptation of Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T22:05:33.707742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:05:33.707742Z digest=sha256:7e6094c27a8785bd03ca0be097cdbca52ab142fcbb331ceb5f082359eb516c8a

Observation f5e818b9-5861-4f14-9567-a44610a4d411 · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T22:05:33.712028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:05:33.712028Z digest=sha256:08fcdf15688394ecb6dc1b1d823cc8aebd6666d599d88a49c91095fc3103dd58

Observation e9e63d8c-3d57-41ac-b815-1bae8f515038 · outbound

This paper cites T2i-compbench: A comprehensive benchmark for open- world compositional text-to-image generation.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation T2i-compbench: A comprehensive benchmark for open- world compositional text-to-image generation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:35.077595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.716794Z digest=sha256:d3ab1ac2dc13a50c03c15fc095f9baf611f59630b6291fd511f67f475ae8241e

Observation ff541deb-427e-43e2-9aff-dca559902a1a · outbound

This paper cites Qwen2.5-Coder Technical Report.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Qwen2.5-Coder Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T22:05:33.720697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:05:33.720697Z digest=sha256:6f3f7751dc7c37738acce6045e67ca68dba33d1dcd8539ea0bacf3f587f3604a

Observation ebeef20e-1b99-46c9-831b-0433b079b4aa · outbound

This paper cites Ssmg: Spatial-semantic map guided diffusion model for free-form layout-to-image generation.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Ssmg: Spatial-semantic map guided diffusion model for free-form layout-to-image generation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:35.062015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.724941Z digest=sha256:8ccf83b6257fffb85f5f4c3966484206f3002601db4d01b704f2b955a3922f25

Observation 6ce962de-516e-40a0-989b-345aea679cab · outbound

This paper cites YOLOv11: An Overview of the Key Architectural Enhancements.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation YOLOv11: An Overview of the Key Architectural Enhancements

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T22:05:33.728754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:05:33.728754Z digest=sha256:267d51337ea7412ac834d80f2499a5ecd43e65cdbfa9188cac370b7cab5e5e7a

Observation fe405240-e763-4be0-83b9-f9c7b5e1fb35 · outbound

This paper cites Auto-Encoding Variational Bayes.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Auto-Encoding Variational Bayes

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T22:05:33.733418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:05:33.733418Z digest=sha256:a3eb6ab98780eab1c4fc6b1c557c2a35d8a616eecd5b5da3275763640bb1fff2

Observation 5593edec-d8c1-40ad-ab7e-e07140ed461a · outbound

This paper cites Segment any- thing.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Segment any- thing

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:35.045145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.738062Z digest=sha256:9056f9a5200ef792c0bc4423325a5512ea409b64bd0d685d97b0d40f889934d1

Observation 7e29d7ed-5f3a-48b3-ac39-ee60a6b858e5 · outbound

This paper cites Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T22:05:33.742011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:05:33.742011Z digest=sha256:dba23a1ccbd69cd4ac3f519cd4f9a8aeb2e6660b7eb18e4cabf1d5b7b73069a6

Observation 4de00f28-f2d1-437a-904e-17ca85bb337e · outbound

This paper cites an unresolved cited work.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:05:35.025822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.746292Z digest=sha256:8900142efd5aadf864dad02be13020dfe748b4373b6bf75562c5df3acc015b13

Observation e0c1d53d-6763-4ae4-8cbb-ea2978498e56 · outbound

This paper cites Laion-aesthetics v2.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Laion-aesthetics v2

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:35.007692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.750188Z digest=sha256:2b05ca3b28058b0968502bcd4ab7b841ffed53dec5dd908151970e42e3fbbaea

Observation 9f399cbd-1d45-420c-8405-9e63cb2088fa · outbound

This paper cites GrounDiT: Grounding Diffusion Transformers via Noisy Patch Transplantation.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation GrounDiT: Grounding Diffusion Transformers via Noisy Patch Transplantation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T22:05:33.754136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:05:33.754136Z digest=sha256:1174797fad1852ae2b18ecd086502812f13a0fd0e14f7a4ba7b68d0bdbe419ed

Observation 4da9a0db-4be4-4185-bc8e-53d50651aa65 · outbound

This paper cites Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T22:05:33.757962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:05:33.757962Z digest=sha256:11b7061abb381b0d590dd61c812cc52b31ce745a34b7cd9a5540ad2f319dd503

Observation a92f2f6a-18c2-45ee-9073-65986c826ef4 · outbound

This paper cites Magicmotion: Controllable video genera- tion with dense-to-sparse trajectory guidance.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Magicmotion: Controllable video genera- tion with dense-to-sparse trajectory guidance

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T22:05:33.761586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:05:33.761586Z digest=sha256:2a240f70372202b758ae464f3556cc9b4abf93397d808f59b78ca045b714fa5f

Observation 91a0041f-f01a-4d61-9f25-fd59fcf9c256 · outbound

This paper cites WebVision Database: Visual Learning and Understanding from Web Data.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation WebVision Database: Visual Learning and Understanding from Web Data

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T22:05:33.765315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:05:33.765315Z digest=sha256:b4a14c02f032e088b4bfba9602fc67d84e2052e3cd615fd311d5fe659efe8df7

Observation 673e0759-53a0-4dc1-a458-cc4905230d1a · outbound

This paper cites Gligen: Open-set grounded text-to-image generation.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Gligen: Open-set grounded text-to-image generation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:34.989702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.769068Z digest=sha256:e49e69152af5fd18d921a5f3976eced5e3a962e78c36892de21f7ca42cb63ece

Observation c273de19-dd59-4a5c-81c9-e0b9d4f7d568 · outbound

This paper cites Hunyuan-dit: A powerful multi- resolution diffusion transformer with fine-grained chinese understanding.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Hunyuan-dit: A powerful multi- resolution diffusion transformer with fine-grained chinese understanding

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:34.977019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.773168Z digest=sha256:4b12cb33760a66cf7694d7a270ca0b6cd5eeddddd9ea446a340dbedfa6283965

Observation 6cfdf3d3-5561-4be1-91ee-2b8700674404 · outbound

This paper cites Llm- grounded diffusion: Enhancing prompt understanding of text-to-image diffusion models with large language models.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Llm- grounded diffusion: Enhancing prompt understanding of text-to-image diffusion models with large language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:34.963421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.777267Z digest=sha256:ff2446b723eb1153b47474396309217675b5c95bdb14a7ab9aa00d95cae22b2c

Observation 917bae66-e05d-412c-8ba3-ec6ce8b2b8cf · outbound

This paper cites Microsoft coco: Common objects in context.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Microsoft coco: Common objects in context

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T22:05:33.781480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:05:33.781480Z digest=sha256:b7fbb0a9c004bc75e3384e27d94a293136a50057dbee2561d455e92bef782832

Observation a114ed54-6af2-4595-9c78-ef507dfa1625 · outbound

This paper cites Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T22:05:33.785966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:05:33.785966Z digest=sha256:4b6279daf4446f4260860803e13258ca63f70524fd424c08940bf3c6538dc8f1

Observation 43774304-a336-4d34-bf58-96fd872bfd28 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T22:05:33.790107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:05:33.790107Z digest=sha256:6a38c2f6687055fafb5bd9e2bab25552183f3cfdfc538518c2e23dedde667b44

Observation fa15e9e3-a838-40c1-b1c9-66195f7e859f · outbound

This paper cites Hico: Hierarchical controllable diffusion model for layout-to-image generation.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Hico: Hierarchical controllable diffusion model for layout-to-image generation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:34.934378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.794641Z digest=sha256:e234e3d3d24f89e8639012c8ca219369509bee02254b6ecc48ce07c42d0baad2

Observation 1968cad9-8563-4d8a-be55-f05bc08ef31e · outbound

This paper cites Llama-3.1-8b-instruct.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Llama-3.1-8b-instruct

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:34.920248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.798832Z digest=sha256:6c5de018c3f099fb5b633143a5e69396611bd3be06b0cf4544ce6c0f98e6d73a

Observation eb20fdf2-a8b7-4011-80cf-922fb65ab373 · outbound

This paper cites Nerf: Representing scenes as neural radiance fields for view synthe- sis.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Nerf: Representing scenes as neural radiance fields for view synthe- sis

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T22:05:33.803104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:05:33.803104Z digest=sha256:3b63b11ebf1ed748a460cee91f65ac9ae8155be1109ea4eec56e965388dce706

Observation 183e40a0-0d4c-46e1-b30a-0000493de82b · outbound

This paper cites Compositional Text-to-Image Generation with Dense Blob Representations.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Compositional Text-to-Image Generation with Dense Blob Representations

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T22:05:33.808311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:05:33.808311Z digest=sha256:95b4dc6fb71a9685034b77a0ef768527b5744922e2da701998c0410d67964014

Observation 7c8928a9-1c35-418d-825e-4f9e38715e10 · outbound

This paper cites Minicpm-v-2.6.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Minicpm-v-2.6

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:34.896324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.812929Z digest=sha256:7ee0bba41f5b379ae1ca0176f4897e04aa800166ecd897935a37461927e6084f

Observation 5ecf1184-0071-4c1b-8cdc-598515f6b6d4 · outbound

This paper cites Scalable diffusion models with transformers.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Scalable diffusion models with transformers

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:34.881590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.817454Z digest=sha256:67bde57d13e6b0f06298e09aad87989c7da2c3d12c9c16a28919d66d9291584a

Observation 141d6be0-9ff5-4368-9e0f-4cbc3a2698ca · outbound

This paper cites Grounded text-to-image synthesis with attention refocusing.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Grounded text-to-image synthesis with attention refocusing

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T22:05:33.823023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:05:33.823023Z digest=sha256:63e6ce667d85058747ff34d81a9fe92f9c144c3f7be19fab481c97afa7e255c0

Observation cb58164b-b9f3-447b-bf49-f0cf1900ae76 · outbound

This paper cites Sdxl: Improving latent diffusion models for high-resolution image synthesis.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Sdxl: Improving latent diffusion models for high-resolution image synthesis

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:34.859874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.828162Z digest=sha256:88876390a45ecf6bd93d19aeed826613200d094dd5260de0f29ee410afda1695

Observation ce8fe400-5101-4bfa-9fc1-670791425991 · outbound

This paper cites Layoutllm-t2i: Eliciting layout guidance from llm for text-to-image generation.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Layoutllm-t2i: Eliciting layout guidance from llm for text-to-image generation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:34.844373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.833468Z digest=sha256:49335ccd0e51fb3b1ae72135b7c43cf6c1f23ea9271b1e0f95ccbc9b44170e6d

Observation 18ce0ed1-09d8-4704-bbd9-3023903deb95 · outbound

This paper cites Learning transferable visual models from natural language supervision.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Learning transferable visual models from natural language supervision

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:34.830187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.839111Z digest=sha256:1ac4b39c42d7f7c1ac6ee8d61fc2483fd9616dae9faa545a32d5503d7c5486bf

Observation c2163677-835d-44b3-84f6-48af9125147f · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T22:05:33.844464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:05:33.844464Z digest=sha256:02bd39d7d5aba8b13cc9b173271381a973f78a28b870fcc28c67cfb9e73f2068

Observation c71bbf37-d996-4972-ba74-bba306e8b3ad · outbound

This paper cites Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T22:05:33.848961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:05:33.848961Z digest=sha256:b17d7227ce7c1b0e1629ee4f0e51aef00c1a5f83b2d425f94a83eb11baea36b9

Observation ed319180-1018-4fcb-974b-d9e7ca5b12a9 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation High-resolution image synthesis with latent diffusion models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T22:05:33.853806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:05:33.853806Z digest=sha256:2a281b04a6aeb03c62cb2af0d3f537123398a147a8c001ccaadaf0b6734599ac

Observation f89ec6df-a78f-49fd-b36e-d9e7cb582568 · outbound

This paper cites Pho- torealistic text-to-image diffusion models with deep language understanding.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Pho- torealistic text-to-image diffusion models with deep language understanding

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:34.792154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.859052Z digest=sha256:556411c3512f5373ac190e5f45fcd932e261216c891f5709156e34dde6912580

Observation 590c86bc-bd44-434b-9c13-ceec6ae60ff7 · outbound

This paper cites Improved techniques for training gans.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Improved techniques for training gans

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:34.773788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.864002Z digest=sha256:2591f7fa85539dadf1cb52b1e894e662d4832858cfecc87f3c635b6a899be897

Observation a8809984-29d2-49d4-b8f0-8d8f1040824e · outbound

This paper cites Laion-5b: An open large-scale dataset for training next gener- ation image-text models.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Laion-5b: An open large-scale dataset for training next gener- ation image-text models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:34.760276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.868297Z digest=sha256:22300687a25e8afa6881aa4596e7fd0ddb4aeb1e9ce0ed1ffbbe3bd7cc3c0476

Observation 15356828-e118-4e34-aaf6-7d2fdff56f9d · outbound

This paper cites Noisecollage: A layout-aware text-to-image diffusion model based on noise cropping and merging.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Noisecollage: A layout-aware text-to-image diffusion model based on noise cropping and merging

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:34.744123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.872985Z digest=sha256:ff2099f8465c2e0ee0928d80b1a7236b334a9d222426fbb43be7d745f0a361f4

Observation 86f215b8-fd06-43dc-a8b8-f12d29134a7b · outbound

This paper cites Denoising diffusion implicit models.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Denoising diffusion implicit models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T22:05:33.877060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:05:33.877060Z digest=sha256:7d558ead31ac4eb0a666cb9111f6df5d2310b999e41f2309d2b9e0311a835fb6

Observation 57ab6d1c-8c72-47ab-9269-21a6900362f6 · outbound

This paper cites Stable diffusion 3.5.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Stable diffusion 3.5

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:34.719693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.881055Z digest=sha256:596d803d54b146aa9c35b530db5fec35b22b76cc8ad72947f7e89ea57dd48d8c

Observation 6fd0e363-cb66-4685-ba11-f537d274d88b · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Gemma: Open Models Based on Gemini Research and Technology

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T22:05:33.885027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:05:33.885027Z digest=sha256:2780b7e9709b9040697a496c1b52722bb1b455d6838e2fca60a916616e2fa469

Observation 87a96ed2-df30-4622-90fa-87a0abd5e49a · outbound

This paper cites Omost github page.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Omost github page

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:34.707951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.889224Z digest=sha256:9e3f4c49399c944acec393149abd0325674276b005ddd45d7eefe265286a33d2

Observation 7529171f-8a64-4c7a-bce7-ccd3c5170588 · outbound

This paper cites UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T22:05:33.892978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:05:33.892978Z digest=sha256:ab79509bf8d44cc568f301bcdbe34a07c7efff47132327e9bffa3d804419fc9c

Observation dbafb880-3815-4c74-8328-98761a31635a · outbound

This paper cites Instancediffusion: Instance-level control for image generation.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Instancediffusion: Instance-level control for image generation

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:34.693466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.897507Z digest=sha256:c544d63695741d61600a7af3f46bbf996d0ff192f1df2600f77dfc7c19957236

Observation 5bf8faf5-f475-4580-b945-46f2226ea138 · outbound

This paper cites Desigen: A pipeline for controllable design template generation.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Desigen: A pipeline for controllable design template generation

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:34.677121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.902123Z digest=sha256:911f801fbe9fc173756a8525ac3c0987597d1ea1db95bf5c5131fcd58d97a8f9

Observation 04250b7e-024e-4945-a136-202eb064c248 · outbound

This paper cites Self-correcting llm-controlled diffusion models.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Self-correcting llm-controlled diffusion models

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:34.654520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.906891Z digest=sha256:62d39431f6aec50b8cf646aa21f035ec14246024bbae80ce555f9210e7331e75

Observation 7fed15b3-479f-4821-be91-fd4d2259133a · outbound

This paper cites IFAdapter: Instance Feature Control for Grounded Text-to-Image Generation.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation IFAdapter: Instance Feature Control for Grounded Text-to-Image Generation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T22:05:33.911518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:05:33.911518Z digest=sha256:d7fa65f4221397b0ec8b445c5dd2f6cb3ec3e2c136f5e471b59ceb75de42b5fd

Observation d245cd8c-5fba-462d-aff4-78cccdbf6981 · outbound

This paper cites VIDiff: Translating Videos via Multi-Modal Instructions with Diffusion Models.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation VIDiff: Translating Videos via Multi-Modal Instructions with Diffusion Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T22:05:33.916043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:05:33.916043Z digest=sha256:6b801458abcc86dc71efcb4a3a6b7d70855efe3bae1544cf9623beb37ae01659

Observation 0a399c3b-4058-47c9-9419-442a4feb5a27 · outbound

This paper cites Simda: Simple diffusion adapter for efficient video generation.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Simda: Simple diffusion adapter for efficient video generation

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:34.637301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.920822Z digest=sha256:fb4a10755e2030ef6f3dc20bbda906254344e431c226541bc0c401e9080b28ca

Observation 3be4c486-dc70-4ef5-9862-1835e5a33be5 · outbound

This paper cites AID: Adapting Image2Video Diffusion Models for Instruction-guided Video Prediction.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation AID: Adapting Image2Video Diffusion Models for Instruction-guided Video Prediction

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T22:05:33.924765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:05:33.924765Z digest=sha256:3437f84f655a23a49bcaec18eaa86381753d371083eceefa3372dddc6425aaa1

Observation 1e6362bf-01c0-46eb-93b3-fa7088e21d83 · outbound

This paper cites A survey on video diffusion models.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation A survey on video diffusion models

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:34.623164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.930216Z digest=sha256:38a10b56c1d6a9aa73aaec155e055b5808e9af0efda92bf40c71770456e6b4b7

Observation 0338289c-af01-433e-bca9-405a9937e1c4 · outbound

This paper cites Imagereward: Learning and evaluating human preferences for text-to-image generation.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Imagereward: Learning and evaluating human preferences for text-to-image generation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:34.607588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.936573Z digest=sha256:96e0b48c47e61ef9eeb1215a6105fba2524d29d6c9c50924c824e76025bdbc28

Observation c8a0465f-6fbb-44bd-9c49-29e1e81a0c20 · outbound

This paper cites Freestyle layout-to-image synthesis.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Freestyle layout-to-image synthesis

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:34.588757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.940846Z digest=sha256:a68464b8698f815dfa843adf58069be517915f70bbc4e8a470487c896cff210f

Observation a605c30e-9520-406e-9a5c-a15313a0e538 · outbound

This paper cites Law-diffusion: Complex scene generation by diffusion with layouts.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Law-diffusion: Complex scene generation by diffusion with layouts

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:34.569900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.944624Z digest=sha256:60dca5ee07311278ea4073a34a6fbfc5f1657d0a35208781003b8514961aca49

Observation e94a0421-4d8a-4b71-8085-c6ea8ef84593 · outbound

This paper cites Mastering text-to-image diffusion: Recaptioning, planning, and generating with multimodal llms.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Mastering text-to-image diffusion: Recaptioning, planning, and generating with multimodal llms

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:34.553344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.948565Z digest=sha256:83518e580d0fbc4bea0c58a5c63ab95c64184264935241acdbcf2883f5ccc39f

Observation 466c2556-3d30-4979-9910-7bfbc9a38232 · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-11T22:05:33.952489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:05:33.952489Z digest=sha256:0c2f70cfc4e19078941d630f011013ee6a307ce77dd9eedbf4851d6c7d717561

Observation adaf16fa-094e-45a8-9c93-d968faa6d153 · outbound

This paper cites Layoutdiffusion: Controllable diffusion model for layout-to-image generation.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Layoutdiffusion: Controllable diffusion model for layout-to-image generation

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:05:34.537773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.957317Z digest=sha256:97dc4641850e707897b46687ade0406324006c2164afb3e925f4ae1489ec3624

Observation 99544d28-def6-46d8-bcf8-dafb204db096 · outbound

This paper cites Layout Image Text SiamLayout-FLUX Global Caption TokensImage TokensLayout Tokensℎ!ℎ.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Layout Image Text SiamLayout-FLUX Global Caption TokensImage TokensLayout Tokensℎ!ℎ

Reference 81

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T22:05:34.074029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:05:33.961490Z digest=sha256:7baaeab87ab20c5ff425c3032ac839a78311f186adba6b73a7c11718fa255bdd

Pith citing papers

Observation ac8abb00-9e24-45e4-9d47-8739112168cd · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.739050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:1780f77ac4796903db736cfa394cc4b1dc188efe2e6fd7c06baefb28d1bf91fa

Observation 089b1b99-3426-4303-bf25-1d647def6bff · inbound

ComposeAnything: Composite Object Priors for Text-to-Image Generation cites this paper.

ComposeAnything: Composite Object Priors for Text-to-Image Generation CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:43.078627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:40:43.078627Z digest=sha256:451ba1248a93ecaa75303c58b8e2cd713cf1109b638b77b724338e379c677cf7

Observation 92ca400d-5d62-4101-a4f9-5de2536ec658 · inbound

DreamPoster: A Unified Framework for Image-Conditioned Generative Poster Design cites this paper.

DreamPoster: A Unified Framework for Image-Conditioned Generative Poster Design CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:29.010130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:29.010130Z digest=sha256:252dd7979826926a55e7a563366078e791bb36eac286a03bccb8420a088d3480

Observation 980fec80-ed73-4571-a9b3-f7c778860428 · inbound

Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking cites this paper.

Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T16:43:53.304734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T16:43:53.304734Z digest=sha256:ce7d9837eccb46e044ac66dcdc1aaccdf9d3456481ec518a5f4f64231d2b654b

Observation 380659ca-630c-49bb-b413-9395a8638c4c · inbound

RefineAnything: Multimodal Region-Specific Refinement for Perfect Local Details cites this paper.

RefineAnything: Multimodal Region-Specific Refinement for Perfect Local Details CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:20:53.845509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T19:11:43.172296Z digest=sha256:790f228b4b7cf7614a6d311138b07ad7694787277423a7055f433027a8c3f0d2

Observation c06090a0-57d8-4db8-ae76-3157feff1802 · inbound

Benchmarking Layout-Guided Diffusion Models through Unified Semantic-Spatial Evaluation in Closed and Open Settings cites this paper.

Benchmarking Layout-Guided Diffusion Models through Unified Semantic-Spatial Evaluation in Closed and Open Settings CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:26:21.818643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-07T16:58:28.380188Z digest=sha256:319943edf0d1e853c8f918ab3660a47ae41b4eb8afcc7d3adbd7d92984c591a9

Observation 90e92c45-3511-40f5-a257-f526b0da77af · inbound

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling cites this paper.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation

Reference 98

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:28.822243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:f3a77f245c575f2a10981a990c4e1ecaadd88f84d53e28ece074a9b41bae44ca

Observation b2376fc3-d09e-491b-ab3d-ff9d59f8ace6 · inbound

Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization cites this paper.

Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:23:28.291397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T13:15:24.299457Z digest=sha256:57487b7fa6e2dfaa665625b2e5f44f045d786b45d6a98acb60eaef9a2cd05170

Observation 7d35254a-a1fc-4346-b477-e36c11b4937d · inbound

MetaPoint: Unlocking Precise Spatial Control in Agentic Visual Generation cites this paper.

MetaPoint: Unlocking Precise Spatial Control in Agentic Visual Generation CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:26:47.757884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T06:07:06.056441Z digest=sha256:6364e0dac919d99a0a0228d8969be3b8bdb4398e4ff3d4895b30b5c79b6693a4

Observation e0e5a2c1-0688-43dc-a5ca-ca7e2a97a114 · inbound

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models cites this paper.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T21:08:01.026456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:08:01.026456Z digest=sha256:c1991d8c010c8ce69035dee7da6402985cd656d2f19d32e9e5fe0cb8aad50080