Pith. sign in

Paper Citation Record · LEDGER

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs

As of 7 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2507.19939.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.19939 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:58:05.169969Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T09:52:01.768852Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact2
  • verified fuzzy35
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 07e9ae56-ddf1-447f-bc70-bf34fe543b0c · outbound

This paper cites Diffit: Diffusion vision transformers for im- age generation.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Diffit: Diffusion vision transformers for im- age generation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.338680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:58:04.946931Z digest=sha256:8874dd005160ffa759f247586357cf055685c9a850b2fae13f30099a38a956e2

Observation ba332586-5090-4171-81ed-75404db30ac2 · outbound

This paper cites Sketch-guided text-to-image diffusion models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Sketch-guided text-to-image diffusion models

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.325945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:58:04.951193Z digest=sha256:637dfba4afbd78a84ea5afc35b151cc5a799807042f9046a97553d3d4dcf4b97

Observation c4ee3e43-266a-467d-95fb-8239ef05321d · outbound

This paper cites Language Models are Few-Shot Learners.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Language Models are Few-Shot Learners

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:04.955410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:04.955410Z digest=sha256:37716c57aa82036e176a8953fd41cfabfe67c1cb379e3ec623dea0e874e45084

Observation b7146fd2-fcec-4aa8-aa49-eb21f363e065 · outbound

This paper cites Emergent Abilities of Large Language Models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Emergent Abilities of Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:04.959610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:04.959610Z digest=sha256:c3f3f5af3f68adf13536f3e6bb0c5dabe45e98e2b9b44b06fb9428cf8900a9d5

Observation 117e6c31-1f2c-4410-bbcc-8af04532b50f · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs LLaMA: Open and Efficient Foundation Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:04.964121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:04.964121Z digest=sha256:4f20acb7abc453be081705cc184142ba340f5739e30c83a8e06a472b6101e9dc

Observation 084b781f-ebe9-42a4-8549-550deea7ee0c · outbound

This paper cites Instruction Tuning with GPT-4.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Instruction Tuning with GPT-4

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:04.968184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:04.968184Z digest=sha256:1b97fe1c0a9eaadadd41291d0b6ce69b34fe31370184dc7370863eb7f88f6bcc

Observation e7c8796a-5a87-450c-9b6f-9242b661bb1f · outbound

This paper cites UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:04.972222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:04.972222Z digest=sha256:a984d63c4f053788106da5de593264011bd25bddfbe9f789adba325c9f148de3

Observation 171efb84-11c5-48a5-b5d8-9d20abe572af · outbound

This paper cites T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.313044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:58:04.976692Z digest=sha256:d282c178d6a01593c0f3e96ed2658eceb24780f75687920a9d4ebaf06bfc789c

Observation bcd86d95-aa16-4441-ae02-c68103fd6549 · outbound

This paper cites Migc: Multi-instance generation controller for text-to-image synthesis.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Migc: Multi-instance generation controller for text-to-image synthesis

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.299741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:58:04.980592Z digest=sha256:5dcbf58485ee99ee266e2a165fc3fd4c201ba6f205d5863b595a5ed9852dbd96

Observation 52a95384-79c8-48c5-9f2d-ef7c6610dcd7 · outbound

This paper cites Minigpt-4: Enhancing vision-language understanding with advanced large language models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Minigpt-4: Enhancing vision-language understanding with advanced large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.285704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:58:04.984214Z digest=sha256:7c7e79683bbc839f01f31542d8e80cdeb99b74d1668ecdbede7c54d75b068e74

Observation 6eaf496b-d8b5-457e-bbea-6eca8fcde988 · outbound

This paper cites Cogview: Mastering text-to-image generation via transformers.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Cogview: Mastering text-to-image generation via transformers

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.272597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:58:04.987777Z digest=sha256:7d008f8e111bbf3b7d2d73132bc1258fa60947e934deeb4ddeedd01296d33667

Observation 6900dd29-b12c-4755-a336-471b5f0335be · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:04.991565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:04.991565Z digest=sha256:05d4f611472157c515f030a3b9cb1cab1c55bfec776e77e68b8019b2743144af

Observation 13a8f93a-9671-44cd-9f4c-7d768096f52f · outbound

This paper cites Freeman, Fr ´edo Durand, and Song Han.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Freeman, Fr ´edo Durand, and Song Han

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.259576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:58:04.995851Z digest=sha256:aea3cc60e584dba26018458ac6aa9a0f0a43fdee82fb05dab8e80c1ab3ab054d

Observation e372bee8-fdf3-427a-a54e-9430c798aad7 · outbound

This paper cites Freeman, Fr ´edo Durand, and Song Han.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Freeman, Fr ´edo Durand, and Song Han

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.245961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:58:04.999510Z digest=sha256:17b547395bb6715df76fd3e9b4e151d054788c6e241194257ba0765e026b3744

Observation a026925c-8916-426c-af89-1f7a18dfca60 · outbound

This paper cites Cross-modal contrastive learning for text-to- image generation.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Cross-modal contrastive learning for text-to- image generation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.231757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:58:05.003315Z digest=sha256:4d393b5b7981227017d1a8812c383f2c2100577abd6a131303cf8117576f7dc5

Observation d7ddf754-1a80-441a-9475-41b3333b530c · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.006991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.006991Z digest=sha256:7baf8273642362b94677c2fd3f9318e4258e4ccb856334f90946d2284301798d

Observation c672f31d-3987-41dd-ac11-5cb1fdbadc9d · outbound

This paper cites an unresolved cited work.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Unresolved cited work

Reference 17

Resolution
verified exact
raw_fallback, observed 2026-08-06T13:58:05.764736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:58:05.011442Z digest=sha256:59cf4e85e74170099a0778b56aa88b2d6f3805f2d66fe6dd0eab771fb78558f8

Observation f49346e4-5003-42e9-b883-139052201eca · outbound

This paper cites Scaling autoregressive models for content-rich text-to-image genera- tion.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Scaling autoregressive models for content-rich text-to-image genera- tion

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.015533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.015533Z digest=sha256:af61d7d66075a5e7142f54738626fda8f6f31524a90c0ec6b55a702022a3a109

Observation 49cc29f7-ae85-47f5-9d8d-d0354b7ef22d · outbound

This paper cites Denoising Diffusion Implicit Models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Denoising Diffusion Implicit Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.019334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.019334Z digest=sha256:1ed42a4efd4d4e3baaadb643c266b78da6a2eb96d58b94b8ce93bbf3e0322baf

Observation 387a71a9-9a1e-4380-965d-b4a3df88279d · outbound

This paper cites Open-vocabulary panop- tic segmentation with text-to-image diffusion models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Open-vocabulary panop- tic segmentation with text-to-image diffusion models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.218527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:58:05.023001Z digest=sha256:905906600041b9e2762ca536511fd4c06aa5800d4e63a4a16563ae3910a99b5b

Observation bb990f3b-b75d-4482-b143-6cdd2f770cd2 · outbound

This paper cites Classifier-Free Diffusion Guidance.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Classifier-Free Diffusion Guidance

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.026909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.026909Z digest=sha256:2084aab014b37b7f2fb1f262670da0acb3a699ce2af96aaf0d21476ca6e5a4ee

Observation 045cb68c-5a5e-4c06-8745-e09cd8738412 · outbound

This paper cites Perception priori- tized training of diffusion models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Perception priori- tized training of diffusion models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.205937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:58:05.030582Z digest=sha256:68347ee121957ab4312362a7808c367dae5a5b456278ba8da53035123852d10e

Observation 01ca8309-71cb-4dc4-9625-b066f51b68a3 · outbound

This paper cites Training Compute-Optimal Large Language Models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Training Compute-Optimal Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.034238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.034238Z digest=sha256:32dca58687769b648baf2b60b85744381c47140fd03b0c6e588522f278b269d0

Observation ba033c7d-4d55-45e4-9218-fd745383be60 · outbound

This paper cites GPT-4 Technical Report.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs GPT-4 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.038403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.038403Z digest=sha256:dcbea364816b9c8ec11ae6e711f852bb46e34bfbe3c44c30a9ccdf492bf0f9ab

Observation 96d40ce9-fc3e-47cc-8936-c5206b6cfb6e · outbound

This paper cites Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.193167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:58:05.042052Z digest=sha256:cf9ffa747afd2b4172a3ba878caed60c6ce0fb65a47b38dde2bb87cbe6a78dcc

Observation dec473ca-5274-436a-8276-a27034d4da20 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.179252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:58:05.045817Z digest=sha256:04eee80fa166e19b5d163ccd3078db40b23e814183ea569328eb28e8b18adbfa

Observation aedc824d-b6f6-442a-a7bd-ae989798ff91 · outbound

This paper cites Composer: Creative and Controllable Image Synthesis with Composable Conditions.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Composer: Creative and Controllable Image Synthesis with Composable Conditions

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.049275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.049275Z digest=sha256:61449abbb9ea1f4522b8a85595fc98cfaa8c92909b64dff1b372195b7c997640

Observation 0a970c1c-d1f8-4434-9689-93710fb5c2ff · outbound

This paper cites LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.053158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.053158Z digest=sha256:4b071f12d15da08ddb759b325c68ec297dfe6ee42985d3cb8ab603c88574085c

Observation fd0473fa-31b1-4309-9aeb-72ca7fb8cc91 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Adding conditional control to text-to-image diffusion models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.165248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:58:05.057314Z digest=sha256:d3d80b60a202cf28778d375394b49511fd50e13004350223993bbbdcd7afe57b

Observation d0195274-c817-4e8b-a3ba-0f901497dd7d · outbound

This paper cites Fourier features let networks learn high frequency functions in low dimen- sional domains.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Fourier features let networks learn high frequency functions in low dimen- sional domains

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.152391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:58:05.061006Z digest=sha256:43a50a26d90b3ff1ff165deddfe8bb86ee5cb011694a4c8f85c2ec9a4904eab0

Observation 9d934d97-73ad-40df-b623-a83d283a0493 · outbound

This paper cites Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.139406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:58:05.064778Z digest=sha256:7e96be92916db4b757c487027cdcf525489ab693b892fc6995a5effbf748fac8

Observation 5dd8cdb2-bf75-499e-b609-e59e05b3f2c8 · outbound

This paper cites Hyperdreambooth: Hypernetworks for fast personalization of text-to-image models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Hyperdreambooth: Hypernetworks for fast personalization of text-to-image models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.127185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:58:05.068589Z digest=sha256:df9a4b28c53563b6687197359f7d1870820bf4f2f64d523fad6ecafd210c115d

Observation d3c4b848-ad70-490f-9bbc-fdcc948ddf11 · outbound

This paper cites Spatext: Spatio-textual representation for con- trollable image generation.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Spatext: Spatio-textual representation for con- trollable image generation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.114259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:58:05.072153Z digest=sha256:ae918aafd942ac8e3de8a40320570e3e830e8354d646ff6319ff8f653f6ecaa3

Observation f00b8a3d-e113-493b-ae07-9bb2c96cdf12 · outbound

This paper cites Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.075777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.075777Z digest=sha256:70d3a77b8d51b86082918bf0d2a8cc496465c389fd057089e0222704fd000f7d

Observation 5f9b0f5c-a565-44a1-91c0-e3870f289e9c · outbound

This paper cites Diffusion models beat gans on image synthesis.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Diffusion models beat gans on image synthesis

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.101424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:58:05.080272Z digest=sha256:b8fbc8ca5218386aae9104f7b8633a8ccaebf5c1c23bdf9d7e0bc68d2aed8408

Observation a820c14e-3831-45c8-b988-edf429f9c255 · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.084309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.084309Z digest=sha256:08d523110e41c745c4b63ea968f76ae4d5ba540c3808d31400dccad2114bb054

Observation e18275d2-9c57-48f6-a3e7-aca56999271f · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs High-resolution image syn- thesis with latent diffusion models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.089138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:58:05.088465Z digest=sha256:6c8ddcbde802dd933278613698f38df319cf71a73bad21bb5fb4f7bdc72d2f68

Observation 1590a588-bddb-4a99-9aa1-282788ff9afe · outbound

This paper cites Generative ad- versarial text to image synthesis.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Generative ad- versarial text to image synthesis

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.076676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:58:05.091929Z digest=sha256:b0301be7802324d39a4825ca577f118c0500a2e32d8e5760e5e16192fc2c7e54

Observation 2882007e-6483-443f-b211-b9d621daf465 · outbound

This paper cites Younes Mirinezhad.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Younes Mirinezhad

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.063281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:58:05.095580Z digest=sha256:6996967fc2221cfb6b74f2f646fe933726d3f19d195c3d3147da28b04d988525

Observation aa69289a-f977-41b8-8bce-40f4f1818b70 · outbound

This paper cites an unresolved cited work.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-06T13:58:06.050460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:58:05.099809Z digest=sha256:9475aabf981fb54afcf20e54f37341aa1c34761cb35fe1eb36a58bffef45deec

Observation 8b607d78-345b-402f-9019-477285db138b · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.103489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.103489Z digest=sha256:9ede8f97c769e1c23b79f9de0d981ab0b3b4e850d3882d4dc7e4d7eaf6888265

Observation dff75b55-9d13-4164-81f5-79445e019411 · outbound

This paper cites Freecontrol: Training-free spatial control of any text-to-image diffusion model with any condition.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Freecontrol: Training-free spatial control of any text-to-image diffusion model with any condition

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.037096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:58:05.107430Z digest=sha256:32e6b802df3b686b6ca49cf1d2f99b1436c21e77e16199364ffe2aa7c42160b4

Observation 9f202afb-5277-4256-9229-d79520cce8a1 · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Deep unsupervised learning using nonequilibrium thermodynamics

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.024283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:58:05.111372Z digest=sha256:33cc2d613bfb0411cff8fd7a0e37eda342aaae8e988154c286715b1623da4680

Observation 365b3635-be53-442c-bcc9-5976687a29ed · outbound

This paper cites Score-based 9 generative modeling through stochastic differential equa- tions.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Score-based 9 generative modeling through stochastic differential equa- tions

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:06.011333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:58:05.115135Z digest=sha256:471de74114a9fa7988bb6f1fb880ce6854cd08a9dcf078b4f19493408eea0837

Observation f4a4d41f-fe1c-4a7e-9c56-368377856526 · outbound

This paper cites It’s all about your sketch: Democratising sketch control in diffusion models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs It’s all about your sketch: Democratising sketch control in diffusion models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:05.997752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:58:05.119314Z digest=sha256:910767350d85eb8478c6475f8c052888aca25b756b9dfd108983d5c6136cfb47

Observation ad81617a-d9db-4dc4-867e-6a734b62bd54 · outbound

This paper cites Attngan: Fine- grained text to image generation with attentional generative adversarial networks.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Attngan: Fine- grained text to image generation with attentional generative adversarial networks

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:05.984472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:58:05.123447Z digest=sha256:ad9890347e4ec9c9191729f03912629adc08fe9144511d89056449ea4a3d0b43

Observation 042213c5-e384-4891-81bd-13bd131496f3 · outbound

This paper cites Plug-and-play diffusion features for text-driven image-to-image translation.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Plug-and-play diffusion features for text-driven image-to-image translation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:05.971428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:58:05.127180Z digest=sha256:7cde7651871153aca8f503de33ea70c860f32515fcc93708f98482ddd152ded3

Observation 40174456-e7e8-4dc2-aaae-aec564f43fd0 · outbound

This paper cites Compositional Text-to-Image Generation with Dense Blob Representations.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Compositional Text-to-Image Generation with Dense Blob Representations

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.130972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.130972Z digest=sha256:0cb74229aa8e317fb55aea9c4e9b0202c6e05d54ce7798855fc9bd1d5d236616

Observation 85990893-f0a0-45e3-b0ac-bfd9e6f7ab99 · outbound

This paper cites Layoutgpt: Compositional visual plan- ning and generation with large language models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Layoutgpt: Compositional visual plan- ning and generation with large language models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:05.957221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:58:05.135424Z digest=sha256:d9fcfca6ebd8a9cbd826e926391d1c06a49ef37661fbfc357f4785405207852c

Observation 6625bbde-6465-4073-9aa9-78dc70909b7a · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.138991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.138991Z digest=sha256:f12c0a8701c2d2f8ac5ffa6fe86ae33802a7f351d29cffe07fbed1c0627ee36b

Observation dac461c6-ad6a-48a2-86ce-d515f6dd62ed · outbound

This paper cites Humansd: A native skeleton-guided diffusion model for human image generation.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Humansd: A native skeleton-guided diffusion model for human image generation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:05.943248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:58:05.142871Z digest=sha256:870dd51955883289a745c81e0f654aef80d7abe9ba2da5592920bb507dbd41b5

Observation 2be2634a-9867-456f-97c1-6ceb237ba7e2 · outbound

This paper cites eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T13:58:05.146539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:58:05.146539Z digest=sha256:9685223bddc1c5fec9d9e05c115531bb7e808ddf67d283abf2604327377586bf

Observation 8b6efc69-2ffd-4ae7-bccb-0251a4f3104e · outbound

This paper cites Lafite2: Few-shot Text-to-Image Generation.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Lafite2: Few-shot Text-to-Image Generation

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:58:05.211626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:58:05.150622Z digest=sha256:3a14a8459e02425a6d02f90148e6eaa0f98393e6574fadc823b0cb1aac608a05

Observation b1d736c3-2c9e-4aed-b529-7dc8e21c151e · outbound

This paper cites Gligen: Open-set grounded text-to-image generation.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Gligen: Open-set grounded text-to-image generation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:05.929027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:58:05.154643Z digest=sha256:1dd2eeb168e347c5bc8eba6abb984428131edecc43f0862041086c50ae44352c

Observation 64e4b6a5-2c47-43dc-998c-bf7fb7cd7e00 · outbound

This paper cites Ranni: Taming text-to-image diffusion for accurate instruction following.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Ranni: Taming text-to-image diffusion for accurate instruction following

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:05.915734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:58:05.158515Z digest=sha256:8bc5939c1d4b2bccd1b7025bcd884a017f91e5b0d7de9f2d89656092fc9f719c

Observation c05889d0-fa51-4631-a53c-603f36e5e8fb · outbound

This paper cites Smartedit: Ex- ploring complex instruction-based image editing with multi- modal large language models.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Smartedit: Ex- ploring complex instruction-based image editing with multi- modal large language models

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:05.902359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:58:05.162408Z digest=sha256:f4cc3dbc1c3e8530bb6d556da344e519d0c424e943b887b5384bdb7b1e2a8b3c

Observation 0fbbccd9-48a4-463e-90e8-709a42ec2a94 · outbound

This paper cites Layoutdiffusion: Controllable diffu- sion model for layout-to-image generation.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Layoutdiffusion: Controllable diffu- sion model for layout-to-image generation

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:05.888090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:58:05.166417Z digest=sha256:9c896a5292b7b8b8c8e6a57e62176c2257c138c6cd86a4b260ed3c8c9715506c

Observation aa617503-f3f4-4e57-a2f1-b75cbf59137a · outbound

This paper cites Reco: Region- controlled text-to-image generation.

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs Reco: Region- controlled text-to-image generation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:58:05.875093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:58:05.169969Z digest=sha256:b7d65f2e00210d6bce96b8cce38b6b2daa45e14aa12a811d39abad3dd090b86c

Pith citing papers

Observation 136e6b49-1b99-449a-b459-2878b5d978c5 · inbound

EventOD: Event-Aware OD Flow Generation via LLM-Guided Semantic Modulation cites this paper.

EventOD: Event-Aware OD Flow Generation via LLM-Guided Semantic Modulation LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T09:52:01.768852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:52:01.768852Z digest=sha256:603d669bb578ffcc2969125a0476b960ce36ea9dfd17b1bcebe81584a52c3b7e