Pith. sign in

Paper Citation Record · LEDGER

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models

As of 21 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2607.16409.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.16409 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T21:08:01.413854Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved48
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2ebf8f38-836b-4075-80e2-bdf23294a8ef · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T21:07:56.778827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:07:56.778827Z digest=sha256:ef47803d93fb3770ec5b612b34c2115c24fe95c0ca8c61210f749d9794bb90a3

Observation dbe82edf-c02a-4661-a88e-efd3394e3473 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T21:07:56.846680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:07:56.846680Z digest=sha256:df6232c658294563d3f044a4fa287c547387b8cf529ba1e8a4f9c91b5af218c4

Observation eeb42eef-8544-4ac2-8af6-eae91e580566 · outbound

This paper cites Qwen3-VL Technical Report.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models Qwen3-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T21:07:56.933862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:07:56.933862Z digest=sha256:086e6252642182c799215e2c27b3881df0d2090933a9c53476516e07ec6bd79c

Observation 3f2f0e2d-9c25-493c-8a7a-0a695b88da38 · outbound

This paper cites InstructPix2Pix: Learning to Follow Image Editing Instructions.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T21:07:57.019322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:07:57.019322Z digest=sha256:75e993a4a2cc1d09c115a8c500849b15dc61c21557efb911ac6ee0d622b7df38

Observation 6f0a6611-1f27-4d66-a6bf-16183931a5fe · outbound

This paper cites HunyuanImage 3.0 Technical Report.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models HunyuanImage 3.0 Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T21:07:57.128157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:07:57.128157Z digest=sha256:6fd73942ef2d1fbab8a0b68ab9d325186948b42581cb94c74f6e4af6b0db5997

Observation e826a919-f622-432c-b162-5a1fa3240afb · outbound

This paper cites Pixart- α: Fast training of diffusion transformer for photorealistic text-to-image synthesis, 2023.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models Pixart- α: Fast training of diffusion transformer for photorealistic text-to-image synthesis, 2023

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T21:07:57.189066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:07:57.189066Z digest=sha256:8439a2ef9b5d35b5f73b1c35a8cc9c0d534dd1f7e9cc0aa0548cf7f2c3fad263

Observation e2d7e3fe-1265-4068-8106-3d420f5b21dc · outbound

This paper cites Training-free layout control with cross-attention guidance.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models Training-free layout control with cross-attention guidance

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T21:07:57.280070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:07:57.280070Z digest=sha256:247c11b0257968b05f284e12c3b6f53012c1f08539176786e615189f1debf2ad

Observation ee573901-cbdf-44d2-9340-0aec73a8638f · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T21:07:57.402970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:07:57.402970Z digest=sha256:921fb3c43dcf800dfcb8daa9f4798a1ba73ba97b56872a78a0e34383223ebade

Observation 7c51226a-d37e-4988-b3e9-f154fa03f3f0 · outbound

This paper cites Visual Programming for Text-to-Image Generation and Evaluation.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models Visual Programming for Text-to-Image Generation and Evaluation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T21:07:57.494609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:07:57.494609Z digest=sha256:7c3253b48a9a0d772eb2dd5b7be10924301f4b170f97f3e669cdbd8927d6522f

Observation 3e558af5-b06b-4306-864d-64b13569f507 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T21:07:57.570181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:07:57.570181Z digest=sha256:ca9714d1581dbcd4cfeb615431f44dcb2a33a46b1141381277ed1bd8eba6ef3f

Observation a1c4875e-6fbc-45b9-9f2b-eb7b5bcd348e · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models Emerging Properties in Unified Multimodal Pretraining

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T21:07:57.616596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:07:57.616596Z digest=sha256:06b1dbca295783bc8ca5c5ad237ace357ddcb8d9e11c17d1c7ac07c41e8e43d5

Observation e698a7fd-a5a6-4515-84b0-1fcef70a4259 · outbound

This paper cites Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T21:07:57.686090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:07:57.686090Z digest=sha256:42475d0a599885c42c8386b9cfdd14bcd3f093ed344bacd4c127d59a11e31c91

Observation 08b4b17e-d836-40a1-b6ff-038dadcdbdd3 · outbound

This paper cites LayoutGPT: Compositional Visual Planning and Generation with Large Language Models.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models LayoutGPT: Compositional Visual Planning and Generation with Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T21:07:57.741701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:07:57.741701Z digest=sha256:601534348a35899d7b06e2cca3fed36707f41727529940489313e613ce6ff91e

Observation 28f6922f-77bd-4ac0-abf5-2f7d213d6661 · outbound

This paper cites Guiding Instruction-based Image Editing via Multimodal Large Language Models.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models Guiding Instruction-based Image Editing via Multimodal Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T21:07:57.941825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:07:57.941825Z digest=sha256:69996c4b6ad41592fbcb0ae0fa349b601e4adfce0557c7344b145c61200792bd

Observation 78358746-fa3b-43c9-ac3f-d883cbf6e34e · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T21:07:58.084967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:07:58.084967Z digest=sha256:a5244cacd0e57a2dc20f0a135f8724755094c84b938f74e6f96d6a9f5428bd89

Observation 5747dfdb-e0ce-44f3-a579-a3898db8b249 · outbound

This paper cites Thinkmorph: Emergent properties in multimodal interleaved chain-of-thought reasoning, 2026.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models Thinkmorph: Emergent properties in multimodal interleaved chain-of-thought reasoning, 2026

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T21:07:58.183372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:07:58.183372Z digest=sha256:a9aee31881d2ca25fcd5667259768f0341856ad52837baf2264ac87b37b94970

Observation daeefbf0-0480-4e24-9dbd-0d08d37efcee · outbound

This paper cites GLDesigner: Leveraging Multi-Modal LLMs as Designer for Enhanced Aesthetic Text Glyph Layouts.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models GLDesigner: Leveraging Multi-Modal LLMs as Designer for Enhanced Aesthetic Text Glyph Layouts

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T21:07:58.295635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:07:58.295635Z digest=sha256:e4ffdf43a045472fbb741d92773e00d778eaaeb5c7be72e86a059533032dd367

Observation b6c98946-d524-487f-ae0a-aa19a7b15b74 · outbound

This paper cites Plangen: Towards unified layout planning and image genera- tion in auto-regressive vision language models.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models Plangen: Towards unified layout planning and image genera- tion in auto-regressive vision language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T21:07:58.446324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:07:58.446324Z digest=sha256:e97fb667710dc26e3111573cfbcdfee6a5fb486f084d1471da82cef5aab350d7

Observation 1abc0145-2f89-4cde-b60d-272a7892e0be · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models Prompt-to-Prompt Image Editing with Cross Attention Control

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T21:07:58.614635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:07:58.614635Z digest=sha256:198523c0f512a239e3d131f804d5ec506ab3c3edb99073ca4c17abd74404fdf5

Observation d92e1202-1a7d-4e3f-a27d-2abe02fc04db · outbound

This paper cites T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T21:07:58.733061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:07:58.733061Z digest=sha256:8c188d4f5344203fe1de7ed9254d562d7ea5127ee60cfc5d97a7fec7cd5cba9a

Observation 38a7ba4b-a4bb-43e8-a350-e078fbd97463 · outbound

This paper cites GPT-4o System Card.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models GPT-4o System Card

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T21:07:58.891632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:07:58.891632Z digest=sha256:c30d158d6257f17812b8900e8c85217e9aceebd81d3b31a43e75f8acda88c22f

Observation b8476964-01ed-4ee9-987f-055bcb89098f · outbound

This paper cites FLUX.2: Frontier Visual Intelligence.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models FLUX.2: Frontier Visual Intelligence

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T21:07:59.051269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:07:59.051269Z digest=sha256:25163bb8973036b72f5d796f0a78e1cb656012eeb8b0c223e24bef8fe410f7fc

Observation 866bec57-0524-4aa8-a248-034ec9417631 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models Evaluating object hallucination in large vision-language models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T21:07:59.217166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:07:59.217166Z digest=sha256:30ca4ac88ec4e1914221a592f0c27414ea372c133ea0a5f5ebccf2f108f27b0d

Observation 3ba5d23e-15f0-4971-a3d3-34d39f4fa50f · outbound

This paper cites Gligen: Open-set grounded text-to-image generation.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models Gligen: Open-set grounded text-to-image generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T21:07:59.385173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:07:59.385173Z digest=sha256:0798f7509bc8461fdfbee9293612caf0cbe5038f8448ebfa3e2e3790a894d837

Observation fc7c65c8-d4f5-4c41-b41c-6eba13ff0928 · outbound

This paper cites Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T21:07:59.483320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:07:59.483320Z digest=sha256:2c380cd9ce77806c2f6272ad81117f4c355a7a37ed40c26d8ef36de708c0ff70

Observation d14b4f97-c2a7-4db3-acd5-ff77c4eea622 · outbound

This paper cites Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T21:07:59.570641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:07:59.570641Z digest=sha256:a3de5b85e1065ec2ddcf05b48add92799307629624a7f402165d2efee63bc7d5

Observation da169950-af7f-49c6-b69f-58d4624f80ac · outbound

This paper cites Visual Instruction Tuning.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models Visual Instruction Tuning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T21:07:59.627471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:07:59.627471Z digest=sha256:5f6c8d49649e3ef3d66e8f4295728d695128a5cbddc0b9b156c71cb2bef018d8

Observation 52ef7976-b369-42dc-a007-e48ceec4686f · outbound

This paper cites Flow-GRPO: Training Flow Matching Models via Online RL.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models Flow-GRPO: Training Flow Matching Models via Online RL

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T21:07:59.699656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:07:59.699656Z digest=sha256:02ff89363525fef2b5d8d7337296560e20745a9e1354d66f1fdd8208ee641438

Observation 6f5c496c-9496-48a7-acea-52fa375fb198 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T21:07:59.782941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:07:59.782941Z digest=sha256:8a41ba8cf0298061825da1f393aeacd09708742151f44dd84ed67d0c89dd313b

Observation 30ead908-e6fe-4f57-a722-814dd1dfae74 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T21:07:59.848307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:07:59.848307Z digest=sha256:17087f6acd66beb606f13168011cbbbbf97d1d02ea89ed107acf2f861516fdc6

Observation 32c959b0-3f83-490b-a9ae-e295a9b81ce9 · outbound

This paper cites Chameleon: Plug-and-play compositional reasoning with large language models.Advances in Neural Information Processing Systems, 36:43447–43478, 2023.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models Chameleon: Plug-and-play compositional reasoning with large language models.Advances in Neural Information Processing Systems, 36:43447–43478, 2023

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T21:07:59.922080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:07:59.922080Z digest=sha256:8df6d0af2db23041b7b044cc4730dd28997dedb14997d9dd175935c97f6b09a6

Observation 2330adab-80e0-4d3f-a802-d4bc4be40eee · outbound

This paper cites T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T21:08:00.007546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:08:00.007546Z digest=sha256:25a7712ac4abdf59b9947619577f6b8860209f54048f1e629f339df65f5142c4

Observation 13fb5e99-7538-49a4-bc78-4194a94710d3 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T21:08:00.086880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:08:00.086880Z digest=sha256:627a6156380ab52d75dac5d2e6ecf4107e915ffedc426ecd0c2e26f8657c90b6

Observation ef147ebe-b2d8-4010-84a7-a0742bf7956f · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T21:08:00.135096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:08:00.135096Z digest=sha256:e5de983316a4546139ce4eded4379fc8376247025c35f6657b7f262de4392a93

Observation 6b143778-ca7b-47c4-aa8f-c0d7a7e6aa09 · outbound

This paper cites Emu: Generative pretraining in multimodality,.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models Emu: Generative pretraining in multimodality,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T21:08:00.211440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:08:00.211440Z digest=sha256:8dd38201d6273e14cf332efd61ea5f1b834351e0db3160f147b5255037d6e9e8

Observation be6ed78c-0ce3-4585-9c2f-2e4abee221d9 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T21:08:00.356603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:08:00.356603Z digest=sha256:8718fb5e649eec79bf46eb0b4661c0d67244bb23015aec8ff860ac1bc0ca745a

Observation 64fba9d9-cc0f-4872-a406-6befed4a4bdb · outbound

This paper cites GenArtist: Multimodal LLM as an Agent for Unified Image Generation and Editing.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models GenArtist: Multimodal LLM as an Agent for Unified Image Generation and Editing

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T21:08:00.430993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:08:00.430993Z digest=sha256:aa2468170bbdcea9977addf89a86ee8252a1a03d83fe5fedde3e207193763f86

Observation 7f3dedc3-9750-4c7d-b097-016502994e39 · outbound

This paper cites Boxdiff: Text-to-image synthesis with training-free box-constrained diffu- sion.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models Boxdiff: Text-to-image synthesis with training-free box-constrained diffu- sion

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T21:08:00.509482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:08:00.509482Z digest=sha256:e09dbd4f6f0d3500c8f441a5a4ac0f4d1497ab6b24040d030b52777b6749d7e0

Observation 0e7fba52-72f6-43e1-b5be-db58ac2f4328 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T21:08:00.612135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:08:00.612135Z digest=sha256:ff380df754d09a236662f16ab2ad74f68e017fd2d8e0177bdbf146a42abb12a1

Observation c2b4ea7e-ac01-40a4-bf5e-1f4d994afd71 · outbound

This paper cites Show-o2: Improved Native Unified Multimodal Models.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models Show-o2: Improved Native Unified Multimodal Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T21:08:00.675288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:08:00.675288Z digest=sha256:086162858e27ccf2e4d595f7d1292dbbaa70742ff4610dcab3213cc41f1e94c9

Observation 9da7ad30-90f1-49d0-bac6-a8c2b3da2368 · outbound

This paper cites ImgEdit: A Unified Image Editing Dataset and Benchmark.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models ImgEdit: A Unified Image Editing Dataset and Benchmark

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T21:08:00.788019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:08:00.788019Z digest=sha256:e9aed8c3cce6f7e40beccd5446fffacd9665c1a79398dd718f58d955597a7b82

Observation 91696826-7d3a-4d6a-a37e-74a820d0c2b7 · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T21:08:00.891370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:08:00.891370Z digest=sha256:2c6089e48ee0daf9aea435e3125e14cccc96aa3a47a9ad5d8aa9575325fb567c

Observation e0e5a2c1-0688-43dc-a5ca-ca7e2a97a114 · outbound

This paper cites CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T21:08:01.026456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:08:01.026456Z digest=sha256:fc54ce6df117f05781da7c84fd7dee00651aeb89a9e2eaed31cd1863910d4414

Observation 5c44f5aa-1795-402b-9ad5-dc08f2fe0155 · outbound

This paper cites Adding Conditional Control to Text-to-Image Diffusion Models.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models Adding Conditional Control to Text-to-Image Diffusion Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T21:08:01.108535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:08:01.108535Z digest=sha256:3dd7f2267a5f56ccb71e1ad78449e7e7ef808faec6863f5282914d5ebc462fcf

Observation ff129d0e-4250-4f2f-bd72-7f41f79b9627 · outbound

This paper cites Layercraft: Enhancing text-to-image generation with cot reasoning and layered object integration, 2025.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models Layercraft: Enhancing text-to-image generation with cot reasoning and layered object integration, 2025

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T21:08:01.186249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:08:01.186249Z digest=sha256:c6c810ea1187f9dce231a2e0375e23d53dd26d48e3000b909303f4edd45f86ab

Observation 16cddb43-7088-4ddf-b420-4df7f2efdd5b · outbound

This paper cites Divide and Conquer: Language Models can Plan and Self-Correct for Compositional Text-to-Image Generation.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models Divide and Conquer: Language Models can Plan and Self-Correct for Compositional Text-to-Image Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T21:08:01.263081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:08:01.263081Z digest=sha256:cbe9ef7051a2bd23bce52b02b1bc71264bc2ae993052395ede6ec73141a85e02

Observation 0a32ab58-fd34-4b37-8091-5a0deda9eb7e · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T21:08:01.413854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:08:01.413854Z digest=sha256:593158f01b6d6dfbd63df88598260ac9c9a1b2dcfea77f744307370a2fe2b858

Observation 7ad1bacd-2c8a-45aa-b0da-bedc64957a99 · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models Emu: Generative Pretraining in Multimodality

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T21:08:00.285957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:08:00.285957Z digest=sha256:43fd0ecbbdcc2b90a433f1b5897c28c9434288aea56cbd1bba586abbbf3ea849

Pith citing papers

No inbound Pith citation observations are available.