Pith. sign in

Paper Citation Record · LEDGER

Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 49 inbound Pith citation observations for arXiv:2409.10695.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.10695 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 49 of 49 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T01:00:03.278624Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T05:54:33.573861Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 980e8f8d-e3ac-469e-800e-3510a6600510 · inbound

Emu3: Next-Token Prediction is All You Need cites this paper.

Emu3: Next-Token Prediction is All You Need Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:56:09.215592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T10:56:06.418360Z digest=sha256:17c689c54d4c10742af303b85833f647da1b7aca299df569ef134c32a095660a

Observation 07ee6c8e-8d94-4799-a876-ac6f425c9e61 · inbound

SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers cites this paper.

SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:56:50.092999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T00:56:50.009149Z digest=sha256:c00ff532a591a60203c61f72e40d03922685539a7754922e874ebac152da7175

Observation 2977e68d-84b7-41bd-ada7-f78dfcb54bca · inbound

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control cites this paper.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:38:24.494101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:b9c57de83f6f8bfb27450c74c94aed4bf85ca03198a5a62edee7d5cdcab53d87

Observation 01114a68-abbe-4329-91d7-036312b2936f · inbound

Open-Sora Plan: Open-Source Large Video Generation Model cites this paper.

Open-Sora Plan: Open-Source Large Video Generation Model Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-23T08:42:45.153879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T08:38:27.946746Z digest=sha256:d77fe6e4d5b8f1bedd05ecf222228abf2163427042e2586fcc6327de66ff4fbc

Observation 355e14f9-495f-4100-942e-305553f47ac2 · inbound

IQA-Adapter: Exploring Knowledge Transfer from Image Quality Assessment to Diffusion-based Generative Models cites this paper.

IQA-Adapter: Exploring Knowledge Transfer from Image Quality Assessment to Diffusion-based Generative Models Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T01:00:03.278624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T01:00:03.278624Z digest=sha256:24fe75c70502afef79757f54cd47bdc7b82b36693c52f8a5588915dc44a41d43

Observation 142d523d-929d-47ef-826e-9084208b1384 · inbound

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models cites this paper.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.636337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.636337Z digest=sha256:7a8c5a939e792958e17d5f61e0cc992683688f67e3fb5fd9a297deb15860162e

Observation a114ed54-6af2-4595-9c78-ef507dfa1625 · inbound

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation cites this paper.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T22:05:33.785966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:05:33.785966Z digest=sha256:4b6279daf4446f4260860803e13258ca63f70524fd424c08940bf3c6538dc8f1

Observation d68d6a65-9d14-42f4-86ce-3d4ff418d0bb · inbound

Chimera: Improving Generalist Model with Domain-Specific Experts cites this paper.

Chimera: Improving Generalist Model with Domain-Specific Experts Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:47.588552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:47.588552Z digest=sha256:2c3ce9b39180c0bcf4119b1890aeb4a9fb0464782aaea3d3563933b57d1212a5

Observation cdcbe4cc-da18-4798-bece-32dd43acc6ff · inbound

EasyRef: Omni-Generalized Group Image Reference for Diffusion Models via Multimodal LLM cites this paper.

EasyRef: Omni-Generalized Group Image Reference for Diffusion Models via Multimodal LLM Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T16:55:34.988363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:55:34.988363Z digest=sha256:6110f47cc43777e2b18c7930e705bce80f79e229dad744f7a01cad2456fa8671

Observation 1a8ebef6-b50d-4aee-ab40-3d7f7d1f899e · inbound

SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training cites this paper.

SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:01.711945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:01.711945Z digest=sha256:42ac0132f678745b0416a07b456e07c9fe8b8f4cf9b77a871b3a7a4fabb4a16b

Observation a64fa0b6-6fba-4423-8591-c0e820380f6f · inbound

Autoregressive Video Generation without Vector Quantization cites this paper.

Autoregressive Video Generation without Vector Quantization Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:07:39.859781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:06c2a7393f3ff1c80b69daa1a445b4154b4d0cd33dec6df3bff9f0fcab67c5a4

Observation 2b7d4d9a-5f22-48e9-8f3c-87bf35b52059 · inbound

1.58-bit FLUX cites this paper.

1.58-bit FLUX Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T04:39:36.630126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:39:36.630126Z digest=sha256:129e4af043c70d24978cd2bb2c712a99cd01ff7a9cf42f15daac69f0fd5dd95e

Observation 1a4f5851-5d71-4b49-9184-deeaefb6e187 · inbound

UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image Generation cites this paper.

UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image Generation Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T04:24:20.039877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:24:20.039877Z digest=sha256:4807711279d5e0fc4776d5d4e9ffad40561aa2d520d60f0b92aa13fceb9f0e1e

Observation 8a02f3cc-7f46-4de5-8ec3-04ece55250d7 · inbound

Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models cites this paper.

Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:34:14.355306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:34:14.355306Z digest=sha256:2411d8014ee7c71ff7e3b67fa3b163ac5f157b37e956d0cf10eb9c8076f30b45

Observation dac03e54-6e18-4ed8-ae35-d7cff5977a2c · inbound

AnyStory: Towards Unified Single and Multiple Subject Personalization in Text-to-Image Generation cites this paper.

AnyStory: Towards Unified Single and Multiple Subject Personalization in Text-to-Image Generation Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T20:01:21.490589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:01:21.490589Z digest=sha256:815338554025f93edd5e9b7d0e23cf24028349aa99eee72b6b03d6d78b5c08ff

Observation 39954511-0346-4579-987d-00fb6dd8ff98 · inbound

MSF: Efficient Diffusion Model Via Multi-Scale Latent Factorize cites this paper.

MSF: Efficient Diffusion Model Via Multi-Scale Latent Factorize Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T16:18:12.220781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:18:12.220781Z digest=sha256:6bcfa90d095ad8c4492a3ca129cd18d12a0ba50924fa16ea2daf3053e7c0e1a4

Observation d6399c26-8129-41ab-b7d8-37bb20ad8fcb · inbound

Diffusion Generative Modeling for Spatially Resolved Gene Expression Inference from Histology Images cites this paper.

Diffusion Generative Modeling for Spatially Resolved Gene Expression Inference from Histology Images Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T14:10:43.637061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:10:43.637061Z digest=sha256:fb7dfd1acd2189e13165fd531b2e785d0de9808c45d52aac4e652d0716168382

Observation e88ed9b2-121a-47e5-9f2c-726b37ef1b9e · inbound

SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer cites this paper.

SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T23:37:12.431235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T23:37:12.431235Z digest=sha256:d5c81f84b16c73e7256372a2428780a5afa93d361a3552a02832b50e2d4a0a49

Observation 4ed10bf5-00e1-4e10-b85f-b2740dd74a1c · inbound

I Think, Therefore I Diffuse: Enabling Multimodal In-Context Reasoning in Diffusion Models cites this paper.

I Think, Therefore I Diffuse: Enabling Multimodal In-Context Reasoning in Diffusion Models Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T10:24:41.559300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:24:41.559300Z digest=sha256:cb549ec8186a6a483baf19534851ca1fe382d48a5658a79e4b13e8e6a6fe5165

Observation 2f33ea93-ae88-49b2-a526-072b95ab80fc · inbound

Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation cites this paper.

Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:13.514247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:13.514247Z digest=sha256:5574b74bf28667e934747c48f06db91ee193f199361cc94bc4c25e22809b53ae

Observation 5b7f518a-5f1d-4917-824d-cf9e5d79b6bf · inbound

RePrompt: Reasoning-Augmented Reprompting for Text-to-Image Generation via Reinforcement Learning cites this paper.

RePrompt: Reasoning-Augmented Reprompting for Text-to-Image Generation via Reinforcement Learning Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:49:04.771360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:49:04.771360Z digest=sha256:ce3548c4045af03562eb201997d490ed786cdb223baf0bddb1d00ef3f6450fad

Observation 5c9985a9-f60c-436a-9496-d462c89a3dda · inbound

Normalized Attention Guidance: Universal Negative Guidance for Diffusion Models cites this paper.

Normalized Attention Guidance: Universal Negative Guidance for Diffusion Models Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:42:36.902137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:42:36.902137Z digest=sha256:0bd313e1e20010c906b3d7e15e408d50cd907063ce43b1f55ebd6536423cf52b

Observation 25002f22-4f2d-4fbf-9222-f49370bb2c88 · inbound

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction cites this paper.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:04.804019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:04.804019Z digest=sha256:04ca1d1faa20f6a57214e66c1eb3525d6f36adee1db405b841d581e281f3e842

Observation 715e4c1a-801f-4b0a-ba64-1693121f496e · inbound

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation cites this paper.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:05.725545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:05.725545Z digest=sha256:3439e9e596458401ff3cc0f8c0c5cf17b6d2072e7c56a160b1304442db155b24

Observation b09a834a-4aeb-4896-b560-eb95b88b368c · inbound

Ultra-High-Resolution Image Synthesis: Data, Method and Evaluation cites this paper.

Ultra-High-Resolution Image Synthesis: Data, Method and Evaluation Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:50:42.929835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:50:42.929835Z digest=sha256:693cace2bd7a40ffa9b13e6d59e6f5e565447592cc75ad5235dd298cdd499532

Observation 7c750e13-1c5c-4021-aa34-6f2c99e27283 · inbound

UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation cites this paper.

UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T17:34:27.139415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T17:34:26.951644Z digest=sha256:b218af17ea74a43bbfce18720ed5b688da5044c092aedbec0badbcb7f7fe0d86

Observation 78c10748-bbdd-4373-b842-1ba1ad57c327 · inbound

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation cites this paper.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.115963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.115963Z digest=sha256:ae656691343bc9d6a41eda2a836a62b1288893dc671b2e406751bc863110e609

Observation d796aa71-2a65-4bfb-8a60-b5db591ac07a · inbound

PosterCraft: Rethinking High-Quality Aesthetic Poster Generation in a Unified Framework cites this paper.

PosterCraft: Rethinking High-Quality Aesthetic Poster Generation in a Unified Framework Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:42.752815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:42.752815Z digest=sha256:887d399e30ec33253433f3e4871dc8d12a7487817ddc34d243d9369cc2fadf68

Observation 82e9cc49-9856-4f30-86c5-0eab0a8f9263 · inbound

CycleVAR: Repurposing Autoregressive Model for Unsupervised One-Step Image Translation cites this paper.

CycleVAR: Repurposing Autoregressive Model for Unsupervised One-Step Image Translation Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.492175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.492175Z digest=sha256:26bce4663e6075b01178771334de079741ce0dfbaa920fb1f494612d103a5503

Observation 7a3d0dcb-6129-494c-8159-78cd2e95eefa · inbound

DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid Tokenizer cites this paper.

DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid Tokenizer Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:23.713635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:23.713635Z digest=sha256:e71b311eb2f0bc4142c6f6eeb109ebd58fb8f53261cc67affa92e93555c97cd4

Observation d8ca69b2-8667-48f6-b93e-685d041ce13e · inbound

Towards Evaluating Robustness of Prompt Adherence in Text to Image Models cites this paper.

Towards Evaluating Robustness of Prompt Adherence in Text to Image Models Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T18:50:43.618673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:50:43.618673Z digest=sha256:4e6e8f42f73999777423fd3b38a307e780f6f5888eebfacbeb54779c98948bf6

Observation 43827c62-29e4-4125-abcf-d0cb4571445f · inbound

PoemTale Diffusion: Minimising Information Loss in Poem to Image Generation with Multi-Stage Prompt Refinement cites this paper.

PoemTale Diffusion: Minimising Information Loss in Poem to Image Generation with Multi-Stage Prompt Refinement Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T16:22:49.672354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:22:49.672354Z digest=sha256:d86702b24ddad0f0568d467e20214e1fe4201a09305d8d8fe31cfffe2e24ce02

Observation 6dfb72e3-ca11-4f38-bf33-48afc8e80cc3 · inbound

Evaluating Uncertainty and Quality of Visual Language Action-enabled Robots cites this paper.

Evaluating Uncertainty and Quality of Visual Language Action-enabled Robots Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:25.641058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:25.641058Z digest=sha256:83ac4cb07792d4206c189c6b2725fd147dab3b6b192a4dfcae9eec79b1b3f40d

Observation 2799051a-0be3-4f43-b590-fc8bc919b5c3 · inbound

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning cites this paper.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:36.402785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:36.402785Z digest=sha256:fedf065fa59895d397784070ecde5cb5473958fa8eebef4332a6656350036007

Observation 9a3cf6d0-8645-4517-8060-6e46deb526bf · inbound

Reconstruction Alignment Improves Unified Multimodal Models cites this paper.

Reconstruction Alignment Improves Unified Multimodal Models Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T22:36:08.047169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:36:08.047169Z digest=sha256:2a4f872ecba657ae0e029053306291eb41195fc3b32dab708550ada96c95dcb0

Observation 55a9c309-62ac-4231-b2ca-f588b283ca0e · inbound

PixelDiT: Pixel Diffusion Transformers for Image Generation cites this paper.

PixelDiT: Pixel Diffusion Transformers for Image Generation Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:31:31.192384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T04:30:07.417197Z digest=sha256:4ff85d466f9ea3abfd70bd046aad5359bf1ede8ded32474b5f9a2cd4d246e879

Observation 66178e1a-860b-4bc5-8d26-ebf20949ed5e · inbound

LTX-2: Efficient Joint Audio-Visual Foundation Model cites this paper.

LTX-2: Efficient Joint Audio-Visual Foundation Model Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:06:20.664484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T07:06:20.470686Z digest=sha256:e2f594d54bdcfa8b2e7ce48092d555d60d5a10a7fb7e3bf0904da43a9534fe80

Observation 32dab576-631f-471d-9d62-61a173be5449 · inbound

SnapGen++: Unleashing Diffusion Transformers for Efficient High-Fidelity Image Generation on Edge Devices cites this paper.

SnapGen++: Unleashing Diffusion Transformers for Efficient High-Fidelity Image Generation on Edge Devices Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T10:56:13.748789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:56:13.748789Z digest=sha256:2d24bf54c9c7b6a03ac617d7c385fdabb8a56976db963aed2b196bdffc1127b5

Observation fa2482c6-3010-442f-898e-3e3ff850af99 · inbound

WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing cites this paper.

WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T18:25:55.817521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:25:55.817521Z digest=sha256:c4040d397c27d0ed36f6e570ca271eb164c1785fd630300abd67aac630121065

Observation 81dce721-b284-41cc-8f68-c2ec56106a31 · inbound

Self-Adversarial One Step Generation via Condition Shifting cites this paper.

Self-Adversarial One Step Generation via Condition Shifting Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:11:01.983027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T15:37:36.994169Z digest=sha256:195b9e264b12eb9c18233bdcc494c523b8c30d8a101deff69f401103e0ffcf20

Observation 6cdca740-2263-4059-b665-c36a3b372e93 · inbound

Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation cites this paper.

Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:28:39.561935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T05:15:22.907880Z digest=sha256:0a45f8e7d5fd862de8c4122cbe39871e9fad8a98510367b37fef3c5272f942b0

Observation 32535488-c446-47fe-a0f8-0ef98a2b0971 · inbound

ElasticDiT: Efficient Diffusion Transformers via Elastic Architecture and Sparse Attention for High-Resolution Image Generation on Mobile Devices cites this paper.

ElasticDiT: Efficient Diffusion Transformers via Elastic Architecture and Sparse Attention for High-Resolution Image Generation on Mobile Devices Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:03:43.913986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-20T20:00:27.987481Z digest=sha256:b3885300c98a362800fcba7124fdbf7d256d21f8404a5dbba1f554557f2512bc

Observation c1a8333e-c3ff-4c31-9b29-2753e52145dc · inbound

PixVerve: Advancing Native UHR Image Generation to 100MP with a Large-Scale High-Quality Dataset cites this paper.

PixVerve: Advancing Native UHR Image Generation to 100MP with a Large-Scale High-Quality Dataset Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:23:03.899503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T05:19:30.372528Z digest=sha256:f429b2dc268712ab931c7c747c077c35fa3ac45a5d7ab068b6b118afa684aa98

Observation 08d28f16-20af-41e7-9b0c-9632f7f59757 · inbound

Rethinking Cross-Layer Information Routing in Diffusion Transformers cites this paper.

Rethinking Cross-Layer Information Routing in Diffusion Transformers Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:29:39.396388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T05:29:19.111071Z digest=sha256:ce09b30df14f7aca7f1605da74da1b996acbbbbb53e33fbf4067c0f62fd797ef

Observation dedc7eea-de42-41a8-a0c0-99e448820e77 · inbound

Rethinking Cross-Layer Information Routing in Diffusion Transformers cites this paper.

Rethinking Cross-Layer Information Routing in Diffusion Transformers Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:54:57.989038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T17:50:22.258541Z digest=sha256:f8b47ddeb6bf517ccae08de3043b03a67ee2b87d680fdfc2c5d66b3b1307ab7f

Observation 1fa295a3-67e7-41c1-b558-572677447c3a · inbound

MRT: Masked Region Transformer for Layered Image Generation and Editing at Scale cites this paper.

MRT: Masked Region Transformer for Layered Image Generation and Editing at Scale Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:33:50.289517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T18:32:37.613906Z digest=sha256:3ac6da2594f7cdbd38063b76d25228fb46b6907c38bd41278a8134af5261d20a

Observation 768b2830-16b2-463c-95a0-c5816281b7a7 · inbound

OctoT2I: A Self-Evolving Agentic Text-to-Image Router cites this paper.

OctoT2I: A Self-Evolving Agentic Text-to-Image Router Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:26:22.528525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T14:22:57.822123Z digest=sha256:5dbe862abc2ba4d2eb503846c6e67c9d9f85e0ef2ea2d53ff045f1b9371305b3

Observation 45069593-a195-4503-8515-5c61b7404054 · inbound

Token-to-Token Alignment of Text Embeddings for Semantic Blending cites this paper.

Token-to-Token Alignment of Text Embeddings for Semantic Blending Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:39:46.159754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T08:34:23.842302Z digest=sha256:52e26fb27dda64078d645189687f402b9ecb80c636f39b0c8faec41aa412b9bb

Observation e9ae30cd-31e1-49b4-b55f-ea999ddbf0a9 · inbound

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders cites this paper.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T05:54:33.575294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:368af32d96bd78589f2b6c6e524640da7f709b2bf9597365a52070815ff73baf