Pith. sign in

Paper Citation Record · LEDGER

Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 52 inbound Pith citation observations for arXiv:2409.10695.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.10695 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 52 of 52 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:48:43.723577Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T05:54:33.573861Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 980e8f8d-e3ac-469e-800e-3510a6600510 · inbound

Emu3: Next-Token Prediction is All You Need cites this paper.

Emu3: Next-Token Prediction is All You Need Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:56:09.215592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T10:56:06.418360Z digest=sha256:815d330d5eaa88c60a4542d7ccb0d5ee49b8ac628034594b3e911d5801ec44b2

Observation 07ee6c8e-8d94-4799-a876-ac6f425c9e61 · inbound

SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers cites this paper.

SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:56:50.092999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T00:56:50.009149Z digest=sha256:58c94ea606ef52c5b7122e729c10a113ed8459793611f2de009b9c8b438fb98d

Observation 2977e68d-84b7-41bd-ada7-f78dfcb54bca · inbound

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control cites this paper.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:38:24.494101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:c9cf286dc6e847dc5815c41622fa2c7e72b1c37bb5a0db5d424458e39284fe67

Observation 01114a68-abbe-4329-91d7-036312b2936f · inbound

Open-Sora Plan: Open-Source Large Video Generation Model cites this paper.

Open-Sora Plan: Open-Source Large Video Generation Model Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-23T08:42:45.153879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T08:38:27.946746Z digest=sha256:7aa76f077aaa5532347ab6068276d13736850b79ecb3519e7e9c3cab950354a8

Observation 355e14f9-495f-4100-942e-305553f47ac2 · inbound

IQA-Adapter: Exploring Knowledge Transfer from Image Quality Assessment to Diffusion-based Generative Models cites this paper.

IQA-Adapter: Exploring Knowledge Transfer from Image Quality Assessment to Diffusion-based Generative Models Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T01:00:03.278624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T01:00:03.278624Z digest=sha256:030a0cee6dadcc0327405f905101544c4b0f9c218dcd7dd82bc73822b7ea713e

Observation 142d523d-929d-47ef-826e-9084208b1384 · inbound

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models cites this paper.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.636337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.636337Z digest=sha256:6feed43f234ad9afd00a1fb5ba000c4a75090d330346d3475b1fea06721f3156

Observation a114ed54-6af2-4595-9c78-ef507dfa1625 · inbound

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation cites this paper.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T22:05:33.785966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:05:33.785966Z digest=sha256:79b464026c7dcfed2948fe7663616c77a761e7b6ac1ac0c5144c876f12cbfdf6

Observation d68d6a65-9d14-42f4-86ce-3d4ff418d0bb · inbound

Chimera: Improving Generalist Model with Domain-Specific Experts cites this paper.

Chimera: Improving Generalist Model with Domain-Specific Experts Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:47.588552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:47.588552Z digest=sha256:10e53e633393e32bb31c7832721fca974268bc34c5776135925c79574e014e95

Observation cdcbe4cc-da18-4798-bece-32dd43acc6ff · inbound

EasyRef: Omni-Generalized Group Image Reference for Diffusion Models via Multimodal LLM cites this paper.

EasyRef: Omni-Generalized Group Image Reference for Diffusion Models via Multimodal LLM Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T16:55:34.988363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:55:34.988363Z digest=sha256:f9fa25b0715ab025abb76a91031347ba60ada6652a328b33689a146d7f25ab23

Observation 1a8ebef6-b50d-4aee-ab40-3d7f7d1f899e · inbound

SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training cites this paper.

SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:01.711945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:01.711945Z digest=sha256:10cf7fb9779588c9bc1c3ea7453152feadf776383332a007496e212237b256fc

Observation a64fa0b6-6fba-4423-8591-c0e820380f6f · inbound

Autoregressive Video Generation without Vector Quantization cites this paper.

Autoregressive Video Generation without Vector Quantization Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:07:39.859781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:0e746ce4f54348865ddbf5bccd1bbfdb9318bf27cefb216b4f1ac43eb9ba6c5e

Observation 2b7d4d9a-5f22-48e9-8f3c-87bf35b52059 · inbound

1.58-bit FLUX cites this paper.

1.58-bit FLUX Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T04:39:36.630126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:39:36.630126Z digest=sha256:b816c78bcb99b3c6a84f357ce8ebc1cfed18782d67363c58134fe8a039a3625d

Observation 1a4f5851-5d71-4b49-9184-deeaefb6e187 · inbound

UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image Generation cites this paper.

UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image Generation Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T04:24:20.039877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:24:20.039877Z digest=sha256:45a43772c98bf0e813035b75723fb8bd790dcee8fae342a650d0f734c6d8d2fe

Observation 8a02f3cc-7f46-4de5-8ec3-04ece55250d7 · inbound

Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models cites this paper.

Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:34:14.355306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:34:14.355306Z digest=sha256:cc1b3d734729f18f4b91013971e685b190c28ac8b2d19b2cbbaad51b5299159b

Observation dac03e54-6e18-4ed8-ae35-d7cff5977a2c · inbound

AnyStory: Towards Unified Single and Multiple Subject Personalization in Text-to-Image Generation cites this paper.

AnyStory: Towards Unified Single and Multiple Subject Personalization in Text-to-Image Generation Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T20:01:21.490589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:01:21.490589Z digest=sha256:daee7e2e90ffe0e049f8013fa4bf2f55c1f50fa2ee1aad40c63f8e881f921492

Observation 39954511-0346-4579-987d-00fb6dd8ff98 · inbound

MSF: Efficient Diffusion Model Via Multi-Scale Latent Factorize cites this paper.

MSF: Efficient Diffusion Model Via Multi-Scale Latent Factorize Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T16:18:12.220781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:18:12.220781Z digest=sha256:b0bf8aa6efc42a91690c49c6e343c3ae19b640788c8face17d1787ed052a7a68

Observation d6399c26-8129-41ab-b7d8-37bb20ad8fcb · inbound

Diffusion Generative Modeling for Spatially Resolved Gene Expression Inference from Histology Images cites this paper.

Diffusion Generative Modeling for Spatially Resolved Gene Expression Inference from Histology Images Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T14:10:43.637061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:10:43.637061Z digest=sha256:427f02f268d8feb3f8c26b85407fd4e2c1a3fb20a58a0f0381e92a62235180bb

Observation e88ed9b2-121a-47e5-9f2c-726b37ef1b9e · inbound

SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer cites this paper.

SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T23:37:12.431235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T23:37:12.431235Z digest=sha256:1f2c614109ab9fb76a4ba459740193a4322e24f67a1fdb38e81f7841dd66c3aa

Observation 4ed10bf5-00e1-4e10-b85f-b2740dd74a1c · inbound

I Think, Therefore I Diffuse: Enabling Multimodal In-Context Reasoning in Diffusion Models cites this paper.

I Think, Therefore I Diffuse: Enabling Multimodal In-Context Reasoning in Diffusion Models Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T10:24:41.559300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:24:41.559300Z digest=sha256:8da3adfde87f46b70b5587157c8ab85120f58425210afc8bc5f729af62230d09

Observation cd739ac7-a0df-484f-898e-50b74e96a432 · inbound

RepText: Rendering Visual Text via Replicating cites this paper.

RepText: Rendering Visual Text via Replicating Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T05:48:43.723577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:48:43.723577Z digest=sha256:82f5e7b460b8c86df34c99ba736ca5ca5943104ff9d21be5cb234d40c0966ee2

Observation ddee4606-d36e-48aa-820d-b6ad6d242bd4 · inbound

X-Fusion: Introducing New Modality to Frozen Large Language Models cites this paper.

X-Fusion: Introducing New Modality to Frozen Large Language Models Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T05:18:15.277537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:18:15.277537Z digest=sha256:bc1fcf85f84360663d9a3ef8d046c58d0ff00e5d63c9da9d62d636882107acd5

Observation d548729c-5054-4608-bfba-5b440b84d94b · inbound

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis cites this paper.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T21:21:22.400326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:21:22.400326Z digest=sha256:5d8cc63d3ef7b75662bbfe053f679089c96c87b9416232c977921a29cabc250b

Observation 2f33ea93-ae88-49b2-a526-072b95ab80fc · inbound

Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation cites this paper.

Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:13.514247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:13.514247Z digest=sha256:696fb52ea74f67d4ac9afe7cdd4282738252706b90030f6f417f130dc056564d

Observation 5b7f518a-5f1d-4917-824d-cf9e5d79b6bf · inbound

RePrompt: Reasoning-Augmented Reprompting for Text-to-Image Generation via Reinforcement Learning cites this paper.

RePrompt: Reasoning-Augmented Reprompting for Text-to-Image Generation via Reinforcement Learning Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:49:04.771360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:49:04.771360Z digest=sha256:2b5360d490e1de85c98e1aafea5cae3c9d4a607636f32500f29a21062fbe5e5a

Observation 5c9985a9-f60c-436a-9496-d462c89a3dda · inbound

Normalized Attention Guidance: Universal Negative Guidance for Diffusion Models cites this paper.

Normalized Attention Guidance: Universal Negative Guidance for Diffusion Models Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:42:36.902137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:42:36.902137Z digest=sha256:fef5e7a2ffa1fbc256428d745b43577f4590fdcd91d83b4a48b1f9affc0313e5

Observation 25002f22-4f2d-4fbf-9222-f49370bb2c88 · inbound

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction cites this paper.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:04.804019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:04.804019Z digest=sha256:8a5fd67e360ccfd364c572534c9235120c34b29b828d19b38fb89f88e15b95b7

Observation 715e4c1a-801f-4b0a-ba64-1693121f496e · inbound

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation cites this paper.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:05.725545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:05.725545Z digest=sha256:bd85cc67a9b2bcc40be558aa4f93f9b69337e88700397830653e5bad2a680760

Observation b09a834a-4aeb-4896-b560-eb95b88b368c · inbound

Ultra-High-Resolution Image Synthesis: Data, Method and Evaluation cites this paper.

Ultra-High-Resolution Image Synthesis: Data, Method and Evaluation Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:50:42.929835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:50:42.929835Z digest=sha256:ccf7b274b7e6a63123fd39265baa4fe0af18586da1733a7b25e4839d5a2b83c9

Observation 7c750e13-1c5c-4021-aa34-6f2c99e27283 · inbound

UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation cites this paper.

UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T17:34:27.139415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T17:34:26.951644Z digest=sha256:bbe7d4b65632f6f94dd170fe637541d30b4eb1ccad606b0da3787fef7739dc16

Observation 78c10748-bbdd-4373-b842-1ba1ad57c327 · inbound

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation cites this paper.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.115963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.115963Z digest=sha256:fe092bfc1f350bb0a0cb2ae0e5cefeddc7611161f5a3cb525ecb0fb128cec7f3

Observation d796aa71-2a65-4bfb-8a60-b5db591ac07a · inbound

PosterCraft: Rethinking High-Quality Aesthetic Poster Generation in a Unified Framework cites this paper.

PosterCraft: Rethinking High-Quality Aesthetic Poster Generation in a Unified Framework Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:42.752815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:42.752815Z digest=sha256:480dd562938bc1b3b9fb25caa439805072362cbdc157c7c2bda06b73d0ca255d

Observation 82e9cc49-9856-4f30-86c5-0eab0a8f9263 · inbound

CycleVAR: Repurposing Autoregressive Model for Unsupervised One-Step Image Translation cites this paper.

CycleVAR: Repurposing Autoregressive Model for Unsupervised One-Step Image Translation Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.492175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.492175Z digest=sha256:7166cf00b29c2f8482b94907b31f31ab997c1d4232d3e48457461c7829ee818f

Observation 7a3d0dcb-6129-494c-8159-78cd2e95eefa · inbound

DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid Tokenizer cites this paper.

DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid Tokenizer Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:23.713635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:23.713635Z digest=sha256:400bc42173ff07f35741622c93dcf697e7b45e6d59429bba92b7e68ba9d077df

Observation d8ca69b2-8667-48f6-b93e-685d041ce13e · inbound

Towards Evaluating Robustness of Prompt Adherence in Text to Image Models cites this paper.

Towards Evaluating Robustness of Prompt Adherence in Text to Image Models Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T18:50:43.618673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:50:43.618673Z digest=sha256:801bb84244b3051b4d2af8bdd2d7c694ac864d38e7d1a38d5044aaa655bc8713

Observation 43827c62-29e4-4125-abcf-d0cb4571445f · inbound

PoemTale Diffusion: Minimising Information Loss in Poem to Image Generation with Multi-Stage Prompt Refinement cites this paper.

PoemTale Diffusion: Minimising Information Loss in Poem to Image Generation with Multi-Stage Prompt Refinement Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T16:22:49.672354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:22:49.672354Z digest=sha256:c53ddbca6fb62681a9a6de584b4c6c39fc080f537b32c41d2fe0b4af162db9d2

Observation 6dfb72e3-ca11-4f38-bf33-48afc8e80cc3 · inbound

Evaluating Uncertainty and Quality of Visual Language Action-enabled Robots cites this paper.

Evaluating Uncertainty and Quality of Visual Language Action-enabled Robots Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:25.641058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:25.641058Z digest=sha256:6c5141aa3f23979031a9de27e81373232ae9b0d2e78defa9f8fafcc9d4f16568

Observation 2799051a-0be3-4f43-b590-fc8bc919b5c3 · inbound

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning cites this paper.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:36.402785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:36.402785Z digest=sha256:8e641a2a6f008f2980a23d1d722aaed64b43ab45231f6e92e3e2829f4acd9b47

Observation 9a3cf6d0-8645-4517-8060-6e46deb526bf · inbound

Reconstruction Alignment Improves Unified Multimodal Models cites this paper.

Reconstruction Alignment Improves Unified Multimodal Models Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T22:36:08.047169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:36:08.047169Z digest=sha256:0167ab376d74221e5ac77df9dd411b8f1fcf9877bf997db879526eea3d74ac82

Observation 55a9c309-62ac-4231-b2ca-f588b283ca0e · inbound

PixelDiT: Pixel Diffusion Transformers for Image Generation cites this paper.

PixelDiT: Pixel Diffusion Transformers for Image Generation Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:31:31.192384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T04:30:07.417197Z digest=sha256:6a05dd01f47be05f2feb87ef695c31b3f90cd0e76c34e6a8db20a7a94f62968c

Observation 66178e1a-860b-4bc5-8d26-ebf20949ed5e · inbound

LTX-2: Efficient Joint Audio-Visual Foundation Model cites this paper.

LTX-2: Efficient Joint Audio-Visual Foundation Model Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:06:20.664484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T07:06:20.470686Z digest=sha256:ea3792ab581a7176c16ee4d25fdb5a23aa4f76f08e562a5b4d2d2ffa96d88650

Observation 32dab576-631f-471d-9d62-61a173be5449 · inbound

SnapGen++: Unleashing Diffusion Transformers for Efficient High-Fidelity Image Generation on Edge Devices cites this paper.

SnapGen++: Unleashing Diffusion Transformers for Efficient High-Fidelity Image Generation on Edge Devices Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T10:56:13.748789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:56:13.748789Z digest=sha256:870cd0e1dabecfb28a142345bbd7f50d83dacd1aa5dc19c9cb5119a001fcaaca

Observation fa2482c6-3010-442f-898e-3e3ff850af99 · inbound

WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing cites this paper.

WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T18:25:55.817521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:25:55.817521Z digest=sha256:bb79f4d1fe828a8356c5811c138086e1537bf7c5121b51e8d327d9422a4c5f2a

Observation 81dce721-b284-41cc-8f68-c2ec56106a31 · inbound

Self-Adversarial One Step Generation via Condition Shifting cites this paper.

Self-Adversarial One Step Generation via Condition Shifting Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:11:01.983027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T15:37:36.994169Z digest=sha256:83b1612ad33155476e8205352d2d4d487e51954bb6cee0a8853593569dbba7a3

Observation 6cdca740-2263-4059-b665-c36a3b372e93 · inbound

Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation cites this paper.

Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:28:39.561935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T05:15:22.907880Z digest=sha256:ee8577285593e55eada6657c24b7e4f9f82fcf70d20c838fe9a9c4fca56f3713

Observation 32535488-c446-47fe-a0f8-0ef98a2b0971 · inbound

ElasticDiT: Efficient Diffusion Transformers via Elastic Architecture and Sparse Attention for High-Resolution Image Generation on Mobile Devices cites this paper.

ElasticDiT: Efficient Diffusion Transformers via Elastic Architecture and Sparse Attention for High-Resolution Image Generation on Mobile Devices Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:03:43.913986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-20T20:00:27.987481Z digest=sha256:d0344183067126b35be609fe2ba1f9f066a4773c5edc4732aef5dccfe42fdd96

Observation c1a8333e-c3ff-4c31-9b29-2753e52145dc · inbound

PixVerve: Advancing Native UHR Image Generation to 100MP with a Large-Scale High-Quality Dataset cites this paper.

PixVerve: Advancing Native UHR Image Generation to 100MP with a Large-Scale High-Quality Dataset Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:23:03.899503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T05:19:30.372528Z digest=sha256:3eb633c0d02ab219298fdbd991881629a23a240b2058b930aa1fca93d15a1c63

Observation 08d28f16-20af-41e7-9b0c-9632f7f59757 · inbound

Rethinking Cross-Layer Information Routing in Diffusion Transformers cites this paper.

Rethinking Cross-Layer Information Routing in Diffusion Transformers Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:29:39.396388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T05:29:19.111071Z digest=sha256:12fc8f14f0ef86619023dcae1998f50aed59cdacda5c69089b70dea331cd4169

Observation dedc7eea-de42-41a8-a0c0-99e448820e77 · inbound

Rethinking Cross-Layer Information Routing in Diffusion Transformers cites this paper.

Rethinking Cross-Layer Information Routing in Diffusion Transformers Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:54:57.989038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T17:50:22.258541Z digest=sha256:6fa79ea55a05cd5f3e49e98a714fedcefc8e006bae0a742562def35c46c3109f

Observation 1fa295a3-67e7-41c1-b558-572677447c3a · inbound

MRT: Masked Region Transformer for Layered Image Generation and Editing at Scale cites this paper.

MRT: Masked Region Transformer for Layered Image Generation and Editing at Scale Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:33:50.289517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T18:32:37.613906Z digest=sha256:9a735941c057b329e475d2a1119b8aaf7e0ed7a8c0732e8ae070b65305b19aac

Observation 768b2830-16b2-463c-95a0-c5816281b7a7 · inbound

OctoT2I: A Self-Evolving Agentic Text-to-Image Router cites this paper.

OctoT2I: A Self-Evolving Agentic Text-to-Image Router Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:26:22.528525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T14:22:57.822123Z digest=sha256:9e6a918e61fbed1c9ed7f9ec7ad552ae5ae0642a701c64da3b9218b411d1deaa

Observation 45069593-a195-4503-8515-5c61b7404054 · inbound

Token-to-Token Alignment of Text Embeddings for Semantic Blending cites this paper.

Token-to-Token Alignment of Text Embeddings for Semantic Blending Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:39:46.159754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T08:34:23.842302Z digest=sha256:781fd5c6696d11a63177bcac6894eb822cc59a84da326be8bf6649479fcc5f39

Observation e9ae30cd-31e1-49b4-b55f-ea999ddbf0a9 · inbound

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders cites this paper.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T05:54:33.575294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:b972c777bd8995acb1aa960fd9113431ac77c964cf9298285ebe6cd01cb93473