Pith. sign in

Paper Citation Record · LEDGER

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation

As of 8 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2506.08210.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08210 v1

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:20:51.244779Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

67 of 67 outbound references displayed

  • verified exact1
  • verified fuzzy29
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 398c47e8-fb10-4e8a-a0e9-f818b192f0c9 · outbound

This paper cites Edify Image: High-Quality Image Generation with Pixel Space Laplacian Diffusion Models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Edify Image: High-Quality Image Generation with Pixel Space Laplacian Diffusion Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:50.977218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:50.977218Z digest=sha256:aee34e56237f2149bbd8d8d5cdf4d512e9cfa4fbd84439e7b3f8a4db6f1e3576

Observation aa460d2e-26ff-4f00-8a19-b5c09a34d2db · outbound

This paper cites eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:50.982334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:50.982334Z digest=sha256:995ab30468dc85e1e635cb7c3246ae4ea094df9501898c1f42518acde95bc6cc

Observation 61d7c6d3-68f9-41d1-a5cb-65495b3ba30f · outbound

This paper cites Imagen 3.arXiv preprint arXiv:2408.07009, 2024.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Imagen 3.arXiv preprint arXiv:2408.07009, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:50.986447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:50.986447Z digest=sha256:cc7071171c5a216c40b5bb414f58521364bb71a715ca02279364209152933e9b

Observation 1634a2a4-269c-4570-bf60-f9799e1fad0e · outbound

This paper cites LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:50.990041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:50.990041Z digest=sha256:89897a0dc6d4793125e1b6e2dad63a51ac36da025245cb0f7387c72c0a8b1745

Observation e697754a-d932-4d75-9ef7-6b3d7692d37a · outbound

This paper cites Scalable Performance Analysis for Vision-Language Models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Scalable Performance Analysis for Vision-Language Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:20:51.597266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:20:50.994543Z digest=sha256:115334919223ef3338ea30866639eaaf12776fd0f82fa6dc30bcd01c56bb12a1

Observation 9b84a85a-28d1-424a-ae5d-98817c55b024 · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:50.998688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:50.998688Z digest=sha256:7fe9c8a903bddc6e79b30bee6ea9e233732d7cd4b02be94dfb1007aeeb284e2d

Observation c575c6f6-6984-4190-8a71-a26d056f684d · outbound

This paper cites Textdiffuser-2: Unleashing the power of language models for text rendering.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Textdiffuser-2: Unleashing the power of language models for text rendering

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:52.027527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:20:51.006114Z digest=sha256:e625eeb416a03a613b716d9a6f13fb6b2fe29dfd656e4f6fb39af53da23b804a

Observation 853b019a-3335-47a9-8701-ce431034a2e2 · outbound

This paper cites M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.010211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.010211Z digest=sha256:3d37b2db369ebae44cb2c35ea5fd66e7cecc7b6ff74b14601234f86eb021baf8

Observation c497bdec-1a28-48ba-876e-9ebb948eda0c · outbound

This paper cites Visual pro- gramming for step-by-step text-to-image generation and evaluation.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Visual pro- gramming for step-by-step text-to-image generation and evaluation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:52.015758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:20:51.014921Z digest=sha256:1a84869f1e9939cec9fb188eeef756104498eedefa3cf79eb73c377883497c0a

Observation 23195d1b-b58d-4436-8aed-e94a525398a1 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:52.004536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:20:51.018690Z digest=sha256:c54349c0f28d78cbc8363029f07769fd35cd204360bfd122a51cfcf183e91e6b

Observation 9924718c-6b1b-4cff-a924-0ff9529cfff8 · outbound

This paper cites Analyzing Transformers in Embedding Space.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Analyzing Transformers in Embedding Space

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.022450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.022450Z digest=sha256:f209facff36063b240f047afd78d51387904081fd79fdb268f974afaabf8dbe5

Observation e5739900-abb2-4b2c-a8d6-cef0fd7ec774 · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.993821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:20:51.027375Z digest=sha256:674b8b2ffe559ff8d4295553231dc237f098590015e742c61b95238a67cf5b20

Observation becaf3ad-2e97-42fc-9820-e757bec27070 · outbound

This paper cites Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.030846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.030846Z digest=sha256:202f75d5c48424d03435a8febe34d6a9fd49f3d51f0a56d8b8678b9d26aa03d2

Observation 941aed29-e666-4cc1-ae4c-391ecdcb3f49 · outbound

This paper cites Visual fact checker: en- abling high-fidelity detailed caption generation.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Visual fact checker: en- abling high-fidelity detailed caption generation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.982389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:20:51.034740Z digest=sha256:42367694dc2e090d4f7d985bb6d20561edfc3e1e33bcc57c7198adfbe47e5658

Observation c4abb752-e600-45c3-8439-85cbfcb3b568 · outbound

This paper cites The Llama 3 Herd of Models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation The Llama 3 Herd of Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.038814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.038814Z digest=sha256:e70e260a2b9159902d4c4739ce684494a11defcd12a835a6ee11cb7359d66092

Observation fbc237b3-b337-479e-8e9a-c9f999690795 · outbound

This paper cites MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image Synthesis.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image Synthesis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.042386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.042386Z digest=sha256:f253fb790fc7fc58282ca87b758720d9212124408e7a10c2102ef7f48c747119

Observation ea91496f-857e-48f3-9b51-e200fe3e3686 · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.046257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.046257Z digest=sha256:dc01134cbd4500f95565f2debaebf88f369b97dbb55f6f0d83dbaa09dd953b32

Observation 6d8bb99b-0573-4499-ae53-3369c5a78347 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.972571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:20:51.051035Z digest=sha256:79cda74116275a1ffbe69ecc9d7c0b18d8dbb3d6844f1ff77bf06b887114fcb2

Observation 039cbba8-4e40-41f8-b62d-971cdaf3469e · outbound

This paper cites Classifier-Free Diffusion Guidance.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Classifier-Free Diffusion Guidance

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.054665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.054665Z digest=sha256:bde17805267285cba1952742036d6991931c2bd416b6231e22c3cc2bb98bfb12

Observation 927c9c3c-c01b-44c8-bdbb-8c2e2d661b09 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.058500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.058500Z digest=sha256:14d0b131475c8c8637a404a2503a26b6a85381348eb3556a0f5c1618ee1edd41

Observation 6f26b31b-c794-4861-bf2b-2305a5e532db · outbound

This paper cites Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.961439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:20:51.062379Z digest=sha256:ae26824a4c948bae9ec6ef503e1aea8adcce9edbf50ef06d93854cbde54b4b40

Observation 87b63063-c52e-4bc7-bce0-91687dd6dc4d · outbound

This paper cites T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.950410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:20:51.066724Z digest=sha256:8e904c03ba93ff133f10c3ed5c1a17b27a6dc42d9401a24aa0033c0570869149

Observation 23a02f78-89ab-48d2-941a-1acf24767390 · outbound

This paper cites GPT-4o System Card.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation GPT-4o System Card

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.070312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.070312Z digest=sha256:886a69f6eff5e0db48130a7d4d505a45933b5c4ebe81ed5f790b5d0f5e2ada93

Observation 5f9e4d79-9629-4745-bedf-a538a28a0496 · outbound

This paper cites What does bert learn about the structure of language? InACL,.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation What does bert learn about the structure of language? InACL,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.938958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:20:51.074586Z digest=sha256:0b7537a1b80799eab15afc19ec53c2bfc515d9abdce9fa2ff1290261d3d69c72

Observation a744e85c-dbf0-40ea-8e1b-74831bdc6daa · outbound

This paper cites Mistral 7B.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Mistral 7B

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.078207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.078207Z digest=sha256:7d7f7d991a9c1eab0ec09f4f4f90669636757e70ac592a13de571c95b543b38c

Observation d6902ad0-3b3b-41c4-8c12-dd1929b6dc40 · outbound

This paper cites Analyzing the Role of Semantic Representations in the Era of Large Language Models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Analyzing the Role of Semantic Representations in the Era of Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.082207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.082207Z digest=sha256:eb4c49abcf61b3732ef2f701e7c710c932676369949bc7c7218f21aba1943f39

Observation 79b400d3-f8ac-4ef6-9dff-60a433fc1530 · outbound

This paper cites Elucidating the design space of diffusion-based generative models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Elucidating the design space of diffusion-based generative models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.927459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:20:51.086581Z digest=sha256:733f6ba799d98364d7ba65d06af5dd82e3539e2a25bf399ba6d838ff8e7ab9b0

Observation b44455d0-a6cc-4098-bd00-f884dce11158 · outbound

This paper cites Analyzing and improving the training dynamics of diffusion models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Analyzing and improving the training dynamics of diffusion models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.916207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:20:51.089829Z digest=sha256:71844509a6aaebcea5b07c6773b2541d550d34431d81c0157a875ab25590d68a

Observation 8f166b0a-9b9b-4618-8442-4e80d4ff018d · outbound

This paper cites NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.093052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.093052Z digest=sha256:62880740f79b0d569e0da63c5ed60e84ba5e4799071f7fbf10207ef86f362674

Observation 4093adbf-4e7f-4333-9802-e5abbf15e686 · outbound

This paper cites GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.096656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.096656Z digest=sha256:edb1e456b0b4c92d1b62408b652eb5e938f4ebc18428e330985f5314c0381715

Observation ce625a39-7946-45c8-88db-e4bd5b49427b · outbound

This paper cites Towards General Text Embeddings with Multi-stage Contrastive Learning.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Towards General Text Embeddings with Multi-stage Contrastive Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.100614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.100614Z digest=sha256:e0146cb63069e2da26a63d7b145fd57000de159c7528388c1bdd00c19f0d1971

Observation 6fe2513e-05dc-4728-b10f-6418a9231af6 · outbound

This paper cites Llm- grounded diffusion: Enhancing prompt understanding of text-to-image diffusion models with large language models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Llm- grounded diffusion: Enhancing prompt understanding of text-to-image diffusion models with large language models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.903817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:20:51.104499Z digest=sha256:ca0b61858a1f6c3572005cfcdf9fa3c081a38cc3ca9173c83a8db6882538e48e

Observation e25b3597-2306-45ac-b073-ecdae4731aad · outbound

This paper cites Common diffusion noise schedules and sample steps are flawed.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Common diffusion noise schedules and sample steps are flawed

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.892766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:20:51.108275Z digest=sha256:be4466fc5b4872f83846ac5dcd6edb07af21a7fadb4f80ef937c6d3d6236d086

Observation 04ead64d-cb04-459d-bb57-c2d51b0d4bf9 · outbound

This paper cites Evaluating text-to-visual generation with image-to-text gen- eration.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Evaluating text-to-visual generation with image-to-text gen- eration

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.882601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:20:51.112353Z digest=sha256:132df7de266a0e8d6d8658cbcea29bac8b2e8b3be1546e8840d28a6e2594f6bf

Observation 78c10748-bbdd-4373-b842-1ba1ad57c327 · outbound

This paper cites Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.115963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.115963Z digest=sha256:626d9b8fdfd36867f54638fd4d527e2225423bdc0801eb5f5a462bc6e48669a3

Observation 46685bd4-1272-4c9a-a0e6-f226a9a9bce3 · outbound

This paper cites Character-Aware Models Improve Visual Text Rendering.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Character-Aware Models Improve Visual Text Rendering

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.120173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.120173Z digest=sha256:5832e248283a54f8a750b784a55e1e67234b80572386b0def5527d971fdfacce

Observation 9494912b-8c09-407f-8ccd-0cbc9de6a8ff · outbound

This paper cites Fantastic Semantics and Where to Find Them: Investigating Which Layers of Generative LLMs Reflect Lexical Semantics.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Fantastic Semantics and Where to Find Them: Investigating Which Layers of Generative LLMs Reflect Lexical Semantics

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.123785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.123785Z digest=sha256:814ab2154f529394c168f8fe1a95462320afa8952d338bbe1042a0939cfe6b39

Observation f5a33c4b-94c4-443c-99bd-e73523f50273 · outbound

This paper cites Decoupled weight decay regularization, 2017.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Decoupled weight decay regularization, 2017

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.870995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:20:51.127292Z digest=sha256:503eddedb1b0b92672d63a31dd41cb2e58104ef2488bf63e6939335a0b30b8f9

Observation 5e1143d1-e17e-46dc-b0ad-588a8094bacf · outbound

This paper cites Salesforce AI Research’s SFR-embedding, the top performing text-embedding model.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Salesforce AI Research’s SFR-embedding, the top performing text-embedding model

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.860167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:20:51.130898Z digest=sha256:75caf352579434a79132a993eb270d65ae6bc339a63f32a45b8408b3771795dc

Observation 08adad73-ee57-430b-936f-99f7ffdda3e4 · outbound

This paper cites MTEB: Massive Text Embedding Benchmark.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation MTEB: Massive Text Embedding Benchmark

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.134528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.134528Z digest=sha256:3bb92a4277a20431bed5b6c804b0b7c39c714096c72072c25ecc3c7917ad8119

Observation b6d0ab25-1849-45fd-b079-deb6eeff7e55 · outbound

This paper cites Scalable diffusion models with transformers.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Scalable diffusion models with transformers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.139116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.139116Z digest=sha256:4665f243187a91f606e5f98bc0b94bbc92b234fd810a893de6e41da33e72c923

Observation 08aee82c-0c41-472e-9dcb-815af9af02ab · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.143069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.143069Z digest=sha256:ca06313a14186d419965d518f6387d6ffa84f31e9e01fd5fa7b46705c42ca64f

Observation be6f771f-b28b-44e4-b94d-fd8c5e1e1042 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Learn- ing transferable visual models from natural language super- vision

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.842843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:20:51.147265Z digest=sha256:a194320ff04f702b6bb1e7ebd714159f187f4c3fa962678693d6fc55a8fc1faa

Observation 30bb8684-6000-4e2c-a8f7-3ee077efd3ab · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.JMLR, 2020.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Exploring the limits of transfer learning with a unified text-to-text transformer.JMLR, 2020

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.832409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:20:51.150903Z digest=sha256:bab1d3dd0c1926741e8f44b3895195e83c9bb23ce664510ed09dae6c1b3c257c

Observation 8f990867-8353-45e4-9157-07fab43f9504 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.154397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.154397Z digest=sha256:7772bc6bd0d76bf4509db8dacbd61c6193c0f202dffabdc5ad50cf1c104fc6a5

Observation d242fceb-18b2-42b1-b1f2-e59bf4045b48 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation High-resolution image syn- thesis with latent diffusion models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.822625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:20:51.158439Z digest=sha256:587e65f211b17e3e3f929fea575af68aa935f0fbbd82042c5f0779cc92e33279

Observation cf875356-8584-4dfa-b89e-7565d2e257a6 · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation U-net: Convolutional networks for biomedical image segmentation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.811734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:20:51.161994Z digest=sha256:991316d4cc8d8ddf6e2c58c4b7c5d468007289315c19c2b61c9e252c5fc16357

Observation 1d1708d4-fdd0-4ca0-8d8e-d1275221830b · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Photorealistic text-to-image diffusion models with deep language understanding

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.800992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:20:51.165570Z digest=sha256:421a57db73f33214adf169bc1a876629e067e3c0729f1a19b1b6379cc3830941

Observation c2cf3cc2-a1d7-4f07-9be7-d6131df0fd71 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.790009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:20:51.169459Z digest=sha256:c66f380e9595740b2815e66de630d974f94a006c369802d970ac6a20eefbf526

Observation 554f8853-d4a0-49a4-9075-33be8ee69c87 · outbound

This paper cites Repetition Improves Language Model Embeddings.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Repetition Improves Language Model Embeddings

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.173283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.173283Z digest=sha256:a70271b1fa60ac61f95c849903660a77dbf2cbc441813eb2062e5a63d2b892ec

Observation 8f9e3754-0106-4a25-8f2d-4963b637f538 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Gemma 2: Improving Open Language Models at a Practical Size

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.181220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.181220Z digest=sha256:52602a6ef061aae662c65671a152327ffce798224d6fc88372803ec60ba13747

Observation ce1863b8-9a5e-46e5-a0ea-e282955f68fd · outbound

This paper cites Stable diffusion training with mo- saicml.Mosaic Research Blog, 2023.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Stable diffusion training with mo- saicml.Mosaic Research Blog, 2023

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.779633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:20:51.184771Z digest=sha256:6aaffd340d11d4f8b4174960169d5e11ffd36b1b95de231f30008aae52d5d582

Observation 650e3c80-0662-460e-a644-f33a1de5cf4b · outbound

This paper cites What do you learn from context? Probing for sentence structure in contextualized word representations.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation What do you learn from context? Probing for sentence structure in contextualized word representations

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.188280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.188280Z digest=sha256:6c22fa363bb94bf4a8c86c131ecbf606644b09fe656d3dc05af3ff58a602972f

Observation 8256558b-81e8-4136-ab26-a8ed29f1c08f · outbound

This paper cites Winoground: Probing vision and language models for visio- linguistic compositionality.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Winoground: Probing vision and language models for visio- linguistic compositionality

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.768198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:20:51.192566Z digest=sha256:2d0915ca666be9c4acc7c87ae4af6c8dc1349a91642414ca207cbe9622efce23

Observation df1e8bc0-cccc-4034-90a9-ac8b7fbfef8c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation LLaMA: Open and Efficient Foundation Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.196243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.196243Z digest=sha256:ec51b94bd6c55ec3ab1eb793c375436594ef170525b879b4fa7af4a25279b08e

Observation c1827cbc-35e5-458e-96d8-66c9a8fe0679 · outbound

This paper cites Diffusers: State-of-the-art diffusion models.GitHub repository, 2022.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Diffusers: State-of-the-art diffusion models.GitHub repository, 2022

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.756013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:20:51.199902Z digest=sha256:8ee759038c150b2042b320d1be4436c5ed51e8b464f86520fec2a60b745cbd64

Observation 5727adb9-dd26-4d48-b90c-70e3e7357347 · outbound

This paper cites Text Embeddings by Weakly-Supervised Contrastive Pre-training.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Text Embeddings by Weakly-Supervised Contrastive Pre-training

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.203423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.203423Z digest=sha256:b441a32a7073a6988fc8f774c04d97dbdc80797f4339044eb054df9c74135b98

Observation 3189fe56-8ea4-442d-b5f1-2e8540898742 · outbound

This paper cites Improving text embeddings with large language models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Improving text embeddings with large language models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.744839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:20:51.206990Z digest=sha256:cc23ffd2b8a831c172f3386f36f35ae023761036f14fff6deeb6691eae393369

Observation f715e3cf-606d-4af0-979d-491ffc7800f5 · outbound

This paper cites Revisiting Text-to-Image Evaluation with Gecko: On Metrics, Prompts, and Human Ratings.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Revisiting Text-to-Image Evaluation with Gecko: On Metrics, Prompts, and Human Ratings

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.210596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.210596Z digest=sha256:aafb27379a05568eb758adc6d0d1b02744e2af6175b46a44b31ec1e44b3c3f7d

Observation fd00a9b6-0f84-43ba-831f-ebc5df508ea0 · outbound

This paper cites SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.214409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.214409Z digest=sha256:85dfaca393cc5964ecb05b1f7cfb08e97e270f4debeb03bc334d2c857358e854

Observation eca2e838-1853-4e5e-aa53-4efa5424ae8d · outbound

This paper cites ByT5: Towards a token-free future with pre-trained byte- to-byte models.TACL, 2022.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation ByT5: Towards a token-free future with pre-trained byte- to-byte models.TACL, 2022

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.733137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:20:51.218461Z digest=sha256:c6ea9093054c2a480471fefaf05187f173b985a8967100e74be06311f8900f8a

Observation 1f715680-9d55-44af-af67-261e85e2dc28 · outbound

This paper cites Qwen2 Technical Report.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Qwen2 Technical Report

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.222983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.222983Z digest=sha256:0ee23eaa96ec79458cf131829c515f432ba79a025776ff49d3fb4ca6d5d23843

Observation 8ca66ea2-ab10-4a4d-8f5a-92033c24c6b7 · outbound

This paper cites What you see is what you read? improving text- image alignment evaluation.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation What you see is what you read? improving text- image alignment evaluation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.721893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:20:51.226960Z digest=sha256:9b86c96ec05d8d3ab76f021dcf60675d51343daa49b261e1b39be9801f28829c

Observation c5623f59-da6b-48f4-81dd-6f7bd2954a2d · outbound

This paper cites Investigating Layer Importance in Large Language Models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Investigating Layer Importance in Large Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.230669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.230669Z digest=sha256:0f003b0363dc723abe691beb6cdcb6eeffd0f2e5097dd5ec60e251c3a6a86ae7

Observation d7afd82b-89f1-4642-9c27-61a1bdf32373 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.235247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.235247Z digest=sha256:bfa366f5e891a8bf58ba2d37abd08fcd3e87d4829444db2f7d6411136407589c

Observation ae66d45d-f3f5-45a0-91f6-994ca9ad6a26 · outbound

This paper cites Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.240632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.240632Z digest=sha256:14c5d6281ccb3d194910e130a805c346a92faa98a8c4b518e1f8fbb32390674f

Observation 275b4bab-9df2-45e8-a30a-4291d5d7f937 · outbound

This paper cites a beautiful morning in the woods with the sun peaking through the trees.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation a beautiful morning in the woods with the sun peaking through the trees

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.711029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:20:51.244779Z digest=sha256:5aee4a54a8887a12765481aed8c0ac1c37c25472988df32030605def0a5072e7

Pith citing papers

No inbound Pith citation observations are available.