Pith. sign in

Paper Citation Record · LEDGER

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

As of 11 August 2026, this Paper Citation Record lists 99 of 99 outbound references and 100 inbound Pith citation observations for arXiv:2206.10789.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2206.10789 v1

Coverage vector

measured 99 of 99 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T04:49:30.873360Z

measured 199 of 199 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 100 of 171 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:39:36.759706Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

99 of 99 outbound references displayed

  • verified exact22
  • verified fuzzy69
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch7

External citation measurements

340
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation bc14b55e-31fb-492b-b253-d76c946e8db3 · outbound

This paper cites Introducing pathways: A next-generation ai architecture.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Introducing pathways: A next-generation ai architecture

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.644567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:d574b1018729ab83e510e1b590e8e21e2ccd17beff9e2a14c941d0852c839532

Observation c98725d8-a553-4cc4-85d5-465ec6b806ba · outbound

This paper cites Zero-shot text-to-image generation.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Zero-shot text-to-image generation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.587287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:910539db41b7e621b4661d508364f55c516653a8a95336a5b18163c810076162

Observation 3c2b7fd0-e299-4b4c-b971-f21e8358ef59 · outbound

This paper cites Cogview: Mastering text-to-image generation via transformers.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Cogview: Mastering text-to-image generation via transformers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.591715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:ab9417453c045faeeedd8d7bb7cb44de4207c4742f95f16095a1af6d39b9880c

Observation 3bcc3d32-7601-4dec-9514-e5552e66bfc6 · outbound

This paper cites Attention is all you need.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Attention is all you need

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.596999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:520bfc359ea3bc40ae96cbeff0ff03826085a7b4d85b2215eb145853bc39937f

Observation 936e10fd-d240-4a20-b291-8099e71d41df · outbound

This paper cites Discrete Variational Autoencoders.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Discrete Variational Autoencoders

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:49:31.033281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:6c365998a5784576910e916b958e7eb8a875566aa8cb44cf85146dcb4c9235bb

Observation 1149fe5e-a39d-400b-86e7-44e55f73db09 · outbound

This paper cites Neural discrete representation learning.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Neural discrete representation learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.609894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:d20e0f1524ec62853b71136f1992af2e7a2f2d438ecf421492da52a8eacf3e77

Observation 828bd678-9ec1-45f1-b097-7da471cd18b5 · outbound

This paper cites Improving language understanding by generative pre-training.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Improving language understanding by generative pre-training

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.618358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:9406d29d9046b21730cb1363e187b9a1e1eb06599f2adcc8ae98f3ba71f335a9

Observation 4d4d2ce9-baf4-4c3a-9bd3-b4dbe0aabd2e · outbound

This paper cites Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.622593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:b2ac69c06f78d42bae0232d3c1d284a562fd2cc548b5bd0a6554f2d7efd97e6e

Observation 1f394c72-f280-492d-813a-35c6bf0e1ae0 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Scaling up visual and vision-language representation learning with noisy text supervision

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.629644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:bdfbddc3091f26dcc8234f90d33f6367dad4f10747c2e103970cbd19d57c7182

Observation 69afec39-e1a1-4aad-8f38-061e7d1ca95f · outbound

This paper cites Make-A-Scene: Scene-Based Text-to-Image Generation with Human Priors.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Make-A-Scene: Scene-Based Text-to-Image Generation with Human Priors

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T04:49:31.051437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:21d39fe9d9da8e206ca802c349f6780426b0dbf12a68aa7ff1d1b38be0d349d2

Observation af9d287a-d262-45be-812f-29ab0cc11cc6 · outbound

This paper cites GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-12T04:49:31.059778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:085f0a44aec3a43acab36f4412fc7b07190e1767d9f2b978047e452070c10ece

Observation 45786d5f-b4f9-4f79-b8fa-2cef30ee3539 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-12T04:49:31.067347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:7c7f4f041dfbfc49566fa4a43d4bd2ce3d2292ba1cfb2e85271feff0f1fe1a8c

Observation 7e9dd7e4-914e-435a-95a3-1c5ab55d28f1 · outbound

This paper cites Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:38:54.138183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:33ade15b2a3fe390973cd06d2064412474742651f01cb97cc3b934482c9ad0af

Observation 3bd4ea1e-c251-4b61-9239-40ac06581b8c · outbound

This paper cites Denoising diffusion probabilistic models.Advances in Neural Information Processing Systems, 33:6840–6851.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Denoising diffusion probabilistic models.Advances in Neural Information Processing Systems, 33:6840–6851

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.655922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:2914e3c784c298b3e4b036efa98ee934f551f324e527ed652d02265165784b41

Observation 498d97f3-828e-43d4-b139-db3d53c746a5 · outbound

This paper cites Diffusion models beat gans on image synthesis.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Diffusion models beat gans on image synthesis

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.663780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:e80c5dcd359d31c6f9f223e8e2e7c27a7a8bcd6bbce8d49da69d55e769a4143e

Observation f80132c1-4209-4e97-b47c-e624b046e93c · outbound

This paper cites Microsoft coco: Common objects in context.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Microsoft coco: Common objects in context

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.669823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:5465974cc5e1785a1eab5f0bde2c03e1e7cd709bbc67caf90a0a7350d183dcc1

Observation cd5cbc39-ffc3-4855-82bd-4e725b357203 · outbound

This paper cites Language models are few-shot learners.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Language models are few-shot learners

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.673381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:ad9290f744c9450293ad5c783d35328e5f5c86f202d5004049ccdfb01d1a4932

Observation 1ff6f802-2b3d-47f3-b04c-411acc00d009 · outbound

This paper cites LaMDA: Language Models for Dialog Applications.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation LaMDA: Language Models for Dialog Applications

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-12T04:49:31.087810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:6cda46dbce3d08546cff2dd3e2ad74086f9f0aa13523a95152871a6eec119e5b

Observation 545b5eed-f442-4098-ae5f-386a3b60a53f · outbound

This paper cites GLaM: Efficient Scaling of Language Models with Mixture-of-Experts.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T04:49:31.100632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:844b106d4d4239e82c0e8e5bda194e3b773bcd9e52443a6ab49f1c4cfa3d6ecd

Observation 50324ecb-9037-446b-ab62-313e7939e255 · outbound

This paper cites Vector-quantized Image Modeling with Improved VQGAN.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Vector-quantized Image Modeling with Improved VQGAN

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-16T18:40:37.571629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:b6e859fe21130977ece33f00f761d0966bf72f377f647e419d78bfb4dbb47c04

Observation 05a7cb37-7bfd-4930-9c01-21472b2178b6 · outbound

This paper cites Transformer transducer: A streamable speech recognition model with transformer encoders and rnn-t loss.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Transformer transducer: A streamable speech recognition model with transformer encoders and rnn-t loss

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.694164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:3d3fa1ca41198c107c01a5c30f61ff3f9acc8902999293489fb2940e5c471bb4

Observation 04016ebf-6b02-4846-90c0-323f59c02367 · outbound

This paper cites Conformer: Convolution-augmented Transformer for Speech Recognition.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T04:49:31.123367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T15:38:39.405362+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:f1fef7e71b229f187cdbff95006922a0158039a4dda829267f8371f8b335f2ac

Observation bd23425f-c046-4e9b-84e8-d2ff0a2b1a2f · outbound

This paper cites Towards a Human-like Open-Domain Chatbot.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Towards a Human-like Open-Domain Chatbot

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T04:49:31.138132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:5c0c22af7f3c20f20921e9a7c5efee858f6c42a34aaf500da6a9eec1734b2014

Observation 8131ddf4-a86e-403d-bc15-313b37c409a2 · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:53:08.606863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:82968f7729694b2a208a17844ea8bc84f8330d4b98190349a9dfb120f28d5bfa

Observation 143e13db-4f46-4d62-b1ab-0f08c017fa60 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Taming transformers for high-resolution image synthesis

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.721279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:563d5b88225f53c0c075a1ce9a0634e41fd39b35529aa3b716ceedda93ee1f23

Observation fa79bcbc-fab3-422d-932a-f0ae2ed8c317 · outbound

This paper cites Megatron-lm: Training multi-billion parameter language models using model parallelism.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Megatron-lm: Training multi-billion parameter language models using model parallelism

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.724620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:6edf7cc7953a226506cd15f9b4a6ee486dc847602f13dd803c1729377c922249

Observation 7e74ebce-cb32-4e25-bce1-eac787778fe4 · outbound

This paper cites GSPMD: General and Scalable Parallelization for ML Computation Graphs.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation GSPMD: General and Scalable Parallelization for ML Computation Graphs

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:36:36.562059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:e9d6754d9fbe90d768a567ab8d5520a4c9850a8465c72cad02c47870f7a60bbe

Observation 00235bea-dff7-4c0f-946a-797dc8a4f904 · outbound

This paper cites Connecting vision and language with localized narratives.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Connecting vision and language with localized narratives

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.732805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:4d3fd99de6e65ae0368262e33c6505969c34d6d71625ac0518f42098c1071d81

Observation 3dbc4469-1055-4b02-962b-f97599fe5eea · outbound

This paper cites Generative pretraining from pixels.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Generative pretraining from pixels

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.741041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:2686f63092c13ba971fccfe3afeb1b02ac74d26a84fe682c2b5e3ef85c3cc2b5

Observation 137f2d6e-6044-4aec-9c16-0989e9b915ad · outbound

This paper cites Wide Activation for Efficient and Accurate Image Super-Resolution.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Wide Activation for Efficient and Accurate Image Super-Resolution

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:49:31.176535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:c30005bba74c93a32d9b20f66efa396c9438ed4d3493adac8505955629928e1a

Observation 1bb5d6e5-2884-46d7-b5a7-f2085104e115 · outbound

This paper cites Neural Machine Translation of Rare Words with Subword Units.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Neural Machine Translation of Rare Words with Subword Units

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:15:20.428834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:ec1a96a1c729ac2535bd3186b5471cf4f6cb673379f24a810cfec8116162b90d

Observation 877876c2-074e-40cf-8d92-1a77a40abf1a · outbound

This paper cites SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:09:07.599904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:181e4aa92f25c9748ca3254040dbafe0107e887de3c608d18116526e02a9831e

Observation 7eca41e1-ceb2-4e39-aacf-7fd5677feeee · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Generating Long Sequences with Sparse Transformers

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-12T04:49:31.204641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:38a1ad1142a2149e375564dee51c57930969c71722652b3984947f66c2b29e4f

Observation c6204cef-b5c7-47c8-8159-84e249a78260 · outbound

This paper cites Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:37:56.023799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:9f2a73a052581bad29137d9c745b8645ffaf6977b9641b028cb362f7c587c44a

Observation c703b258-e8b4-4ce4-9322-f091c0240866 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-12T04:49:31.222092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:d055b91e258272cf3aecd8b6f74df620a883f7a53539990e580619f69c65f877

Observation 4a03492d-8135-4a75-97d4-7b2f4a20d2f3 · outbound

This paper cites Classifier-free diffusion guidance.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Classifier-free diffusion guidance

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.373070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:88c7ed83e9741a16f89f65b96e010078d98cf0ebd4703f3fb6a574cc549936ce

Observation ef326a36-0001-43ef-ba0a-901466d2783a · outbound

This paper cites Classifier free guidance for autoregressive transformers.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Classifier free guidance for autoregressive transformers

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.377792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:8c632286f01273a3b201f3d631a20b633976578011c12c454a1da08f8f0ee103

Observation c569a808-39c0-4161-b83a-8a148736d3bc · outbound

This paper cites Lingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Lingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-04T23:25:00.242485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:a66d013aa2cd89f5370ecedad23d7c7910b9e4993625a42857e512a92a144f9d

Observation 06281e1a-070e-4deb-9e6a-c4ff25376b5f · outbound

This paper cites Le, Yonghui Wu, and Zhifeng Chen.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Le, Yonghui Wu, and Zhifeng Chen

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.389060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:ea4d54e65957cf6c1f179823824f63fee40e2dea19ba7665158ca06904f24780

Observation 7491df29-2d76-43ea-adc5-8af51716372e · outbound

This paper cites Efficient large-scale language model training on gpu clusters using megatron-lm.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Efficient large-scale language model training on gpu clusters using megatron-lm

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.399425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:6b21cc636063902606a2ee014c8455b8ed241bfca865e16c0c514386235a57f8

Observation 378d248c-f965-4e02-9190-6f8e9a51d07c · outbound

This paper cites Adafactor: Adaptive learning rates with sublinear memory cost.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Adafactor: Adaptive learning rates with sublinear memory cost

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.410042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:8b903742a16550de97eb30c5578445dc95748a012d73b69c11f53811939ec539

Observation cfa0a83a-5a53-407f-8f97-9ff47245f3a5 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:21:01.130308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:2c49f5e0795c9e568653720c2bc143c697e504b9964e019bc711c0e1c31ceaa1

Observation ce2162a6-d9f8-435a-a5a5-e14e7aae0bdc · outbound

This paper cites Scaling vision trans- formers.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Scaling vision trans- formers

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.420494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:af4922dacf6ab37b19ee5881806415971a84a4f4f3d9bd9539a7834be7f6ff4e

Observation eed7dac8-d443-494c-8abe-df58544933c5 · outbound

This paper cites SimVLM: Simple Visual Language Model Pretraining with Weak Supervision.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T04:49:31.252855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:462d7942a8c05662998109aa5e58991151dad2ac9e2f6d55e12aa50450d145b9

Observation 98aa7198-2109-4a10-a9ee-40b344173f80 · outbound

This paper cites Text-to-image generation grounded by fine-grained user attention.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Text-to-image generation grounded by fine-grained user attention

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.432097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:d0b33d687672bf90ae82aec8ce8faee4f43f1f99935c3ba465e4290df9a4436a

Observation 04ec64ab-5066-4e41-948b-ff1fbc4be3f9 · outbound

This paper cites Cross-modal contrastive learning for text-to-image generation.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Cross-modal contrastive learning for text-to-image generation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.436158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:9a8245bd0503a50673fe912b56d288eb9a34cdc5fc84873afab352c2f8bf535a

Observation fc56d227-404f-4e5a-8a2d-519b26719380 · outbound

This paper cites Benchmark for compositional text-to-image synthesis.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Benchmark for compositional text-to-image synthesis

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.443984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:d6eb50fde75a6de1b61624d8b742139d9def2433a1e4fdcae80b888a6b902788

Observation 3ee3cc0d-44c2-4742-a9a1-e5ad6544a394 · outbound

This paper cites Vector Quantized Diffusion Model for Text-to-Image Synthesis.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Vector Quantized Diffusion Model for Text-to-Image Synthesis

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:49:31.260695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:25793d5879e0291e1de44d51da600e7214824bda14d53160ea76fb2e899a540a

Observation fbca2ccf-b9b8-4b30-ad1a-daa809c88995 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Scaling up visual and vision-language representation learning with noisy text supervision

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.455566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:4ca58bc173f5e34a6d93b9ddfdb81bc8361345623394ebb04773d803f7ec39f9

Observation 60cbf6a9-7dbf-48fc-b45c-5f356ef8902a · outbound

This paper cites Accelerating large-scale inference with anisotropic vector quantization.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Accelerating large-scale inference with anisotropic vector quantization

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.464476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:8bc0f35be113287ffeb0fc259a6259ee7ea0af3079e07b0ea7a9f457de865f63

Observation 26047782-7f48-444a-91d8-9ae4f4f01ab2 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Gans trained by a two time-scale update rule converge to a local nash equilibrium

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.468484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:08aeea1c9d0e6bfde304ca5a59e5133964310449f1c4c0aad896b37f170934c5

Observation 0b012cc9-0104-45f7-8d46-91263668a923 · outbound

This paper cites Re- thinking the inception architecture for computer vision.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Re- thinking the inception architecture for computer vision

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.472063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:f88f3445b355ee104ef046798134cdc823588fa30116ee296c4cf52d27ea8c49

Observation f2a79930-87ac-4367-809a-7b1bd23eeae0 · outbound

This paper cites AttnGAN: Fine-grained text to image generation with attentional generative adversarial networks.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation AttnGAN: Fine-grained text to image generation with attentional generative adversarial networks

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.475897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:7fb3ea6c963c482b90f0fbd22f1438dedfe410aced69886ad3145af89c2dc34d

Observation 078c06e0-abb8-4316-ba83-53463e646c7c · outbound

This paper cites Dall-eval: Probing the reasoning skills and social biases of text-to-image generative transformers.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Dall-eval: Probing the reasoning skills and social biases of text-to-image generative transformers

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.486349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:d6ecca1c16adf55beda1eaa1be20fa8fe54b3ed2bf9ce556036dcb320c050e9e

Observation e5cbe754-66c4-4f35-89f7-a4540b90dd5a · outbound

This paper cites Unifying vision-and-language tasks via text generation.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Unifying vision-and-language tasks via text generation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.493587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:aebdc7824beadb12bd529b6dd3876c9e8ffea32f7d528bc4c25fa30a9275cf17

Observation 5e496484-b211-485f-95ad-f4e8ddbabeaf · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Bleu: a method for automatic evaluation of machine translation

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.500028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:f41b77559970b1f403c610a85ce87dde9f7919d104f73aa356a9f9ab840025cd

Observation 79ed99c4-235c-474b-a436-0873eea9d43f · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Lawrence Zitnick, and Devi Parikh

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.506731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:c3790a94ca8ccbd54504a504530d56075905f60025a8e6e40c9c1a785d85a46a

Observation 07f1d6e4-64ba-43b3-a97b-de8d625015d8 · outbound

This paper cites Meteor universal: Language specific translation evalua- tion for any target language.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Meteor universal: Language specific translation evalua- tion for any target language

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.511991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:453afb1cb7343eb69c56c2aa649dc7f3062fad704de6f2a1c064ad4b169e6b02

Observation c0d00522-1d6c-43b4-94f4-317f116f5fc4 · outbound

This paper cites SPICE: Semantic Propositional Image Caption Evaluation.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation SPICE: Semantic Propositional Image Caption Evaluation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.517421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:25c18bec82c668f8f4348cc756062142c6f5ae9a993a39e57f139e6e565cee7d

Observation 8c4d1e65-12d7-4008-90f3-000fd769abec · outbound

This paper cites Cogview2: Faster and better text-to- image generation via hierarchical transformers.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Cogview2: Faster and better text-to- image generation via hierarchical transformers

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.522321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:c0b65bf8459af93d18c449126f078253cca343082859c215a0a79a615a0dab3f

Observation 1f94c185-79d8-4e82-89a2-f87dce69d1de · outbound

This paper cites mindall-e on conceptual captions.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation mindall-e on conceptual captions

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.528124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:b834148dc7a47d025b744c48ef616751130a44e9a88dc6be2db3aaa1b94b9dc1

Observation 351e33f4-d5a6-4193-abe4-e73fce1c3d8c · outbound

This paper cites X-LXMERT: Paint, Caption and Answer Questions with Multi-Modal Transformers.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation X-LXMERT: Paint, Caption and Answer Questions with Multi-Modal Transformers

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:49:31.268347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:10c6b04180ffb74418894432b009536bd8f390f0c3100d26f66a016fe2a46221

Observation 7ed17d4d-3339-43f6-ae5d-3d57d8f1d3c1 · outbound

This paper cites Deep visual-semantic alignments for generating image descriptions.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Deep visual-semantic alignments for generating image descriptions

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.538923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:2aaade648204164250f1b9ec875135ae33bda21f78a83ee724b99cbac12d964e

Observation d3e075f2-626c-4030-9056-4b39d39ed05c · outbound

This paper cites How to marry a star: Probabilistic constraints for meaning in context.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation How to marry a star: Probabilistic constraints for meaning in context

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.543969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:dcabe22cd2b5610c84287480b3d5ba061e9953c766a089c4030edef040a42311

Observation a6ba989b-45a6-4104-9702-77c2e931570d · outbound

This paper cites Wordseye: an automatic text-to-scene conversion system.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Wordseye: an automatic text-to-scene conversion system

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.550200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:7f2948d06ce254aae1d1c8f3fed5eae1ff93c96f77129189f055c08df5b77e37

Observation c9072070-eca4-4f6f-aa20-30a2a823162a · outbound

This paper cites Generative adversarial text to image synthesis.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Generative adversarial text to image synthesis

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.555183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:2c30a6171f4eee098cc5afb4e40a8f543daaa829aa194d687fb787193d4fc0d2

Observation ae0499a9-51b6-4d2b-9bd9-348702e53c03 · outbound

This paper cites StackGAN: Text to photo-realistic image synthesis with stacked generative adversarial networks.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation StackGAN: Text to photo-realistic image synthesis with stacked generative adversarial networks

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.558785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:415ba680f4603c0bf9ef9caec8a2d9b885af3ee78860a4c82363e2d6df997d44

Observation 9ae4e39e-54bb-4ad9-b3c4-0a53f401a474 · outbound

This paper cites an unresolved cited work.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-05-12T04:49:31.562208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:1f836efaaf2ac8766539f8d1df916722eb284546067cdd519fb38b0faec30d92

Observation 39b3f9d0-f299-4dab-9248-614c265f483d · outbound

This paper cites Learning what and where to draw.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Learning what and where to draw

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.566854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:9c4ddb39a06adeaa9b89ef2efe8aee31e1d7fcfc314f4b31649ab5a8e97107fa

Observation b9b2a0fe-b2b5-45ae-ae9e-9c6aa5c53710 · outbound

This paper cites Inferring semantic layout for hierarchical text-to-image synthesis.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Inferring semantic layout for hierarchical text-to-image synthesis

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.573248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:fc5c4458fb572ef66478e01f1c7d6477d8f43da6a2984150f8903018fa85998e

Observation a4d6adf4-3f47-42ce-9abd-f740df48e97a · outbound

This paper cites Generating multiple objects at spatially distinct locations.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Generating multiple objects at spatially distinct locations

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.578256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:8daea9e1659556b68e7242183aa9ffb523dbf5fa9551c7b034adbbfd3ba2d508

Observation 9d15a21e-3156-468b-9f7a-5e0bf16ca814 · outbound

This paper cites Dall·e mini, 7 2021.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Dall·e mini, 7 2021

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.605957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:d553bb81aee7f4fb77af2660955eeadec9f360cf766b88181d6f40c4fc8e9faa

Observation 6e2c1129-6f29-4c6b-a1fd-d6d781f258ce · outbound

This paper cites MaskGIT: Masked Generative Image Transformer.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation MaskGIT: Masked Generative Image Transformer

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T04:49:31.287350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:3c33926df8e3ea591fd513a79cd50438682ea96cb390279221fb4d86bc66d28e

Observation 16e94e7e-e41e-4bb5-b982-80daf11a579a · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation PaLM: Scaling Language Modeling with Pathways

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-05-12T04:49:31.295609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:eb72993bc509dfcfd7e22dc3613a86a6dadc8cf0489a5de7aee5bb51f9717b9f

Observation 7de39fc6-bd99-4eab-bd79-7b2d76bb6058 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Training Compute-Optimal Large Language Models

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-05-12T04:49:31.020705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:90700b9e2aaec8ccada6e0d52c4273a84980cd15aecd115ba7d6d458053157da

Observation 9ecbe2e2-4e51-4c53-97e9-f8eb645105a6 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Learning transferable visual models from natural language supervision

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.650795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:6149e69933f847b9e2b73ea07588003a526c5665aa52cc45c36464c95d061808

Observation 7091d1df-f145-487e-8621-5e347984e964 · outbound

This paper cites Cascaded diffusion models for high fidelity image generation.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Cascaded diffusion models for high fidelity image generation

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.680855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:e6c419407077398ba6b20493e50640697665c0c5342ba43e204c2a1910745c96

Observation b40d0471-b9e2-4fc3-b81c-212509fccde7 · outbound

This paper cites Analyzing and improving the image quality of stylegan.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Analyzing and improving the image quality of stylegan

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.687216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:c502d1073ae15d83132832a3b7f46f06e21eb278b6cf7872aca6bc60d9efe0e4

Observation 216471d8-8cf4-4447-b34d-025fec080f40 · outbound

This paper cites Perceptual losses for real-time style transfer and super-resolution.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Perceptual losses for real-time style transfer and super-resolution

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.690922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:918a9200544ceeb68bfdcd3d38a86f2b49037f56a2c24460328d963b93c3cded

Observation d9e7c3d2-254f-405d-a860-71a6adc8f0cb · outbound

This paper cites The unrea- sonable effectiveness of deep features as a perceptual metric.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation The unrea- sonable effectiveness of deep features as a perceptual metric

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.699162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:08c023a34bf4254b88e3569aa444b1bdf072966d121e4a6417c348d4e2ef5b67

Observation b4c8d938-ab4b-495e-9712-209263673dce · outbound

This paper cites Anonymous paper under review.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Anonymous paper under review

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.710989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:204e7664c8c4a24839f0c2aadbe2f485b3a573ea2c30c30a6e0ce45cd6e92a98

Observation 2887472c-5a7a-4148-bf10-86834a88fc41 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation On the Opportunities and Risks of Foundation Models

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-05-12T04:49:31.301985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:115d59bfd5284b37a4dbd637f98bea6e219950d70abde179e9d1d981de1bc463

Observation 35898be5-69dd-47eb-a369-e889fd4a237d · outbound

This paper cites Towards accountability for machine learning datasets: Practices from software engineering and infrastructure.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Towards accountability for machine learning datasets: Practices from software engineering and infrastructure

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.727890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:7afdee8c71f99f8d1c48bc4b0d232b697f9db3deb1f4f214e0c1e322377a4279

Observation 64c75d76-5a9f-4bc0-9221-e5e6c40dc688 · outbound

This paper cites On the genealogy of machine learning datasets: A critical history of ImageNet.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation On the genealogy of machine learning datasets: A critical history of ImageNet

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.331152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:8701dc7f29729135eb7013eeeeddae514d96ada1b7aa810b2e2f232862372ebc

Observation ad9abd1a-ccbb-45ef-a8dd-0c96a1bec59d · outbound

This paper cites Model cards for model reporting.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Model cards for model reporting

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.338893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:1ba62d81099f2f08f1aff2fee7796eea8cb6abe13b2e2e6651fa393b252b8cd3

Observation fec16499-e3c3-408d-bdf4-0b2f3611aa81 · outbound

This paper cites Datasheets for datasets.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Datasheets for datasets

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.342826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:b36c50bafceccb1f1b742255892a5b47262a0aa30b0cdc80ef77f26acaac9fc0

Observation 7f7eef7f-8a5b-44c4-83d9-854365b75b25 · outbound

This paper cites Data Cards: Purposeful and transparent dataset documentation for responsible AI.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Data Cards: Purposeful and transparent dataset documentation for responsible AI

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.348931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:5307e58150fbc1925e4207d77328111e213c780f1e527ea7d05cfdf878b008f8

Observation 91f1b8c1-bc63-41a7-bbfd-06145b98c72f · outbound

This paper cites Tell, draw, and repeat: Generating and modifying images based on continual linguistic instruction.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Tell, draw, and repeat: Generating and modifying images based on continual linguistic instruction

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.353081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:e86e06ce44f6e18a93c5afe633226b0cc0350d65580ef9d38bf8f897dbeb7601

Observation f8b78fcd-5d49-4a78-a0ea-2406892275e1 · outbound

This paper cites Chatpainter: Improving text to image generation using dialogue.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Chatpainter: Improving text to image generation using dialogue

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.360345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:3ad300bcad88ff02375410a0828a368c3012049c22f9bfd8ea266d1231a9ce39

Observation 529981f9-fae5-4074-a160-7e05d9dcc656 · outbound

This paper cites Disability studies as a source of critical inquiry for the field of assistive technology.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Disability studies as a source of critical inquiry for the field of assistive technology

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.383296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:06aca83252bea253cc6c2d032eba8d9b6e0b5a1e9718d59964cfdb5fe744be8f

Observation 193ebf13-2669-4203-aecc-cb975069b782 · outbound

This paper cites Who (or What) Is an AI Artist? Leonardo, 55(2):130–134, 04 2022.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Who (or What) Is an AI Artist? Leonardo, 55(2):130–134, 04 2022

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.416655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:b8ca7873c8edb4b6a790ffb3e3d0f5d6d5acb75d619bea1c3405ea51262841be

Observation 15cc6739-3416-4b8a-bec5-ea085005e266 · outbound

This paper cites Biases in generative art: A causal look from the lens of art history.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Biases in generative art: A causal look from the lens of art history

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.427582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:503aed9731ab06475c8d16d2cf06088d0d84cd029891f23ce22975ab6d642b18

Observation 602d13d7-e314-413f-a54f-159907a83a75 · outbound

This paper cites Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:21:29.028955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:7c07b6929b4257d225b656f0d7535168f575d65732ad5fb9dbae7e0e5f071020

Observation 6eb280d6-dc86-4fa3-a8ba-65c2ed3f8b33 · outbound

This paper cites Data-driven sentence simplifi- cation: Survey and benchmark.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Data-driven sentence simplifi- cation: Survey and benchmark

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.533990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:f3605312ab2ca60bc21ccbc92c8ffe460d313d44fcb61336c337990acf47b539

Observation 35589fa7-491a-40fc-9ccb-971b4dfcbdf2 · outbound

This paper cites Paraphrase generation: A survey of the state of the art.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Paraphrase generation: A survey of the state of the art

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.633229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:eff421b5a7b8d76dd589daad291467ded3b14127373a744fe08d8b21e83e8e16

Observation d25f3bfd-87b3-4e96-9db8-6fb1efbef896 · outbound

This paper cites No classification without representation: Assessing geodiversity issues in open data sets for the developing world.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation No classification without representation: Assessing geodiversity issues in open data sets for the developing world

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.639303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:fa8c7745483dfc2466f817db723af1d9d6ed75af6783c48baed5d92219599e99

Observation 2cdab732-07e4-4742-adcd-fc368d58528a · outbound

This paper cites Does object recognition work for everyone? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pages 52–59.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Does object recognition work for everyone? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pages 52–59

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.717350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:059c0f60dd7f10a96948542f3beb32b2068c900c022113ef434b3fe3634cf6e3

Observation ba0c2444-b47e-45f3-8891-1a7d1449bd61 · outbound

This paper cites Distortion agnostic deep watermarking.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Distortion agnostic deep watermarking

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T04:49:31.447494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:eaf5952bb19de2010f38773159b36a31142d35618b71f9af22401d5f9350aa65

Observation 204ebc01-8786-4186-960d-4bb4f74cd790 · outbound

This paper cites Multimodal datasets: misogyny, pornography, and malignant stereotypes.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Multimodal datasets: misogyny, pornography, and malignant stereotypes

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:49:31.324161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:df8d3f344a43ccb41903335727a0c4057c7f2f524b012bd418f75536d5d802ed

Pith citing papers

Observation ed759494-4e47-4531-aed0-2fdbfb2b4bde · inbound

An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion cites this paper.

An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:49:31.747550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T18:08:55.311069Z digest=sha256:5785ebd55df21d1dd090657a523141e739b08c729825c7d6748500cae54f4889

Observation 73f522ce-279b-4823-ac7c-cb65e3fada02 · inbound

Prompt-to-Prompt Image Editing with Cross Attention Control cites this paper.

Prompt-to-Prompt Image Editing with Cross Attention Control Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:49:31.747550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T07:00:01.154743Z digest=sha256:35db1332a56a51867c0405b0a3bb5591ca08dbc3d28664b6f3c3c111bed42d93

Observation 2c9909b1-76e8-4c41-809e-253f5f0f18e2 · inbound

Make-A-Video: Text-to-Video Generation without Text-Video Data cites this paper.

Make-A-Video: Text-to-Video Generation without Text-Video Data Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:49:31.747550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T01:13:03.213358Z digest=sha256:7ad84509df4570f256406e31619cb9e6e79e81aa3faa1a4f2a3f2d11c26a7648

Observation bb7fbf59-e28f-48f8-a0b2-545cdeb66394 · inbound

DreamFusion: Text-to-3D using 2D Diffusion cites this paper.

DreamFusion: Text-to-3D using 2D Diffusion Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 160

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:49:31.747550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-11T11:27:10.287358Z digest=sha256:1f78593c35fd4729468c847a1bf5639cfbbe3b5e5d6dcca55b1e0447e4fea026

Observation f9f0375e-9642-413a-bf1d-7ae6887d7bac · inbound

Imagen Video: High Definition Video Generation with Diffusion Models cites this paper.

Imagen Video: High Definition Video Generation with Diffusion Models Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:49:31.747550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T03:31:08.155347Z digest=sha256:2719855210dee9472ecd6dbed8c6da4d277f1ed94da3d38d7a32e418505986de

Observation a6cd0a22-c329-476d-b79c-1222b7d9ed4e · inbound

Phenaki: Variable Length Video Generation From Open Domain Textual Description cites this paper.

Phenaki: Variable Length Video Generation From Open Domain Textual Description Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-17T01:43:34.177769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T01:43:34.024375Z digest=sha256:1b5baaa97e3ad2736d51e22a9dd2b9e136a7e72d857fefcce9fb17fb427ed804

Observation 55fb41e9-13c6-47fe-ab76-b0d73638b582 · inbound

LAION-5B: An open large-scale dataset for training next generation image-text models cites this paper.

LAION-5B: An open large-scale dataset for training next generation image-text models Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 93

Resolution
verified exact
local_arxiv, observed 2026-05-13T14:22:17.446645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T14:22:16.968028Z digest=sha256:60bbc82a0239c11045fde5bf00fba0adbe6d1abcfa07bad9ba915fb527b4be27

Observation 6dd858a3-a606-40fd-bb39-a89b27ce06c4 · inbound

eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers cites this paper.

eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 87

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:44:22.842165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T01:44:22.710205Z digest=sha256:510e90c4a6a202be7eb32baf73ce447820dceebcfc345ced25e21e893e557a87

Observation de0d3ea7-bcce-4e3f-84fa-e765473df9c3 · inbound

Galactica: A Large Language Model for Science cites this paper.

Galactica: A Large Language Model for Science Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:53:21.931338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T05:53:21.810346Z digest=sha256:cda4195d6d7156633a5c232bff433f0c8f653c05c4991e85369fdd8158cdf83f

Observation cc1dfb85-3770-42c6-84d3-7b0887b400a7 · inbound

Continuous diffusion for categorical data cites this paper.

Continuous diffusion for categorical data Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 98

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:30:22.339640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-18T03:30:22.025578Z digest=sha256:f5efde1894baa0e568dc94555dd56edf8380995b8e844ef67038a56e969525c1

Observation 96609aa2-9528-42b2-8e77-8c3b6ba190bd · inbound

Fast Inference from Transformers via Speculative Decoding cites this paper.

Fast Inference from Transformers via Speculative Decoding Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 67

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T22:52:00.223825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-17T22:52:00.101612Z digest=sha256:c4fb7f1c5e9250627fb972a74975e3b774f7fba9201a82e8eb165561597a5e5d

Observation 4a244c0c-9f35-4558-992c-af6ed4c78b8b · inbound

Scalable Diffusion Models with Transformers cites this paper.

Scalable Diffusion Models with Transformers Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:04:05.810349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T06:04:05.434354Z digest=sha256:c2322c60ff48a250580af81b3322bda412d7eefa5b7d52049e3215ec7f4f1916

Observation f0c655f1-1a4a-41ea-be5f-fe1eb4976b81 · inbound

Visual Instruction Tuning cites this paper.

Visual Instruction Tuning Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T04:49:31.747550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T08:22:03.403362Z digest=sha256:027a37a1a849dbb596922dda6d2924576accfc35137cf81d8fb75a107778f182

Observation c437b3ac-3c66-48fc-ba21-e2f2ffe49bd0 · inbound

Shap-E: Generating Conditional 3D Implicit Functions cites this paper.

Shap-E: Generating Conditional 3D Implicit Functions Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:32:06.692336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T15:32:06.563955Z digest=sha256:6d14466775b025eb57cda2bcb34aefa439f8b132ca5a68bb6e96b8a54e112b42

Observation 57b34476-e85b-4cee-91df-5e3fd4d70b43 · inbound

IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models cites this paper.

IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:49:31.747550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T23:57:08.818348Z digest=sha256:10ad59e26e16aebc20ec97d3b648b95d451a155f2e8643354ea14933e151f8ab

Observation 55ba762c-7bae-4b10-bcc2-99ac10481875 · inbound

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation cites this paper.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T20:06:44.602357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:647a8f4eabc548e5b5c9f86f600c728bcae55f20c09776e5eb896f8a06c37148

Observation 96a611e9-5731-4105-b9c7-142c9356f31a · inbound

Learning Interactive Real-World Simulators cites this paper.

Learning Interactive Real-World Simulators Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 172

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T02:15:18.629848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-16T02:15:18.265190Z digest=sha256:e6c1a31284d70d6343be36a9379cb4903184aedb785d38e97fd3a8ed74d6d3ea

Observation a8853952-f1e1-4838-b391-1c392deb9ec6 · inbound

VideoCrafter1: Open Diffusion Models for High-Quality Video Generation cites this paper.

VideoCrafter1: Open Diffusion Models for High-Quality Video Generation Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:40:44.160267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T21:40:43.956642Z digest=sha256:07f575194f6a80b264d2dad8704469ce80b60db8b97a2e5dfd31c3f3579d4dfc

Observation aacf55c7-ee86-4337-8e17-dfe53b47619e · inbound

Gemini: A Family of Highly Capable Multimodal Models cites this paper.

Gemini: A Family of Highly Capable Multimodal Models Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 127

Resolution
verified exact
local_arxiv, observed 2026-05-24T05:03:55.541087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-24T05:00:28.453838Z digest=sha256:be6d5f34790cde70cc4d5f6ca6e26a945d3e002d67c2e8db8f8f9001416cc4fe

Observation 817bb176-c1ba-4fce-9048-b41dca849e31 · inbound

VideoPoet: A Large Language Model for Zero-Shot Video Generation cites this paper.

VideoPoet: A Large Language Model for Zero-Shot Video Generation Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 39

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T17:51:05.614606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:aa2aa69bfcd5136c024cf8ca1612f4897ae2382948180fff726d46ed45a8bfbe

Observation b8e9fb38-b93f-451d-9b04-9905280eb14c · inbound

World Model on Million-Length Video And Language With Blockwise RingAttention cites this paper.

World Model on Million-Length Video And Language With Blockwise RingAttention Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:36:57.239048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T06:36:57.165551Z digest=sha256:acdbe96c78d7d824cf09780283a35a03e177b040f2f3c582d3319b3eeb53fa4b

Observation 0aa3e213-fc17-46af-bd23-d94834f8e3e8 · inbound

Scaling Rectified Flow Transformers for High-Resolution Image Synthesis cites this paper.

Scaling Rectified Flow Transformers for High-Resolution Image Synthesis Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 194

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T08:27:53.634088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T08:27:53.446686Z digest=sha256:8c9b0812683c7a8b450912a86a0fa4b37c4da7668b5eceed960f0c21a40a786b

Observation baa85def-08ce-4b23-a374-c1247054efca · inbound

ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment cites this paper.

ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T04:49:31.747550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T19:43:03.310237Z digest=sha256:25e4bdc1539cf5715276ee5ff0b460035181d7e8b48989c3789fd0eae5c21a27

Observation 8b402d7a-4e61-42f6-92bc-b7801c4c48ec · inbound

Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation cites this paper.

Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:49:31.747550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T22:09:16.622717Z digest=sha256:75e5a94f6c4e77969fddb50fe22edd052f6cacae969aa4a38bfc7bb396caf1ae

Observation 70d2c301-4bec-4eac-b181-731fa5545c0e · inbound

Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model cites this paper.

Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:57:26.993448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T05:57:26.887069Z digest=sha256:9363fb8bc2e121291dde8e0c97e83cdb098783e6c823ec09d211bb277490c1b2

Observation 4c09ecad-160f-4d44-a2bc-0dbbd3b4c369 · inbound

Diffusion Models Are Real-Time Game Engines cites this paper.

Diffusion Models Are Real-Time Game Engines Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T12:04:44.225014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-16T12:04:44.162343Z digest=sha256:3dba60d6bb00e3bd9343e5f730cee61a1da9ffcae6c4f5d076937fa17c7c2255

Observation 74c55ca5-f052-4dc5-a716-2cd715a152f6 · inbound

Emu3: Next-Token Prediction is All You Need cites this paper.

Emu3: Next-Token Prediction is All You Need Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:49:31.747550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T10:56:06.418360Z digest=sha256:55a8b275cfc7bc189894cacb0c0b9643a713ece7a0a9695eefe8510598fb0b3b

Observation 9d8816d3-7bb2-4e7a-9683-94a1d219df76 · inbound

T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive Concepts cites this paper.

T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive Concepts Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-23T08:02:43.331108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-23T08:00:12.781392Z digest=sha256:552ade1044a33a7fc955b6aac676f9f1bc3155752029b49187e0de36b3af843f

Observation a131a26e-66f1-48e2-a545-bc3cf6dc6b16 · inbound

Autoregressive Video Generation without Vector Quantization cites this paper.

Autoregressive Video Generation without Vector Quantization Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-17T15:07:39.829441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:1b01dfc4181ae3275452ff1ee5de2eeff64343d25cec26bd99219915beb7ba09

Observation 06afcad5-ae64-4a94-af83-a593239ac90c · inbound

1.58-bit FLUX cites this paper.

1.58-bit FLUX Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T04:39:36.759706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:39:36.759706Z digest=sha256:7022f0f864a9cc3374ae88adc7dab791af66a406b29383102e79b7be1b5d6994

Observation ada49843-71c3-412e-853f-228e19ea5ce6 · inbound

Is Your Text-to-Image Model Robust to Caption Noise? cites this paper.

Is Your Text-to-Image Model Robust to Caption Noise? Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T00:18:36.849864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:18:36.849864Z digest=sha256:6f0c9f6b4bcd063ea8ecf212c1862a5fa2d7398f9f16178285c1edd117f629a9

Observation b81c66ec-e33c-4b39-9591-18074c91767b · inbound

ReNeg: Learning Negative Embedding with Reward Guidance cites this paper.

ReNeg: Learning Negative Embedding with Reward Guidance Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T00:13:43.677898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:13:43.677898Z digest=sha256:9dc8d99ddbd3fd4535853e2b4df29eab57dfd93d8293a52e0ad316b3a2faec51

Observation 5037f6f1-1b70-42ef-8380-5381a1c8bd05 · inbound

VMix: Improving Text-to-Image Diffusion Model with Cross-Attention Mixing Control cites this paper.

VMix: Improving Text-to-Image Diffusion Model with Cross-Attention Mixing Control Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T23:14:40.708574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:14:40.708574Z digest=sha256:e1ce31a8c41099b995af4cff4d28b84cfdc9da24c57abc9c02106bbb1b8240f6

Observation 5f838876-fc7a-4165-a8d3-daadcbbee38b · inbound

Text-to-Image GAN with Pretrained Representations cites this paper.

Text-to-Image GAN with Pretrained Representations Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T23:05:13.746565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:05:13.746565Z digest=sha256:b72c1b633ab808e0e27621ba9313092d25a9d3dbd0f46bee82b6f642c6b6a27d

Observation 52e35fa3-3958-4a53-9135-94994b442840 · inbound

Text2Earth: Unlocking Text-driven Remote Sensing Image Generation with a Global-Scale Dataset and a Foundation Model cites this paper.

Text2Earth: Unlocking Text-driven Remote Sensing Image Generation with a Global-Scale Dataset and a Foundation Model Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.541656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.541656Z digest=sha256:a0f178d7c8ef5568469d59acdfa15b4221f1c938cac3dfdf9ffff45823d48536

Observation 98c4c712-4d37-4459-8c0e-81c7d8caeaeb · inbound

A Novel Diffusion Model for Pairwise Geoscience Data Generation with Unbalanced Training Dataset cites this paper.

A Novel Diffusion Model for Pairwise Geoscience Data Generation with Unbalanced Training Dataset Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T22:43:31.714764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:43:31.714764Z digest=sha256:c09ee8983743f3918d0b83708b1bb8588ca5ac436f321894b88fda40ae3ade08

Observation 3b5862f0-521a-48bb-9464-57df2ad73499 · inbound

INFELM: In-depth Fairness Evaluation of Large Text-To-Image Models cites this paper.

INFELM: In-depth Fairness Evaluation of Large Text-To-Image Models Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T23:48:02.416943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:48:02.416943Z digest=sha256:bfa4ff3e5b47fcfdd61bd629bcb34b81b247195b7eedbbad310ef7fc8b29f9ea

Observation 6e3fefee-ef75-4a8a-aefa-a8f2b5e92e24 · inbound

DGQ: Distribution-Aware Group Quantization for Text-to-Image Diffusion Models cites this paper.

DGQ: Distribution-Aware Group Quantization for Text-to-Image Diffusion Models Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:07.106704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:07.106704Z digest=sha256:8bfcfa1bd9d8f4e8ba7e0d61163439adc869232bc887d5dadebac64972077ae8

Observation 231f247a-8df4-488a-a358-63b542d6eb62 · inbound

EditAR: Unified Conditional Generation with Autoregressive Models cites this paper.

EditAR: Unified Conditional Generation with Autoregressive Models Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-10T21:29:20.140153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:29:20.140153Z digest=sha256:121a328fee277be78eafd53861a9fc9428fa96ce02bd3dd9c6ebf112ec193b82

Observation ec1ca818-d779-4d23-80bd-47d78f7af62e · inbound

Dissecting Bit-Level Scaling Laws in Quantizing Vision Generative Models cites this paper.

Dissecting Bit-Level Scaling Laws in Quantizing Vision Generative Models Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T22:04:24.330890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:04:24.330890Z digest=sha256:465cf5191d7c9fb19cae7050f14fc0315eb6829825fc214c5f088cfe98427fb5

Observation 0c78d1d8-d5ca-49c4-acda-5720336a061a · inbound

Focus-N-Fix: Region-Aware Fine-Tuning for Text-to-Image Generation cites this paper.

Focus-N-Fix: Region-Aware Fine-Tuning for Text-to-Image Generation Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T21:04:57.554821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:04:57.554821Z digest=sha256:102fb26d1690ce80a6537e45ba417b04b108b5e6ec31be3176f8eebb5179e98f

Observation 8d406714-7fd8-407b-ba3f-d9a77ed1373d · inbound

RepVideo: Rethinking Cross-Layer Representation for Video Generation cites this paper.

RepVideo: Rethinking Cross-Layer Representation for Video Generation Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T20:16:58.637567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:16:58.637567Z digest=sha256:d8a718187ce489dd8c1ecaf1ee7a849112d6093f64634ab711dff4d0d4052f91

Observation 6737044d-731d-4629-8795-d114f4c6a526 · inbound

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation cites this paper.

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:08.234771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:08.234771Z digest=sha256:8eb824c7645d195b8573f255f2db504bba527a8aeb7870678aa16e902cd9bf9f

Observation ba7fd32d-d8a2-40a2-b46b-84d79d47ff3d · inbound

EchoVideo: Identity-Preserving Human Video Generation by Multimodal Feature Fusion cites this paper.

EchoVideo: Identity-Preserving Human Video Generation by Multimodal Feature Fusion Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T16:12:00.043560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:12:00.043560Z digest=sha256:8a6d5dc8a06ad526e25259ddc031452f7cbe1b2d929d258f6ef74da83c4cfe7c

Observation 10880191-3d47-4747-ab61-163f8a62be6b · inbound

Diffusion Autoencoders are Scalable Image Tokenizers cites this paper.

Diffusion Autoencoders are Scalable Image Tokenizers Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-09T22:56:36.227366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:56:36.227366Z digest=sha256:a828247c89319d13a96a02f6b6e10729a2f6e6c130fd7259ee2d2333e9d7edd5

Observation 37fbd40b-1879-4e8a-8833-4db61cd71d61 · inbound

Visual Autoregressive Modeling for Image Super-Resolution cites this paper.

Visual Autoregressive Modeling for Image Super-Resolution Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T21:47:59.592777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:47:59.592777Z digest=sha256:6e50d99d4b3e0c71119af08bd5fd62e88ba7e244fb7f2439b23266ed9ec85a2d

Observation 2b410c8c-d285-4b2b-b65a-cc13d2ce4079 · inbound

CAT Pruning: Cluster-Aware Token Pruning For Text-to-Image Diffusion Models cites this paper.

CAT Pruning: Cluster-Aware Token Pruning For Text-to-Image Diffusion Models Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-09T19:05:52.999296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:05:52.999296Z digest=sha256:32084193c93cc8fe15e66a7df506e32d6f460b90c678f36c40a4802f9352620e

Observation 433754fd-2c16-4c97-8c20-c822fb0e8dbc · inbound

HuViDPO:Enhancing Video Generation through Direct Preference Optimization for Human-Centric Alignment cites this paper.

HuViDPO:Enhancing Video Generation through Direct Preference Optimization for Human-Centric Alignment Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T17:34:10.212570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:34:10.212570Z digest=sha256:5e33142b271b7566ce5e371f8482a97fbdb11167bde4111491bd73ce3d7f9661

Observation 58144a33-5b5d-476b-a589-da6a641d948e · inbound

Variational Control for Guidance in Diffusion Models cites this paper.

Variational Control for Guidance in Diffusion Models Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T04:11:42.108293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:11:42.108293Z digest=sha256:26de600724a50f17d19c0e3320ae23ed7695425a1a2ce9bd883ab9474dd686f3

Observation 93633c06-e13d-4cc0-b4a6-81db2175d86c · inbound

Decoder-Only LLMs are Better Controllers for Diffusion Models cites this paper.

Decoder-Only LLMs are Better Controllers for Diffusion Models Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T23:56:18.756524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T23:56:18.756524Z digest=sha256:ce9c85dc5c6cf7543ae8dd05836184961564748886c86b24f84ea488bbb1e105

Observation eee958a3-ce80-4dc1-aa91-877675b86481 · inbound

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths cites this paper.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.475345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.475345Z digest=sha256:c365aed31312dfd35cb71961fad5b2f4a406aafab9ed423d2a0cf65c1ec19816

Observation 9a02b24a-7b41-4465-a473-9e6fcfb9b7c4 · inbound

EQ-VAE: Equivariance Regularized Latent Space for Improved Generative Image Modeling cites this paper.

EQ-VAE: Equivariance Regularized Latent Space for Improved Generative Image Modeling Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T21:17:55.978646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:17:55.978646Z digest=sha256:8da634dcf98c08378ce71e5bf8939813dbdb0cfdc57a12aee5dc200840e60a05

Observation 841c55ce-4a5c-40b1-9b5e-61b6c7470d7b · inbound

Generating on Generated: An Approach Towards Self-Evolving Diffusion Models cites this paper.

Generating on Generated: An Approach Towards Self-Evolving Diffusion Models Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T20:00:34.088913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:00:34.088913Z digest=sha256:2bb8b781b0d931480ed6d2745291e4a903512c6149884a28f2d3dae72a6c618f

Observation 5edf496e-fb40-4a8e-91f7-514112de257b · inbound

RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control cites this paper.

RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T19:40:21.443079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:40:21.443079Z digest=sha256:8a295144f0a9f53e65e0c5161eab3bb9e7c6f1cf3c7920a4d8c8d79c1510754d

Observation 8420c32b-5fe0-4f3c-b92a-0c0b3959ff58 · inbound

Unified Reward Model for Multimodal Understanding and Generation cites this paper.

Unified Reward Model for Multimodal Understanding and Generation Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-14T00:44:30.783001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:44:30.558048Z digest=sha256:69a02e9786e5ac39b382a2a51080274e6468a9e0ced2bb60f5a7a11ca09e4401

Observation 2c46995d-b210-4425-8d3d-3a86648220c0 · inbound

Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model cites this paper.

Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-17T08:27:36.364653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T08:27:36.242416Z digest=sha256:94012235886ab1dfd5a747d04e92ae873ca2991448a643c9db6369d09822fa88

Observation 9b6dde2a-03da-44fb-b40a-dcc78074431f · inbound

BalancedDPO: Adaptive Multi-Metric Alignment cites this paper.

BalancedDPO: Adaptive Multi-Metric Alignment Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-22T23:42:16.581410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:37:55.154902Z digest=sha256:7a9ea1d1c1ce8706b1b8af274c203ffee153a270d28d3c5167f0e7b67ae1f10a

Observation cc27188b-03e6-4d7d-8ea5-b7d05dbd53ec · inbound

MSDformer: Multi-scale Discrete Transformer For Time Series Generation cites this paper.

MSDformer: Multi-scale Discrete Transformer For Time Series Generation Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-22T13:44:52.871888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T13:44:08.891339Z digest=sha256:188d8928ed0af3084ed4d0d7a76269fd7770e361be08ebe1be472c5bb0738993

Observation dd326d28-e7c6-48f9-b32f-770ebd37379c · inbound

An Exploratory Study on Multi-modal Generative AI in AR Storytelling cites this paper.

An Exploratory Study on Multi-modal Generative AI in AR Storytelling Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 112

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:02.518907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:02.518907Z digest=sha256:5aecf1ed4100548dd3070b1131388686ca77365624dae3451002860b64b187b2

Observation 78e39d37-354c-4b90-8b8a-310c159b4d0f · inbound

DetailMaster: Can Your Text-to-Image Model Handle Long Prompts? cites this paper.

DetailMaster: Can Your Text-to-Image Model Handle Long Prompts? Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:31.570990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:31.570990Z digest=sha256:ff47fc7c39aca49d1a2f1842254c5335ef7c8f1f6d8bd436d0565028e679b064

Observation ac01512f-127b-4625-9706-5d0f8abe1555 · inbound

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation cites this paper.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:21.180996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:21.180996Z digest=sha256:60fcf456ae6530fe8f1bf61155d4366eaf6d08391787ea6e00f5c5c99fc2bcc3

Observation a7d720ca-b63c-439a-8d53-88cc3e268380 · inbound

LlamaSeg: Image Segmentation via Autoregressive Mask Generation cites this paper.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:55.721329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:55.721329Z digest=sha256:0c2cc628b6a97d01c18b2d41680f4cb33a955141e7e16c9cfb73054b2156a22e

Observation 81b0344d-10a9-4bc5-b6e1-085584823e6b · inbound

One-Way Ticket:Time-Independent Unified Encoder for Distilling Text-to-Image Diffusion Models cites this paper.

One-Way Ticket:Time-Independent Unified Encoder for Distilling Text-to-Image Diffusion Models Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:39.862109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:24:39.862109Z digest=sha256:f9d6838880a75374e3551f6f99f8f4b403e94a67d78c11902dcb98a581a2fd0d

Observation 6444d429-66b4-4131-8379-fd37e72ace47 · inbound

Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human Preferences cites this paper.

Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human Preferences Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:58.701288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:58.701288Z digest=sha256:5fce679e4c56a4047b4c549fa111cbb6832f6067d49c72756e52eef2d51ff294

Observation aa682b6b-9d61-4183-9d58-2f8001db09e3 · inbound

Native-Resolution Image Synthesis cites this paper.

Native-Resolution Image Synthesis Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:27.931555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:27.931555Z digest=sha256:62e03d0a2a6f0ebdb634e03f206ce8ae645b13857b9527bf04e65fa515e0eb6e

Observation fbe66dd3-b9f1-4061-9032-6aedf59e49c6 · inbound

How Far Are We from Generating Missing Modalities with Foundation Models? cites this paper.

How Far Are We from Generating Missing Modalities with Foundation Models? Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-25T08:15:33.614999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T08:15:12.947854Z digest=sha256:8ab653704584812ae1777125e007115e35cc29e472d4f74a50eb98acb5578ce9

Observation b7f27687-3010-43f3-bdbe-4f0cc2ef0e23 · inbound

Images are Worth Variable Length of Representations cites this paper.

Images are Worth Variable Length of Representations Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.407403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.407403Z digest=sha256:92a2f02f8c9b095ce5f79f2db9adf308efebf4b1be44be1a9ab8c58b4d947034

Observation 25a8ef6f-7c96-4537-acab-f6ba0e316079 · inbound

OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation cites this paper.

OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:44.500274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:44.500274Z digest=sha256:8dd025dc7137e8e7a6a0a49d111eb6de4317999fbfc37fc7c2eb72db74e670de

Observation 65d4e749-b8d0-4fda-93f8-6157a89cb6a3 · inbound

CuRe: Cultural Gaps in the Long Tail of Text-to-Image Systems cites this paper.

CuRe: Cultural Gaps in the Long Tail of Text-to-Image Systems Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:16.277893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:16.277893Z digest=sha256:16a42529714fa84ea996b677ebb2d2303f41fb611b837a8a690306f81d18ba5e

Observation 3980dc5b-eeb1-416d-93a5-624f7010b676 · inbound

SpectralAR: Spectral Autoregressive Visual Generation cites this paper.

SpectralAR: Spectral Autoregressive Visual Generation Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T04:17:51.824999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:17:51.824999Z digest=sha256:c507c491bb33f70e42f640e533851f108fe41a428246ccacdda89e25936b7283

Observation da2819f2-d93e-4d5c-be32-2e15503fb3be · inbound

MARch\'e: Fast Masked Autoregressive Image Generation with Cache-Aware Attention cites this paper.

MARch\'e: Fast Masked Autoregressive Image Generation with Cache-Aware Attention Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:51:35.831619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:51:35.831619Z digest=sha256:1de225da15b7f1566cb13783ea41c07beb04aaa2fadd828a6da4d1adfbcf6c69

Observation 49f672d8-3b0d-4384-9cab-82399c4fea80 · inbound

A Minimalist Method for Fine-tuning Text-to-Image Diffusion Models cites this paper.

A Minimalist Method for Fine-tuning Text-to-Image Diffusion Models Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:17.476303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:17.476303Z digest=sha256:d5a275ee98ad28e7afefb11a20be5fac0016f5c60f21f5dcacb1b9b7b808f756

Observation e979a7b2-90d9-4ec4-8946-4ae5a8d52129 · inbound

Disentangling 3D from Large Vision-Language Models for Controlled Portrait Generation cites this paper.

Disentangling 3D from Large Vision-Language Models for Controlled Portrait Generation Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:08.113817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:08.113817Z digest=sha256:0dd09b0bfa0d2fe3e666072e0a24dfcc5770abb244b04e37c5a9565e7436852b

Observation 9e9d6904-a56d-425e-8ee8-95ff170090b6 · inbound

FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space cites this paper.

FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:49:31.747550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:36:02.016696Z digest=sha256:08ce10bea29fe4a38607a7f8f5e049ccea0561209b1ee552c0fb2c2549167515

Observation 78361dec-7c30-48e3-bd4c-a52ff3363c04 · inbound

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation cites this paper.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:08.742835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:08.742835Z digest=sha256:fa5ee08ff8365ff9210e50ebb59587250951215797069da8891b56b439c3035c

Observation cc175dbd-a9e3-4cb2-b9fc-a48af257ccf8 · inbound

CPAM: Context-Preserving Adaptive Manipulation for Zero-Shot Real Image Editing cites this paper.

CPAM: Context-Preserving Adaptive Manipulation for Zero-Shot Real Image Editing Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:42:09.177386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:39:02.539294Z digest=sha256:706239abc37adf8133396c87ee5d30a8859eaac027ad31e70360ac324edd2e33

Observation 29cadb39-d80a-4d47-92ca-1aec93da7c87 · inbound

CycleVAR: Repurposing Autoregressive Model for Unsupervised One-Step Image Translation cites this paper.

CycleVAR: Repurposing Autoregressive Model for Unsupervised One-Step Image Translation Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:45.318496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:45.318496Z digest=sha256:a340357ed151d16d100d2febf2b2c5e9c207ade96221914ec2e5946b904ff4c9

Observation 2753b0f3-2511-4949-9bd2-e5eb3bcba558 · inbound

Transition Matching: Scalable and Flexible Generative Modeling cites this paper.

Transition Matching: Scalable and Flexible Generative Modeling Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T21:46:30.140725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:46:30.140725Z digest=sha256:6c2d9600115a82bf9999f8654c21017fc52d5930c5958ac5f0ce0ba534ba1919

Observation b7568541-b979-4092-a208-5c609a598400 · inbound

SD-Acc: Accelerating Stable Diffusion through Phase-aware Sampling and Hardware Co-Optimizations cites this paper.

SD-Acc: Accelerating Stable Diffusion through Phase-aware Sampling and Hardware Co-Optimizations Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:32.534698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:32.534698Z digest=sha256:15ab0212abfed71c2f2a1179baeeb4ca479cdfc8e6ef0faf37aa95f9723a023a

Observation efce1fec-01bf-4eb8-b361-4439f5f9c537 · inbound

Hita: Holistic Tokenizer for Autoregressive Image Generation cites this paper.

Hita: Holistic Tokenizer for Autoregressive Image Generation Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:58.143921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:38:58.143921Z digest=sha256:e043cfe2deb6fed3d1cb1610a9e268b6cebfcba589308c1511306b0f95b6838d

Observation 64a470f1-3ccf-455e-8aaa-69b0484e7651 · inbound

APT: Adaptive Personalized Training for Diffusion Models with Limited Data cites this paper.

APT: Adaptive Personalized Training for Diffusion Models with Limited Data Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:30:26.552436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:30:26.552436Z digest=sha256:2e5410122b3e2e463bd0e1a64ab213aa2cd644c068040c9172d37323daa5153e

Observation 39d586ed-687b-4e0b-aa4d-78c1e53c3648 · inbound

TextPixs: Glyph-Conditioned Diffusion with Character-Aware Attention and OCR-Guided Supervision cites this paper.

TextPixs: Glyph-Conditioned Diffusion with Character-Aware Attention and OCR-Guided Supervision Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:49.550557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:18:49.550557Z digest=sha256:0bac166ae686ef03b8322dc170ec109d83d8359b4bd057a6d395da7b9aa4b93d

Observation 74bafdb0-3aeb-4e06-bed6-c32db1b6805f · inbound

Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models cites this paper.

Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T18:54:11.615639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:54:11.615639Z digest=sha256:67edb5de0f35705de416b3d915e736d833be50e57ff92074be0624ce749e8701

Observation dba5b899-6ab5-4acc-996d-fa96cf81ad29 · inbound

Automating Evaluation of Diffusion Model Unlearning with (Vision-) Language Model World Knowledge cites this paper.

Automating Evaluation of Diffusion Model Unlearning with (Vision-) Language Model World Knowledge Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T19:06:40.558141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:06:40.558141Z digest=sha256:b74f18e0cdc65f10f9a809c4dbe9d437b3739a7d8bbf57d6ff88d31a17d15bbe

Observation 41877635-a953-4336-a07f-49dc001d2e7c · inbound

Hybrid Scandium Aluminum Nitride/Silicon Nitride Integrated Photonic Circuits cites this paper.

Hybrid Scandium Aluminum Nitride/Silicon Nitride Integrated Photonic Circuits Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T10:17:00.900106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:17:00.900106Z digest=sha256:1a6ab452359ead8d4de4e0c424fdc8d01f3a516d594259617a9c2f890a6e3355

Observation 03bc6434-d9c6-47eb-bd98-575a24fee16c · inbound

Steering Guidance for Personalized Text-to-Image Diffusion Models cites this paper.

Steering Guidance for Personalized Text-to-Image Diffusion Models Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T10:18:02.963300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:18:02.963300Z digest=sha256:434c38d42783435e20b8abc94c575d6a17a47761767c26769f070848aa048baf

Observation 75f515f5-2b13-4031-b42c-f47808b2960d · inbound

ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation cites this paper.

ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T05:59:15.646747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:59:15.646747Z digest=sha256:fdebbb28e5dff3ff7541e3fe8335b38e854d4f13ff905ec5c6375eed484cb202

Observation 2c517ff5-5419-457a-a318-b2c96d8ae3b0 · inbound

Zero-Residual Concept Erasure via Progressive Alignment in Text-to-Image Model cites this paper.

Zero-Residual Concept Erasure via Progressive Alignment in Text-to-Image Model Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T00:04:32.821331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:04:32.821331Z digest=sha256:78621be4dc55b9d2b1899d9a23fb8d98d6b085b5310ee90cddb20e3dbac37fe1

Observation 73bcffa3-7f25-404f-b2be-affd62948699 · inbound

SPRINT: Robust Model Attribution of Generated Images via Secret Pixel Reconstruction cites this paper.

SPRINT: Robust Model Attribution of Generated Images via Secret Pixel Reconstruction Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-19T00:42:54.495610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T00:42:01.640233Z digest=sha256:d2157bb1df087d160bc42e63d36c747cc20520d70869a35fe87050caa23f5678

Observation 18cde8cf-8778-47db-b12c-1cf1955b78c1 · inbound

Per-Query Visual Concept Learning cites this paper.

Per-Query Visual Concept Learning Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T21:18:53.874431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:18:53.874431Z digest=sha256:b245088253ad5467152f494ce6bdd3978710b080cc13649fbfe881c0dd8034a9

Observation dd892bda-310d-41cd-bce9-12f7aee75dd0 · inbound

HumanGenesis: Agent-Based Geometric and Generative Modeling for Synthetic Human Dynamics cites this paper.

HumanGenesis: Agent-Based Geometric and Generative Modeling for Synthetic Human Dynamics Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T20:49:53.234073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:49:53.234073Z digest=sha256:774aec80735b36b86353aff12eae829da69ca6f203f14b6f49549f532fc8df8e

Observation a6acb6ac-e0eb-4ca5-9beb-5f7d217b9e78 · inbound

CTA-Flux: Integrating Chinese Cultural Semantics into High-Quality English Text-to-Image Communities cites this paper.

CTA-Flux: Integrating Chinese Cultural Semantics into High-Quality English Text-to-Image Communities Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T18:40:37.210646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T18:40:37.210646Z digest=sha256:93c7aa1515e124a62d9ff29b69a9262158f6f1e10b9e8003078dd2011bcb8c5d

Observation 08260805-81a1-48e1-a37a-25de76424aa9 · inbound

FICGen: Frequency-Inspired Contextual Disentanglement for Layout-driven Degraded Image Generation cites this paper.

FICGen: Frequency-Inspired Contextual Disentanglement for Layout-driven Degraded Image Generation Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T12:59:28.474779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:59:28.474779Z digest=sha256:dc7bb472dbde808f5e8bafea1b00ecd48b650112ff46bbdd5796d24620ac4e2c

Observation d2e44fef-7cac-4680-9705-e9a6e449ee52 · inbound

Discovering Divergent Representations between Text-to-Image Models cites this paper.

Discovering Divergent Representations between Text-to-Image Models Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T19:59:32.102328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:59:32.102328Z digest=sha256:b7bd33f0c75e459d6585e9b1ba18c4d0278ec0727601790266f70600e339530e

Observation d6030455-2072-4d5c-8360-fa3b2aee2280 · inbound

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark cites this paper.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.872017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.872017Z digest=sha256:796b492ead4232d27ba3834a02c3b451404e1d7192e277e4fb85847dc1bb7a69

Observation 9a5a0829-9c42-4048-9b53-2d9376339568 · inbound

Maestro: Self-Improving Text-to-Image Generation via Agent Orchestration cites this paper.

Maestro: Self-Improving Text-to-Image Generation via Agent Orchestration Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T17:40:28.525639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:40:28.525639Z digest=sha256:31420149b14d99ba4e79d6ff3b4d03d3c228fb129e540d526af48dcf88d14033

Observation 1c0875f4-5e9b-44e3-b2cb-6dfe9f159174 · inbound

IAR2: Improving Autoregressive Visual Generation with Semantic-Detail Associated Token Prediction cites this paper.

IAR2: Improving Autoregressive Visual Generation with Semantic-Detail Associated Token Prediction Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T11:08:01.939270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:08:01.939270Z digest=sha256:4de1b0510852c1865f5a92433a40a166d8a1e4673dc208ed56a83192752deac3

Observation dfcec9ce-0fc9-4f58-ad60-e2efb6c7d603 · inbound

RubricRL: Simple Generalizable Rewards for Text-to-Image Generation cites this paper.

RubricRL: Simple Generalizable Rewards for Text-to-Image Generation Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T20:15:31.231928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:15:31.231928Z digest=sha256:20f4fb9c36ccbc706d90ca21883c60f640930cd5c98ee0cd320ca8171179d3e1

Observation 8cbaf14d-58a5-4971-b0b9-a58207b66585 · inbound

GeoLoom: High-quality Geometric Diagram Generation from Textual Input cites this paper.

GeoLoom: High-quality Geometric Diagram Generation from Textual Input Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T17:49:11.421158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:49:11.421158Z digest=sha256:a4c9ee32761d08f47a7c81d2db62436ebf0f33d82c823f6aa79157110bd3779e

Observation a373d2fd-0024-44e0-b6ff-7250cd6a1bdc · inbound

FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive Models cites this paper.

FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive Models Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T15:36:52.177245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:36:52.177245Z digest=sha256:939e453ce9136dfedce58c4c9286e0c3b00884420d623f674764f190e898e7aa