Pith. sign in

Paper Citation Record · LEDGER

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

As of 6 August 2026, this Paper Citation Record lists 100 of 291 outbound references and 89 inbound Pith citation observations for arXiv:2310.05737.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.05737 v3

Coverage vector

measured 100 of 291 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T20:06:44.480769Z

measured 189 of 189 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 89 of 89 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:33:45.572085Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 291 outbound references displayed

  • verified exact5
  • verified fuzzy19
  • unresolved62
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch14

External citation measurements

21
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 0c80a096-a91d-4cb5-8c3a-a40bdc115bc7 · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-05-14T02:23:39.690311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:a2e67d4085fb314ae40fab8c990f6ce90b2935451ae02f3e1692e761af3b5609

Observation b7931924-488d-4b0c-af73-ca8b858e9a17 · outbound

This paper cites Long video generation with time-agnostic.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Long video generation with time-agnostic

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:23:39.657081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:4dc204708b19d328b9a1ceba92381492030c2b6ec53f80e1499aff133f3641d7

Observation b43eb763-81f7-403c-bfa1-92a7e356b3f2 · outbound

This paper cites VISIGRAPP (5: VISAPP) , year=.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation VISIGRAPP (5: VISAPP) , year=

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:23:39.686044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:e08a71f2a5736600170429584c162992d1060e8a63457fd62cf4770938e4b239

Observation 73449078-4699-4e37-a9b0-8569183a8d93 · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-14T02:23:39.694139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:ee5db4cb99f712a48123f72b0a37e2a1c352afbc84776f604077f2097e36c730

Observation b2f1cea4-7721-45ee-b677-c6eda590348c · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-05-14T02:23:39.679663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:7d204faa25f3f916fa0f8626166e408278f453438598a65135bb7dd7958f56b1

Observation ec3178e0-2d1f-4e48-a4f7-3633c8d864b5 · outbound

This paper cites Adversarial Video Generation on Complex Datasets.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Adversarial Video Generation on Complex Datasets

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:06:44.731017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:c5a214457bfc62d952eeba587b1d93d81f9512fe54dd91ebf1365f800d65b029

Observation e4311910-bf4e-46a0-be6a-52caa1ce0a3f · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-14T02:23:39.707721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:0c511887dfa17f915dac831331aa96125b0faa260865a156b1bc3f3d6c86c12d

Observation 562299c7-96a4-49eb-8918-9755296d3976 · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-05-14T02:23:39.668027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:50bc965dca933d1fbaf500b3b377a5bb8f276dcab76ea85fa9e3f48b86705b84

Observation 61733854-0bed-430c-a775-3bf8c20d9551 · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-14T02:23:39.671869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:15b8e0d942134089b8808f35ab3256669f463b3fa7528f77e96ea28363c6651f

Observation 435d4781-8562-4efd-ad1b-287f074bbf80 · outbound

This paper cites Transformation-based Adversarial Video Prediction on Large-Scale Data.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Transformation-based Adversarial Video Prediction on Large-Scale Data

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:06:44.593279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:3633d53a925c963a718986698b43c839e58a2bbc7ad5e17c93f92410cdb61e57

Observation 3783bab3-398c-43bd-af34-e7bddc7a3748 · outbound

This paper cites Transframer: Arbitrary Frame Prediction with Generative Models.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Transframer: Arbitrary Frame Prediction with Generative Models

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:06:44.612652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:e1f0886cc90d527d13e75d4cc1735b47207245567757b8dbf308ba45a575a49d

Observation 85bd2d48-26b8-4b67-8c18-6919797453e9 · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-14T02:23:39.660912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:0580e7da805e728140a255809a1c0cc3b2ba831f2a72e00e8779ec072277c9b7

Observation 69b97509-1433-41ef-a792-cb9db08b0e5a · outbound

This paper cites FitVid: Overfitting in Pixel-Level Video Prediction.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation FitVid: Overfitting in Pixel-Level Video Prediction

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:06:44.737606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:83c4887b533aa83267c25f56bac6b68f959445621530870e1fbfd5226b688ea9

Observation 80785f33-4eb3-4a86-bc2f-8e8fd8fe5100 · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-05-14T02:23:39.697480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:69013c3f89664ab12f37f31c62a744b6f60c46687313b74ba9905498bbe3a2ff

Observation ba1b4fa5-76e5-4efe-b64f-3d2fdfbbb781 · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-05-14T02:23:39.711609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:4e0c103cbc68f0b4188f378d7cb0885080f3f98ee354a179fc307fc9040edbb0

Observation db166f65-efce-4cd9-822e-802c0873a13f · outbound

This paper cites 2020 , publisher=.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation 2020 , publisher=

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:23:39.644414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:dc033ec492d157f137863034ab3bfe56f6ca1ae68fb361b5755e0eb00f440478

Observation c4afdd8b-f6ea-4103-b667-d28d6b56c287 · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:08:13.791471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:498551cffa85978912fd028d6d8a8d8a4a9db7f3f8edcd98395674dd23059920

Observation 76450f03-38f1-4d97-b18f-388c5eef2fc2 · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-14T02:23:39.649281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:31e0d02b0819f4235df5d680303a27a012ac7a9174eed3703ee20fd0e5f506cf

Observation a5987eea-9132-4f1b-8748-5ba02ff7c257 · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-05-14T02:23:39.701241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:9cdbc777a22be508ce86798e6736bd27efe0e85d6320d64f3e958b6017ce9f91

Observation 5ce3c670-021c-421e-9969-f1da665ddca4 · outbound

This paper cites Towards High Resolution Video Generation with Progressive Growing of Sliced Wasserstein GANs.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Towards High Resolution Video Generation with Progressive Growing of Sliced Wasserstein GANs

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:06:44.672072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:a20346c11e4bd4bf1a9cd5deee4f2e2a3ea4643261f21b960f31e42feceb99cc

Observation 360805d3-001c-4b92-b92a-c0666815fde0 · outbound

This paper cites Neural Networks , volume=.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Neural Networks , volume=

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:23:39.640743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:61fc4afff419b7cd1a40dea19b2b39ffb47a637151c63de7beedecdccd206bf4

Observation adbfb8d9-cc2d-4079-a90c-fe407772489d · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:08:13.664583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:beff7a81aa61de1c11159a8ad27041bbfe1522295296e4bcd0e4e94ffa9e9433

Observation 540cabe7-8921-4dfa-b7b3-c23895f2d3a9 · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:08:13.789629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:22405b361f0510915118f58dd405ea9a5cfcf85f0ca9fb8f0a20218a1211c4b6

Observation 51e84f06-e777-4160-9cb3-3227db99bce8 · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:08:13.615524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:63c068fb5ab6902e86a7f40f8de27690f52b99b9ace399b45b4ffb4aa2dd8020

Observation 28dfeb95-a301-4119-9d5b-0330b63dd06f · outbound

This paper cites Image as a Foreign Language.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Image as a Foreign Language

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T20:08:13.649381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:23c19fce808c1d73edb792d76f8bf54454fe058ada7f458f2f07651b887f5d2b

Observation 0cfa03d1-471c-4801-9df3-e31889c6964d · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:08:13.597250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:db5aecfb61641a6e13671a1d628428b6d00f956d9f2bcbce27afa011e54b16fc

Observation 4d7c6c92-b0d4-4f42-b9bb-c5851395a31e · outbound

This paper cites author=.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation author=

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T20:08:13.593448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:ba80e01c300e061aafae8eb65b4ee327dc2cfe506e8581d401081fa3c4e6297d

Observation b56e9092-9f65-4fd7-b848-1a84c3704f26 · outbound

This paper cites A short note about.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation A short note about

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T20:08:13.595432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:15b8aedf7ba2255ea9b13ba341041fec3ffcc0a6db8934af610f3b39b7f2c16e

Observation 0af78fd3-a098-4f32-ba20-93ed967d42fb · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:08:13.605216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:d789b0afdf83af398ea9fd99e26bdbf1e84e7279fca5d0d1cf5d7fcdf03ecc7c

Observation 549e54d8-a871-4012-b56f-466558975023 · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 30

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T20:06:44.619893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:8c89a8d9435bb638e5f75136ec7c953be29b789e070f7308fee73b636c888e89

Observation dd86efd3-3b09-4f04-a9cb-3b0cab702a39 · outbound

This paper cites Quo vadis, action recognition? a new model and the.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Quo vadis, action recognition? a new model and the

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T20:08:13.603321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:e362f4dc9c69260a73470bb7161ac1a953d3b0348dbb19029167db738421f8e3

Observation 5c4f39d2-8cc8-4217-92ce-bdbf63ecc526 · outbound

This paper cites Learning spatiotemporal features with 3.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Learning spatiotemporal features with 3

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T20:08:13.647325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:ba107757e73357e820356021446761c707d58f5d347da742c61fd2d1a1ca8faf

Observation f26a62e3-2eec-4825-9313-6fdf5d9eb374 · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:08:13.659725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:4e5982f505fadb954efcd5179661270898d57f70c94d30a11306cbc595deeb25

Observation 6d7e2ccd-d58b-4187-aab6-79934656e29e · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:08:13.676829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:7dce02038bd613d3c08705cdaa47c9218d833a0b1924fac1b0fcf64230c39df6

Observation 55ba762c-7bae-4b10-bcc2-99ac10481875 · outbound

This paper cites Scaling Autoregressive Models for Content-Rich Text-to-Image Generation.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T20:06:44.602357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:e020a9a183fb7431c99c3b3e3aebd925b34786517b1c98a6359d715e1b8ab1f6

Observation af97548c-a3da-4d33-8a76-6f01edfa55eb · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:08:13.720345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:6ff5d2613f3cd2e0f41ec49804be3b5d35daa77ae9e2efde3fc428592a663682

Observation ebab0b46-9af3-4ddc-8e4b-ab8f78f45f73 · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:08:13.759038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:49db85a62af38b3b5818c71b5f61fc20cec3f815543dea4358f3c46be02fc436

Observation f829d4aa-d66d-47f4-a23b-9b54d7ec30bc · outbound

This paper cites VIMPAC: Video Pre-Training via Masked Token Prediction and Contrastive Learning.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation VIMPAC: Video Pre-Training via Masked Token Prediction and Contrastive Learning

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:06:44.647062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:3a33f60aad0571832c8d6dede72ac0dec84c1cdb0a103cfc312ca30c4b88ae63

Observation 8f1974c8-f3f8-444e-b3be-f278697a64ec · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:08:13.754385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:2c4e8051f135f3ea8c555c9f9281f7faf0e8d66e370a366a7aa752b913a64710

Observation 830c8e2c-af25-4b60-b2a8-be90067e1c72 · outbound

This paper cites Vector-quantized image modeling with improved.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Vector-quantized image modeling with improved

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T20:08:13.784897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:d197d616d9aa4faaf488e37b1d59b89499c93adb50e927a80f7034b4f061910a

Observation 6bcd2896-1e50-453f-a84d-64889a2b9190 · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:08:13.746723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:bcc544169dfff3f7ccbe73faa5e4f518a09cf649e2076ee840995df2239e3a1e

Observation 5b285656-be66-4d39-b807-31c6b992848a · outbound

This paper cites 2004 , publisher=.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation 2004 , publisher=

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T20:08:13.713487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:7ff545d16d72360853986a5c078520a4c42ce02b39cb8143aa27457989826922

Observation 3bdf8527-8951-4b98-af43-50290974b30d · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:08:13.727527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:2220f033593d2317b4a8ce45f6a1a19ba3a8771f0a39022bc96459f12b2efcd7

Observation 0147ad35-08fb-4488-8b58-51a060b82629 · outbound

This paper cites Diffusion Models for Video Prediction and Infilling.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Diffusion Models for Video Prediction and Infilling

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:06:44.605737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:cf19cd7e5dfe8a0cfe336035bf804369fe0c29e7427921e5f1ea1cdc201be52c

Observation 1399ec90-b92c-49f2-b18a-a46103967558 · outbound

This paper cites Large Scale.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Large Scale

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T20:08:13.683351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:1309976ca4c73f9a97d7f1d92c168c3615966e065d3981adcdd4e22c46a47624

Observation 3221e029-9328-409f-85af-9bfad0d2d9e6 · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:08:13.662330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:181cc1897e37da6a80ec25a2ed922ba4d0ddb9ea97d5c3f59988a31c0756bd53

Observation 8d8dc07a-2b83-441b-af5e-60c9bb8c6a33 · outbound

This paper cites Improved Masked Image Generation with.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Improved Masked Image Generation with

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T20:08:13.672527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:e61c8fedfcaa6c3dcc01cd86eead4ca7510384de81b7600561b7f514dc36bd48

Observation 642e89e2-de80-4732-bf1a-537c9ca823c9 · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:08:13.678840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:589f5cec331cab766da1ef92d5c908e67bd41248d236ca3c7f05464137e69c66

Observation a30ce423-1cae-49ed-8e3a-6e729dad1fb5 · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:08:13.681202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:d60b7b2fa12356969402ee94f6bc1a85622e54f04a4f9692b7db664e7a3563ec

Observation beece267-aeeb-4f7d-9df3-dea665378fa3 · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:08:13.716934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:bcc1a0d983221b3957988f3626237ffca3f2040dc421b8de10a6b80308232d40

Observation 02a638fb-162c-440b-a9e8-81f9a0af6a01 · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:08:13.781660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:921d3859c1a84b772654004833b641b0b055c431d3c2979aa05b394a68785e97

Observation 4f0bc0a3-07e9-4707-9069-37c8a0f1b06b · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:08:13.617470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:bc8811484732ea71069101afb3c6740d234626bc50c59e18497f8c43d97cc215

Observation 8e72e09d-826e-4c6d-8f4e-d07e3aa96a47 · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:08:13.623615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:7535f74a9f082fa29dcea85d912c3fe30cd07828ca2e4ca90442ef86684afc4f

Observation 7410faa7-1428-45bf-b946-939e8e33034e · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:08:13.633490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:bd639672954dc36cc063862c78d5c49dc5ae5541519dd8c700b2db634639dbfc

Observation a4b67684-62dd-4fcf-9664-841cf7a40eaf · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:08:13.625383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:0d58dc11397050169ea29fe42e71426a6f08d2c4d23dbda1fcc00a784625f8be

Observation 2a5b21a2-6dd5-41fb-99ee-32d8f3611af5 · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:08:13.589570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:14f213665a3528390ac1144cad460927e0855e84df48ee756d5445b7391fec38

Observation 9d1239c0-420c-4668-a6d4-f55a4ef1d2eb · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:08:13.599241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:c14ce6b32e46dc0918fdda989d1e2393613b1fe5c4f402319cbc71bb47aa7244

Observation c332e76e-de41-4af9-b39e-7435fa4e75e9 · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:08:13.627227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:5282494d39b8278f1772d9b9d4af0c62d0e7949be4b1185f8dedc10b6b920fe3

Observation 3c96234c-ddd9-476f-af1e-48202613838f · outbound

This paper cites Distill , year =.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Distill , year =

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T20:08:13.737618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:33eb440ad74ea85031f0e33660098546a530ecec535b7fe63cf147d13e1e4fcb

Observation 6928b5d3-696b-4290-b352-340ddcf69430 · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:08:13.703839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:fdf6aafe7cc57401dd4514ac0359d52403f0c8bd680c91d27a40e91ca7032d6e

Observation 8d1d9120-f4ed-421d-82ac-27c43095cee3 · outbound

This paper cites Ten lessons from three generations shaped Google’s.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Ten lessons from three generations shaped Google’s

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T20:08:13.724298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:23831b963896e8c52735ca366072f032d7719885d5c24257dad73c30262c0b46

Observation 22cc81ab-cdc2-4b48-8694-c393d50d2ecc · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:06:44.740046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:3474ea316046cbe888f8b3dc30c3dc5d9575b546f439506644b17e968ec40fed

Observation c07b3c99-e177-475e-b952-ff8f20cf261d · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:06:44.742600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:71baf2cbb41a87e53f20cbfd79ab82866865fd426d295f26bc5a6b0271c66cb4

Observation 8d28620f-6d99-42c2-8ed6-bd8b69ab02e2 · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:06:44.744813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:c1ed36272183fd9ddbef632b9b0387f58e42e72a84cc11cfde4361286d1a2257

Observation 13bb4a70-5241-49aa-8840-6694e36bfa73 · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:06:44.747020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:25f782c4f44443a78bdd45e639ade5f5643a08a87d7a164a6c005e1148f7a04a

Observation 164eb0b2-8fc1-46e1-9504-cf95cc2410ac · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Adam: A Method for Stochastic Optimization

Reference 66

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T20:06:44.668646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:6597a5120c7108ca1a2eaa53b54fb49a5c17ed710bf60a70adadd2850adb78e2

Observation d7278b64-b0fc-4e51-978c-ea4eba6e5628 · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:06:44.749051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:7c17bd1f8cbe13031940dc87a06264697b8b899ad8c3cb1c3af3ee1004480a44

Observation 903a5d19-dce7-4342-bc54-48e17815eabd · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:06:44.751337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:9c8bdce1f1a817c4f7b2e5cf1829333082fb7fbf11f73e84851f67deba8fed5f

Observation 8d9c2254-d51a-445a-8868-6dd4d2faf4f9 · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:06:44.753394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:a3f3c131f61e98ca8de2bca455cf5533128cf691741252232b44caefdeb9636b

Observation 40433807-76b3-4146-9ce5-44520110c018 · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:06:44.755461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:2ca19ac8f606836cc1c382672ca14660cf2a2487bf23542007746d17ea31e2f1

Observation d85fef5b-6c70-446f-b26f-e37fa790235e · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:06:44.757723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:ad72868f1148747d5f727241a75fc779a50eeafa924f234c26ae9692d3816e95

Observation 934b4a3a-3e3e-428c-b59d-f2890f2e45f4 · outbound

This paper cites Pixel Recurrent Neural Networks , booktitle = ICML, year =.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Pixel Recurrent Neural Networks , booktitle = ICML, year =

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T20:06:44.759889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:6b1559225d327465dfaa40a5f8e5b3293a61c1934be9c2a2082a949e9377e9f0

Observation 5605f809-485a-4857-a051-298965aca8ea · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:06:44.762144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:a83fb7279c0ff7b0599b3f104100ace605c71331ba0676af888b4b8fff0c362e

Observation 195c0f03-1ee1-4eb9-8824-ea6a7036c4c9 · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:06:44.764302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:608dc8c9d239f772479e6069c39967c1f48e86a9756934277edbf30fe8c2d637

Observation 191d29ba-931e-4fdc-a983-18b5fb3cafab · outbound

This paper cites Kingma and Tim Salimans and Ben Poole and Jonathan Ho , title =.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Kingma and Tim Salimans and Ben Poole and Jonathan Ho , title =

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T20:06:44.766232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:f92cf8b84dca700ac1190b0772a4eb74e10c8abf7cd02d86fbc746f4e82204e2

Observation c6a3b219-d040-4b9d-a721-554d931cb228 · outbound

This paper cites Auto-Encoding Variational Bayes.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Auto-Encoding Variational Bayes

Reference 77

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T20:06:44.658654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:a776668e47acf924dd1d6fea8c3933f8993b92dbf686e1ca573b2a5775b1ede4

Observation 208f4c38-333e-41bf-afa1-727a8dea55a0 · outbound

This paper cites Neural Stochastic Differential Equations: Deep Latent Gaussian Models in the Diffusion Limit.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Neural Stochastic Differential Equations: Deep Latent Gaussian Models in the Diffusion Limit

Reference 78

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:06:44.662299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:22d816281054534220174beef5f7381aa4cce178064258d7c6b649c0af123d84

Observation afcd270b-9be8-4327-98c4-902ba7e7a80d · outbound

This paper cites Denoising Diffusion Probabilistic Models , booktitle = NIPS, year =.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Denoising Diffusion Probabilistic Models , booktitle = NIPS, year =

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T20:06:44.768463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:2c185c8aac4ca656e0b8a624652c80ce6b2984fb13ac20c30fb969dcadcd3e7f

Observation 70d330dc-9dd8-4b44-8d78-c6bd29f1aa36 · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:06:44.770452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:8883e8148481ea79a2bcdad8fc207f6c83a463b8696ce9ddac64cdf90f74cb29

Observation ca7c8ea1-2f4e-4603-b5bc-6db0eceeec75 · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 82

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T20:06:44.692639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:2d4b2bb4687cae59998fda8277d8dc04ed02478735aff721365c8eade20e0ddd

Observation 32f4bd0d-7a33-4a07-8b75-adfacb5357d0 · outbound

This paper cites Phenaki: Variable Length Video Generation From Open Domain Textual Description.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Phenaki: Variable Length Video Generation From Open Domain Textual Description

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T01:43:34.268786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:f5c87c2b60a036bf234d2a5f991d6a9e0fb57043389d93de0fe8b564227b53db

Observation 9078a094-0bff-40e8-90ff-29a36dd503ad · outbound

This paper cites Generating Long Videos of Dynamic Scenes.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Generating Long Videos of Dynamic Scenes

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:06:44.718280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:fc215dbb93a5b584d158bce046784ccac10b07756e754c658f20e65ba20fa2e2

Observation 93422a9c-c635-4b1b-8403-37651c6ec5de · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:06:44.772344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:85ee181279c0b99f48169d8d76a51c201acf93ab9269fd6a003cbd41320f7957

Observation 8b61aac9-a439-40fd-87f5-de5110afafa8 · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 86

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:06:44.774224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:9d2bb8817a12b97422a53cbaaed1db67f7761f7e18711d241e7c21bedcd06db2

Observation 7e285878-c227-463f-bc4d-dcee0577926e · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:06:44.776661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:7700c19a8d533cff81e6ae6ce013799b4d49d676e96fd8d64dcfa3681541584c

Observation 16cb701e-f3df-4202-9383-c2ba931ae565 · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 88

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:06:44.778744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:bbfe743070ca44a3b0e14bf11f4f56ee6f3acb283a86d95ddf0c7b546c812c1e

Observation 8b9fa479-0104-4228-aaa3-44fc7e9a6850 · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:06:44.780557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:d54f1245bd5412b9a96f48f40182d49bf7e8e4faef8806f9836165da7598cf39

Observation f6e37385-1f34-4ad7-8714-806f902e4d9d · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Gaussian Error Linear Units (GELUs)

Reference 90

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T20:06:44.609181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:c16b14d46a21dced275580aa5eaa72226401ce069ec24b1b4a044c2b117dc9e4

Observation daed45de-beb6-430d-b81a-dbda4e5610cf · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 91

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:06:44.782513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:06786ca76b9a7241f6d7a186c9cdbcb267bf606358464892fbc865536d927653

Observation 2d65e868-947e-47d5-b6b2-e821dd911a26 · outbound

This paper cites 2016 , ignorganization=.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation 2016 , ignorganization=

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T20:06:44.784966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:cf55b57d539ff74578c5d0a0f15f69faf221fbcdd813020fc26a11c77304f09d

Observation 1b40324a-7307-484c-98ff-f470b9b48f34 · outbound

This paper cites YouTube-8M: A Large-Scale Video Classification Benchmark.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation YouTube-8M: A Large-Scale Video Classification Benchmark

Reference 93

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:06:44.631852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:f38a34948c4a7f887b7ccea9c58aaf6110ceb8a80bc13e89103d3b276f139e08

Observation c2aa6aa5-0c4f-4451-a527-072689507745 · outbound

This paper cites Visual Prompt Tuning for Generative Transfer Learning.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Visual Prompt Tuning for Generative Transfer Learning

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:06:44.635330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:df98365f33825d506559ca5444ff5d629ce2cc7fd44407d9f8081f36c56e435c

Observation fc2492fe-9c88-404a-a361-afbb789153fc · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 95

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:06:44.787127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:b2c30759b418de6fa72ce009fe1cbd7c5ed69a54ab5cde3813d30e1ae5c2df17

Observation 4fa017c6-bfc8-4ada-b616-45d6bbd50ccf · outbound

This paper cites Language Quantized AutoEncoders: Towards Unsupervised Text-Image Alignment.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Language Quantized AutoEncoders: Towards Unsupervised Text-Image Alignment

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:06:44.655014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:bb06336bdb9d4a2f7e334ada22db33d1098061e730c3a5152a92a1a90c2c5ae5

Observation b0199353-3fb8-4dee-8bb0-bf1ac5200243 · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 97

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:06:44.789169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:c908f8d4b95370367db112d09e3f0cd93fd7f78a8fb7c3895320c1f57099832b

Observation 1806216b-60e4-4f66-8ae9-00cad181b858 · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 98

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:06:44.791217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:e48f3b21ab6ee9419b6f33dbbab545561936af8fc6630fff9b1313e732d92c3f

Observation 625d236b-1196-4ee4-b9ba-a13bf3430083 · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 99

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:06:44.793398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:859215232216ab922d5988bd069a437f7eec8b8ff8624f81d6179d86561fec77

Observation d2cf5f89-2b2f-483d-a146-e3d3aff27b29 · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 100

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:06:44.795228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:2b74d03dac7220c8823b9818635c1ea03314755ca1f0c4a271fdf90145c0938c

Observation 9cc7d358-7aa4-4fbb-804d-10f1f9c1728a · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 101

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:06:44.797106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:5c1d461433b7c3e38afac08b6a3c647069fbd87ebe1110699dd9794726362f0e

Observation 2061d523-f4af-4e49-a001-ce905f7ebd7c · outbound

This paper cites an unresolved cited work.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Unresolved cited work

Reference 102

Resolution
unresolved
raw_fallback, observed 2026-05-13T20:06:44.799011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:44b823c9f46ce3791d94611d458b86a2fc2d288f9d15ade6926d8dc00b8da7c7

Pith citing papers

Observation a62e11c0-6585-4629-b8df-a1cfe76fbb73 · inbound

Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models cites this paper.

Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:06:45.090605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T13:43:11.024069Z digest=sha256:81a364ca501b4aee57c075cfac26e41577ee7c02b866ed78824f58b145f5b2a6

Observation 8269caef-9368-47f3-809b-a6a66bde15fd · inbound

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models cites this paper.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-15T12:26:37.406253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:4c737632243bd7c3386e158a69e4b450f36e829222ef11b2a718f04f946659c6

Observation 15746614-99f9-4b56-8cc4-fe93c6945b72 · inbound

VideoPhy: Evaluating Physical Commonsense for Video Generation cites this paper.

VideoPhy: Evaluating Physical Commonsense for Video Generation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 116

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:34:37.748499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T11:34:37.599691Z digest=sha256:3629ba40e4b9460243469cfaf3bbb345948b5c9067ff94e0d4902eb7d8fdeb00

Observation 34637a02-55ce-429a-a114-635fbfca1a05 · inbound

Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation cites this paper.

Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:06:45.090605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T22:09:16.622717Z digest=sha256:9402c9e7ef7f2ad8a38b281d79cdc697f8011d78dbf015c4d49289a270f21d83

Observation 508b355d-d7e3-4d34-831c-985a8144f9a1 · inbound

CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer cites this paper.

CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:06:45.090605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-10T18:26:22.224924Z digest=sha256:2d24f364d3e1512b7a7552e3de1ee0b8e1cf8d6c326f48f0765592a26f0b1cc6

Observation 7299ffa5-4cf3-4801-8188-88774f622139 · inbound

Show-o: One Single Transformer to Unify Multimodal Understanding and Generation cites this paper.

Show-o: One Single Transformer to Unify Multimodal Understanding and Generation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:06:45.090605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T21:03:33.427939Z digest=sha256:57974ac5c970fcf7bdccdb699cda5f2fbabfbebf1a17755d0ac6299b141e5438

Observation 86105db8-b705-4c88-b1ba-c5b8893b03e3 · inbound

Movie Gen: A Cast of Media Foundation Models cites this paper.

Movie Gen: A Cast of Media Foundation Models Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 82

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:06:45.090605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:f72bd0b06f1dcdf1246df98554d79d2c762bbe6d1761e097614138c78b4ee390

Observation b73c6851-bd68-4f63-9cf0-aff0883c8baa · inbound

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation cites this paper.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:06:45.090605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:5e174e2504a81da56451f8bfefb68326c0ea95172e31848c214329982eae1e52

Observation 59659b5e-f1e0-40d6-9c41-f52e293a3c1b · inbound

HunyuanVideo: A Systematic Framework For Large Video Generative Models cites this paper.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 95

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.594956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:441594aaed55ef3540f05aa7c89220a9010a617793b6a395229689177e86ebc6

Observation 6f2433d4-b724-41b0-9d5d-7983d0457623 · inbound

Open-Sora: Democratizing Efficient Video Production for All cites this paper.

Open-Sora: Democratizing Efficient Video Production for All Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:06:45.090605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T12:01:51.366667Z digest=sha256:8159eca53d522185ab8ab692eab1ca201abe8aa0a1a9fbbd1154ad05396f96b3

Observation 88e2e2a3-6266-4263-a184-e0603270b0ce · inbound

History-Guided Video Diffusion cites this paper.

History-Guided Video Diffusion Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 67

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T12:00:14.778781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T12:00:14.672729Z digest=sha256:c0e3b7d861bda3418ec571b375ff78dbcd823904cd4363b79dd14d67122ae833

Observation f24f2499-65a8-456d-87d6-b5d79340aacd · inbound

DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies cites this paper.

DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-22T23:52:16.744030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T23:51:43.934329Z digest=sha256:4529df0f84103143736872f54a2eff81fc537876342db2a321ea445ee5577540

Observation 9c24c522-6d85-4ffc-adf3-a8c0e566b2d0 · inbound

Long-Context Autoregressive Video Modeling with Next-Frame Prediction cites this paper.

Long-Context Autoregressive Video Modeling with Next-Frame Prediction Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:05:17.319698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:05:17.201790Z digest=sha256:af9fab3668c92751c8db1eb01641b4ea6dd1aa414d4df5930f06b0b26369ef47

Observation 861e91f5-a972-47f1-9302-86ab9090e24e · inbound

Wan: Open and Advanced Large-Scale Video Generative Models cites this paper.

Wan: Open and Advanced Large-Scale Video Generative Models Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-22T23:07:14.230472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T23:05:32.595632Z digest=sha256:d1b06b80dfed4b30f7982718ab4cc410a2d6f96bffd7c47b4fe887d4ad1ed3a1

Observation 47ddc277-904f-43b0-a2d7-987c74eeb0ae · inbound

MMaDA: Multimodal Large Diffusion Language Models cites this paper.

MMaDA: Multimodal Large Diffusion Language Models Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:50:59.838917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:b2fe27e4dbfeefdd0f28f7fe4cafd2e10f7febd3e8e9e21b2fa30012c74944e6

Observation 988a970a-d54e-43ab-a159-202120020f99 · inbound

Seedance 1.0: Exploring the Boundaries of Video Generation Models cites this paper.

Seedance 1.0: Exploring the Boundaries of Video Generation Models Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:06:45.090605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T12:09:56.836351Z digest=sha256:c438840d0a69ef1a3cf487f9b40cc8c6e45517ee9d2621cac4f5c58c40ee1f16

Observation 42a3761b-fb1e-4cd7-9020-57caeba12fd1 · inbound

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos cites this paper.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:45.572085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:45.572085Z digest=sha256:dec2bf1f8faf2c6d0018890750daa9482adeda412c98434dcab4b2948d1598d0

Observation ece70674-f426-4caa-8b23-94f5688ab990 · inbound

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models cites this paper.

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T22:35:50.958123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:35:50.958123Z digest=sha256:612d8621a0bf5c82493ade9ee8ff1b5eb4f20e9b528bee4e3efa9fa72ea8ffc8

Observation e7e379fc-a4a9-4efe-9405-b7485f2192dd · inbound

CharacterShot: Controllable and Consistent 4D Character Animation cites this paper.

CharacterShot: Controllable and Consistent 4D Character Animation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-05T22:13:13.818684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:13:13.818684Z digest=sha256:a123cdcfbef140de38d5bfde638be3a599bc50665d8271acbf10f31f8cc7142e

Observation 085c8fec-7879-4d83-9ba4-3c5a4fa9332b · inbound

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model cites this paper.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:00.299651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:00.299651Z digest=sha256:2778e0bcd4931ccde2ae39248ed5702efbd2e89490027374c2045c7a3de2c6e2

Observation 5cd7c1a4-8aa2-4023-9c60-36f65d93ae02 · inbound

Time delay as the origin of oscillations in anodic Si electrodissolution cites this paper.

Time delay as the origin of oscillations in anodic Si electrodissolution Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T20:50:29.163351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:50:29.163351Z digest=sha256:9632131b565bd3487ede14fe8cd146997aad48accb2dfeb5f420a3eaa4b90f8a

Observation 0ab7a641-5b0a-42ee-8e98-a496394c79d8 · inbound

HumanGenesis: Agent-Based Geometric and Generative Modeling for Synthetic Human Dynamics cites this paper.

HumanGenesis: Agent-Based Geometric and Generative Modeling for Synthetic Human Dynamics Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T20:49:53.496674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:49:53.496674Z digest=sha256:0f64fea3f0335267c2cba7b7c430ca77e2cd6cdf2552c5460008c7d8174c57cc

Observation c0098db9-3dc5-4551-b07b-e52f6010aaf4 · inbound

Representing Speech Through Autoregressive Prediction of Cochlear Tokens cites this paper.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.673271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.673271Z digest=sha256:c9b71a83cd397e0c1bd45e26d999516d7403ff26da828afa94705dc41de73db7

Observation aaea03a9-88b0-4faa-9aa9-509c2580afd1 · inbound

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation cites this paper.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T17:34:01.775720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:34:01.775720Z digest=sha256:062565312872dad34d437a6dceb9d79a81174c891845416d18cd0038a4efaafe

Observation 8521cf50-b046-4cac-a5d8-a67202885bac · inbound

Discrete Noise Inversion for Next-scale Autoregressive Text-based Image Editing cites this paper.

Discrete Noise Inversion for Next-scale Autoregressive Text-based Image Editing Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T12:07:36.996492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:07:36.996492Z digest=sha256:f0632fe1be8baf5962bc419903d0c85322197083df44fee358bcbf8425a60130

Observation bb26dbb5-a894-490c-a3ec-a7245e1dde27 · inbound

Transition Models: Rethinking the Generative Learning Objective cites this paper.

Transition Models: Rethinking the Generative Learning Objective Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-05T10:19:54.524189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:19:54.524189Z digest=sha256:6fca828d5d9b10fee00fd4835cf70e11e7e631f1c50788dfc7fff27735810029

Observation 84bb5c98-9958-4157-bbcd-1e90c5fc36c1 · inbound

Missing Fine Details in Images: Last Seen in High Frequencies cites this paper.

Missing Fine Details in Images: Last Seen in High Frequencies Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T05:27:34.352559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:27:34.352559Z digest=sha256:16c281e91677d22d303a0baecf64e133306533e5a4b24306aebc96eb1cc0cd21

Observation fe512b93-21aa-4a66-ac9c-d1b067709c97 · inbound

WindFM: An Open-Source Foundation Model for Zero-Shot Wind Power Forecasting cites this paper.

WindFM: An Open-Source Foundation Model for Zero-Shot Wind Power Forecasting Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T23:57:19.690271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:57:19.690271Z digest=sha256:8aec67289621747811615e60d6f4d86c64b608bda7d2b057a9db3213e1bb0714

Observation 76786ac9-fcd0-4a3a-ab66-6a0575b633b9 · inbound

Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking cites this paper.

Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T16:43:53.195927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T16:43:53.195927Z digest=sha256:198eb2e55180d686c06c373ccc671a5706f30e248da01d2f06a9a25065424289

Observation 3adcad21-1e27-4228-9f05-d45dfd5735e5 · inbound

FluentAvatar: Flicker-Free Talking-Head Animation via Phoneme-Guided Autoregressive Modeling cites this paper.

FluentAvatar: Flicker-Free Talking-Head Animation via Phoneme-Guided Autoregressive Modeling Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-18T16:42:43.875344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T16:42:25.803856Z digest=sha256:8c8a9c7a76d14820833886d920373f5669de20b42f953170d245214075f3c47a

Observation f015d29f-1849-4513-8d6f-98fb47f7c4e6 · inbound

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs cites this paper.

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:01:24.294582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T12:57:04.450462Z digest=sha256:a69b50d5b33f504b06c3b4e5b14dff68d4af76285abecd0c447866b908fb4bce

Observation 89c5cd24-a3fe-4422-935f-aaae85fb6685 · inbound

Control-Augmented Autoregressive Diffusion for Data Assimilation cites this paper.

Control-Augmented Autoregressive Diffusion for Data Assimilation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-18T09:06:09.268139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T09:02:38.416697Z digest=sha256:e69a3e8c26153c69084ba3382c6a080ad705eb500529fe7112ae58f2a78b45d3

Observation ed0b44ed-8e16-4de4-b403-31258fdda2e3 · inbound

IAR2: Improving Autoregressive Visual Generation with Semantic-Detail Associated Token Prediction cites this paper.

IAR2: Improving Autoregressive Visual Generation with Semantic-Detail Associated Token Prediction Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T11:08:02.602491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:08:02.602491Z digest=sha256:b7c660e074c22b91c61ba112f05b95bec37ea32f4063cd3eb7d13b83f74ab040

Observation a8de84c0-39f0-4236-bdb3-511cd157aaad · inbound

VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models cites this paper.

VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:22:23.910210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T05:22:05.125849Z digest=sha256:512b37f32d4040b81c7a744ac29d79e27238377dae7e2b7b7c64934e13e8028c

Observation fdd71819-bcf5-4be2-842d-7a77301b0a35 · inbound

Distribution Matching Variational AutoEncoder cites this paper.

Distribution Matching Variational AutoEncoder Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T17:55:08.235444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:55:08.235444Z digest=sha256:1a31f5e1ab3c0c0695d836d0d7a0f3088e4ddce9cf2c7d5a82df3dafe527152b

Observation 2abe063b-81db-4fb9-b09c-345f987c8e35 · inbound

HD-Prot: A Protein Language Model for Joint Sequence-Structure Modeling with Continuous Structure Tokens cites this paper.

HD-Prot: A Protein Language Model for Joint Sequence-Structure Modeling with Continuous Structure Tokens Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T16:07:37.936876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:07:37.936876Z digest=sha256:56aa780bf0fd56fe8c463a5a0e94973ee0fb6e21e2668c10602359e892471d6d

Observation ed5a70c3-3676-43e2-a4be-1426cfbdd2fd · inbound

MixFlow Training: Alleviating Exposure Bias with Slowed Interpolation Mixture cites this paper.

MixFlow Training: Alleviating Exposure Bias with Slowed Interpolation Mixture Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T14:53:03.206913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:53:03.206913Z digest=sha256:f7e9ba0f89c204d58a89e99b983aef7a05d4c52b0c627b231fa0c9283341f5ff

Observation 888fe29a-1065-4260-be74-1b29cd10fe46 · inbound

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training cites this paper.

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-03T13:29:53.536777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:29:53.536777Z digest=sha256:a9314d4de43da3c6255281519ea8acf080d5b7d076194d6799cf29eb3d5cfeaa

Observation 6a57b081-e11c-4424-9b81-839f312ebd74 · inbound

LGQ: Learnable Geometric Quantization for Image Tokenization cites this paper.

LGQ: Learnable Geometric Quantization for Image Tokenization Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T22:43:28.446943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:43:28.446943Z digest=sha256:cadc66748967cc740536a2a0b8436d32b4115e87f99172f363f864a14af6a01c

Observation b336320e-fa4a-48fd-83ef-7f48cfaae028 · inbound

Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation cites this paper.

Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-15T21:30:20.846856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T21:22:41.935691Z digest=sha256:05dccebcf8aa37d274015182c7f17938379f61616a7caa292dc03c073f42be58

Observation dca2d457-9a40-456d-9ff3-7f3afbdcf5fa · inbound

ChopGrad: Pixel-Wise Losses for Latent Video Diffusion via Truncated Backpropagation cites this paper.

ChopGrad: Pixel-Wise Losses for Latent Video Diffusion via Truncated Backpropagation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 74

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T09:45:23.318920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T09:42:26.074609Z digest=sha256:15b25a7fadd1dbb62a81e3301f157b2b9bf9ecb1b996f92d5c37039b154ffdec

Observation edb49e94-8b98-44d5-aae1-6ffe96c59bbb · inbound

GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation cites this paper.

GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-13T17:23:44.758186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T17:23:44.758186Z digest=sha256:a3b35fad069c38690734f8ef377c22849a7aeedcf82c16afe25faa506e5a72e7

Observation 0f4bc646-fcdf-40ef-a45e-09190cda0827 · inbound

ExploreVLA: Dense World Modeling and Exploration for End-to-End Autonomous Driving cites this paper.

ExploreVLA: Dense World Modeling and Exploration for End-to-End Autonomous Driving Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 49

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T20:53:15.590642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T20:52:25.139770Z digest=sha256:2211dddb05afa17d69e513ffdad7113942be0dca661cdcc4b0e3e94af4c7f2af

Observation 1686961d-e361-43b4-a16d-4daa5e7666ad · inbound

ELT: Elastic Looped Transformers for Visual Generation cites this paper.

ELT: Elastic Looped Transformers for Visual Generation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:06:45.090605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:19:22.543462Z digest=sha256:dc77c93502514997c01f05bd60361272a3969b795891ea543b661394261186f6

Observation 2138cc31-9281-498c-bc4d-ae00c8605193 · inbound

ELT: Elastic Looped Transformers for Visual Generation cites this paper.

ELT: Elastic Looped Transformers for Visual Generation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-02T16:35:03.423853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:35:03.423853Z digest=sha256:98e927fb8d97530fa11b55dd446169bdb2e14175608dc0c6012b8edc13fbf81a

Observation a4da9180-7516-417e-8d90-f7502fd141d2 · inbound

Prompt-Guided Image Editing with Masked Logit Nudging in Visual Autoregressive Models cites this paper.

Prompt-Guided Image Editing with Masked Logit Nudging in Visual Autoregressive Models Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:06:45.090605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T11:48:00.163951Z digest=sha256:4a01ba47c407cee820cdd20b53823adba07aa2421087a037b3d89cdd26d257d3

Observation 77bafbe1-84b6-488a-acff-59a4c5739655 · inbound

Latent-Compressed Variational Autoencoder for Video Diffusion Models cites this paper.

Latent-Compressed Variational Autoencoder for Video Diffusion Models Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:06:45.090605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:23:19.583713Z digest=sha256:07d7f6ecfdf1a8013b24eb8303fb5a3e255d2f438e680240498832e5d164d4ae

Observation 38bb0a5a-7dd2-4464-b8ce-4ba21c79d784 · inbound

dWorldEval: Scalable Robotic Policy Evaluation via Discrete Diffusion World Model cites this paper.

dWorldEval: Scalable Robotic Policy Evaluation via Discrete Diffusion World Model Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:06:45.090605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T11:45:18.081248Z digest=sha256:e416dff8f608b99575e07c88389839882f53b2b3b1916e0be831d45445c290ef

Observation df0314d0-fa8c-49ba-ab81-30ee5c88deba · inbound

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations cites this paper.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:06:45.090605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:ddbfb8a34db4bc0d88640e09043adc767898042b5c747dba9e5a9ecd68f2b679

Observation b6aa4ab8-6696-4c74-b20a-5c45614a2fa3 · inbound

End-to-End Autoregressive Image Generation with 1D Semantic Tokenizer cites this paper.

End-to-End Autoregressive Image Generation with 1D Semantic Tokenizer Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:06:45.090605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T19:41:03.302303Z digest=sha256:5087c200bbcf3abd822d1496e8e10abad09e928ed441e13d92af2a361f901dbe

Observation 935a52f3-4e73-4203-a007-14b27697ad20 · inbound

Co-Generative De Novo Functional Protein Design cites this paper.

Co-Generative De Novo Functional Protein Design Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:06:45.090605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-09T15:18:49.411291Z digest=sha256:f474bc0cd681bc7bdc39dfd00b30af5b906dbc07da143147462f2a6efdf363c2

Observation 1edb73f1-2eec-447e-914b-f89d4cf37902 · inbound

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models cites this paper.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:06:45.090605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:433030f09796e6bd09b0061b6fc5420888f6abfc03691c39101e41f69c355533

Observation 7e4b04a8-56ee-4e5a-a1c1-edbb41265288 · inbound

MUSE: Resolving Manifold Misalignment in Visual Tokenization via Topological Orthogonality cites this paper.

MUSE: Resolving Manifold Misalignment in Visual Tokenization via Topological Orthogonality Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 159

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:06:45.090605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-08T15:04:41.518195Z digest=sha256:6d26c4f97e4fcdd893d0e6f84d97db2883bf56180b97111c9980eb13a6918a96

Observation 28cbb921-8f5d-4fd2-b5f5-c73a38ecf191 · inbound

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization cites this paper.

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:06:45.090605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:25:52.847432Z digest=sha256:c433973fad1c58d2782bb48d7f0507bf7221c445f96d1094a26a870e08d486bb

Observation 7b51f058-dbde-4dc6-ace1-67255cdc55fe · inbound

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization cites this paper.

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:15:07.766612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T23:14:32.494076Z digest=sha256:511253651e6fb2e3a2fc675408247f1b964b8ec65f22fa070f096f8ce8a63f95

Observation 2583d372-b0fc-4cea-99c1-279d73e86670 · inbound

CASCADE: Context-Aware Relaxation for Speculative Image Decoding cites this paper.

CASCADE: Context-Aware Relaxation for Speculative Image Decoding Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:06:45.090605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T02:08:27.374066Z digest=sha256:c95ee5f0b73c6e7a39faef4b8474f8152b9064e1307209a92ebb2bb48ca7fb3f

Observation 518836b1-b8e2-495b-834e-56e6e130c9aa · inbound

Yeti: A compact protein structure tokenizer for reconstruction and multi-modal generation cites this paper.

Yeti: A compact protein structure tokenizer for reconstruction and multi-modal generation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:06:45.090605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T04:01:05.492969Z digest=sha256:13560c087aa54c3603ce1be6ed5fa102113bd7cc57d16eb41401a69adc30afc5

Observation 514b5810-22c4-47f6-8eeb-59abaf9cc22f · inbound

InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation cites this paper.

InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:13:30.598211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T02:10:18.004724Z digest=sha256:6772c380cbca09662300acc761a90091b41c95b57572a987f25c667fc830b1e0

Observation 2c41be14-9832-4ba3-bf6e-b31ea9c5b618 · inbound

Multi-scale Coarse-to-fine Modeling for Test-time Human Motion Control cites this paper.

Multi-scale Coarse-to-fine Modeling for Test-time Human Motion Control Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 94

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T21:55:06.078194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:48:55.350493Z digest=sha256:ce0d6d9e52fbea9edf798e3c7883e633ebda39793ee09a72e40ed08ee61cb048

Observation 57a7a722-9287-4357-b770-63e6dcf9360f · inbound

Mutual Enhancement Between Global Tokens and Patch Tokens: From Theory to Practice cites this paper.

Mutual Enhancement Between Global Tokens and Patch Tokens: From Theory to Practice Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T22:43:51.089664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T22:41:44.510546Z digest=sha256:eb5bcb6d7274621bc2e0527c0852bc16f507608f0f8663b3dd7b80b956fd9740

Observation ffb30c11-ffcc-44f4-be5c-86670fed6589 · inbound

SRC-Flow: Compact Semantic Representations Enable Normalizing Flows for Image Generation cites this paper.

SRC-Flow: Compact Semantic Representations Enable Normalizing Flows for Image Generation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:03:15.313939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T11:59:54.139888Z digest=sha256:c653023b0feccca6a11a4607f0bc13b894804c68549394dc6a2a18a94c518ec1

Observation 2a2a847e-4a26-4a51-88df-75340501e293 · inbound

SRC-Flow: Compact Semantic Representations Enable Normalizing Flows for Image Generation cites this paper.

SRC-Flow: Compact Semantic Representations Enable Normalizing Flows for Image Generation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:45:00.693501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:39:40.667006Z digest=sha256:4dbc1669dccd4113732f34993e182cb451cc7da6fa5d0cef080fc5ddf7ac5304

Observation f45aacf6-8785-4197-a7b1-8b5502e1f04d · inbound

SRC-Flow: Compact Semantic Representations Enable Normalizing Flows for Image Generation cites this paper.

SRC-Flow: Compact Semantic Representations Enable Normalizing Flows for Image Generation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T13:49:21.531757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:49:21.531757Z digest=sha256:e6901de8a39dbc844a014db101c254de7956bab668d18470e29bc6adc2e7054d

Observation d56a88f3-8438-479c-867a-d87b6a4025ba · inbound

A Dialogue between Causal and Traditional Representation Learning: Toward Mutual Benefits in a Unified Formulation cites this paper.

A Dialogue between Causal and Traditional Representation Learning: Toward Mutual Benefits in a Unified Formulation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T05:43:58.744717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T05:42:48.667216Z digest=sha256:0ed12497aa56af3ea22734efc7e3570bc71be5b3c845ba1e8284534d895f5bfc

Observation d767dc13-57fe-4618-9e43-318bcf8fd947 · inbound

Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models cites this paper.

Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 77

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:34:46.664366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T09:34:14.596976Z digest=sha256:da079e5b5e45a53930d7ff2ddda12fa0540b883acb30a5014464a2386f21288c

Observation 83b96a24-3a42-46d2-8315-bc5cf9be4832 · inbound

PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion cites this paper.

PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-25T04:20:19.314786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T04:18:45.403718Z digest=sha256:2e2b3ecdc6ca8291cb138b2423283698c2947b1198e2d58840f73b0984969973

Observation 601e44bc-c052-4d4b-89f9-e011dd6a7e21 · inbound

VPG: Visual Prefix Guidance for Autoregressive Image and Video Generation cites this paper.

VPG: Visual Prefix Guidance for Autoregressive Image and Video Generation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:03:14.019333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T08:00:16.005187Z digest=sha256:a5cc701bf50d1605ba7bd82c6180ede890d5dd33a2848e8a9c3aabde7c7bda9e

Observation 93cf740d-f35d-42ec-98ba-da75e6a91569 · inbound

ChannelTok: Efficient Flexible-Length Vision Tokenization cites this paper.

ChannelTok: Efficient Flexible-Length Vision Tokenization Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-02T07:06:44.337086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T07:09:25.049534Z digest=sha256:a892fdde1f6a2ab6b2492ac3acaa3a5693fc85dba7998b7909e127524dbbf497

Observation ab2c642d-3fee-4b1b-9fd2-99fc6af91f94 · inbound

MaCo-GAN: Manifold-Contrastive Adversarial Learning for Single Image Super-Resolution cites this paper.

MaCo-GAN: Manifold-Contrastive Adversarial Learning for Single Image Super-Resolution Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 53

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T07:36:44.752572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T06:53:27.042821Z digest=sha256:4312b88e5c654545d05335ec4be4a63dbb2cd038277b1ba55ac73fa8992f7df3

Observation 63658c79-b988-405c-9f83-afbf15a113ce · inbound

Balancing Image Compression and Generation with Bootstrapped Tokenization cites this paper.

Balancing Image Compression and Generation with Bootstrapped Tokenization Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-02T11:46:55.476386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T03:07:33.054518Z digest=sha256:f34c9b701709a079888c9f115a787feb7acb6e9ea50363be7e15ab85604991a3

Observation 78906622-c9d2-4555-9241-f3efd9a81819 · inbound

AdaTok: Self-Budgeting Image Tokenization with Quality-Preserving Dynamic Tokens cites this paper.

AdaTok: Self-Budgeting Image Tokenization with Quality-Preserving Dynamic Tokens Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T16:27:09.097579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T22:43:09.489524Z digest=sha256:1030c7b62cf77fa11330a73c4f33b1eaa6c0bcc171ad472a4e0d868132b01ea8

Observation 55e97e9a-910f-472d-9925-78803d6b634e · inbound

Time Series as Language: A Universal Tokenizer for General-Purpose Time Series Foundation Models cites this paper.

Time Series as Language: A Universal Tokenizer for General-Purpose Time Series Foundation Models Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-06-28T18:02:27.050012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T17:57:07.869965Z digest=sha256:e2d47f0da80b3c5d10bc49035222a93c19258d37b92e1ba9ae65ab9fc71486f6

Observation a4981707-dbb8-4620-8c92-5dfff0059716 · inbound

Self-Guidance: Enhancing Neural Codecs via Decoder Manifold Alignment cites this paper.

Self-Guidance: Enhancing Neural Codecs via Decoder Manifold Alignment Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 138

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T16:08:37.471517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T06:05:26.735340Z digest=sha256:fbdf08c562682a13ad8c34300c3e54061a5f6e784b09958784afe35ad908f68d

Observation e9d36f68-0d8f-4812-9bec-57ab91a2194f · inbound

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers cites this paper.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 46

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T14:38:28.909903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:db3498b105303b1e3c857dbbbcd931a97a9f41c703fcae82803f7998662ed903

Observation 23f9c086-50ec-48a0-8512-9b65556db51e · inbound

ARP: Enhancing Quantized Skill Abstractions via Visual Alignment and Iterative Refinement for Robotic Manipulation cites this paper.

ARP: Enhancing Quantized Skill Abstractions via Visual Alignment and Iterative Refinement for Robotic Manipulation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-04T09:19:43.447089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T10:14:42.111428Z digest=sha256:19fbfa287a07804807c4bb0aca7fe1857db84b91dd5ce989f625b1a0ede2e6e0

Observation 2ed7dea8-8e8b-4a1a-9d63-584d3d265d23 · inbound

Lighting-Consistent Object Transfer Across Radiance Fields cites this paper.

Lighting-Consistent Object Transfer Across Radiance Fields Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 265

Resolution
verified exact
local_arxiv, observed 2026-06-26T09:49:18.026808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T09:46:46.216810Z digest=sha256:b01c3839ae3622fa8696e6181f7f2dc5491778368bd935bfe89176b64167a2ea

Observation b8cdd610-6234-4765-81cb-daa427a2d524 · inbound

DiffusionBench: On Holistic Evaluation of Diffusion Transformers cites this paper.

DiffusionBench: On Holistic Evaluation of Diffusion Transformers Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T16:59:58.103164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T00:06:11.951205Z digest=sha256:39f2211a7483c2c6591437b2aa9603de97b05d75d2bf979c9157f651207761b6

Observation 0e2d65ef-42fe-49e3-92c2-f5832fd50a3d · inbound

Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis cites this paper.

Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-06-30T08:14:26.629603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T06:07:55.600338Z digest=sha256:6a4ac974c7310636535373998adae999779550c545067696ce6a6f8ebcc2dd96

Observation 855dce2e-4ccf-4494-8121-13b6e690c3e5 · inbound

Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis cites this paper.

Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T09:39:34.321400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:39:34.321400Z digest=sha256:82bc31804e8351eb1180c0713637502453a136f326bd096c1844bc7b7ae367e1

Observation 3c219fdf-a4a5-4e4e-9497-efd3339d6dd2 · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 154

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T11:55:42.292952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:c35a292fee34d0c2fb2cf82815438141b994da7db91d706f6b0fef1acbaccc4f

Observation 8ec98d85-a776-4087-9667-32ecea3de434 · inbound

GEAR: Guided End-to-End AutoRegression for Image Synthesis cites this paper.

GEAR: Guided End-to-End AutoRegression for Image Synthesis Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-07-01T10:35:42.656650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T05:19:21.647714Z digest=sha256:63cb5f5e318a28f332f043a53eb9edcbbe1d4f90ec52a50103eab4d36c053c56

Observation 7c0c8cb3-a989-42e4-9360-8062e2ceca53 · inbound

Arachne: Orchestrating Cascades for Efficient Text-to-Video Model Training cites this paper.

Arachne: Orchestrating Cascades for Efficient Text-to-Video Model Training Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 54

Resolution
malformed identifier
local_arxiv, observed 2026-07-03T06:37:42.063124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T06:27:48.935774Z digest=sha256:c074b39457e8e4f7c9c9a9dbc605b9651d6455cf14d8d02b335c11f53c3caabc

Observation ed704826-f99b-473b-a598-18623d53d049 · inbound

MentalThink: Shaping Thoughts in Mental SVG World cites this paper.

MentalThink: Shaping Thoughts in Mental SVG World Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-12T01:50:59.184754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:50:59.184754Z digest=sha256:c8626debe2c5661ff1e0dbc17f637ae5711dc6f7412a9374908d13b3f739462b

Observation f26a8f91-2e8d-4c2b-b5b0-b2da464e77bd · inbound

From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence cites this paper.

From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-14T03:51:24.547781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T03:51:24.547781Z digest=sha256:af196630e16e1b7acff91e63ec24ee50e19825dcd512f287455ac9624d2dda39

Observation 168a649a-0a32-4b07-a546-4b80edc34ebd · inbound

Three-Body Scattering for Generative Modeling cites this paper.

Three-Body Scattering for Generative Modeling Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-01T15:49:34.744699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:49:34.744699Z digest=sha256:4904cd8215a4d39597326394f462f608d179ba29000a004e7bf9b8173ee8384f

Observation 9870d400-c33e-4aa4-9102-eee85c43cac1 · inbound

dRAE: Representation Autoencoder with Hyper-Spherical Codes cites this paper.

dRAE: Representation Autoencoder with Hyper-Spherical Codes Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-01T05:45:53.082144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T05:45:53.082144Z digest=sha256:3099639a6683f1c0af2db57c864699ea9614a3bfa9ad54a004e03a3957b2f9f2

Observation 93c76633-4f2c-4a27-a6f4-4e89acadfbcf · inbound

PhiZero: A World Model Built Around Physical Language cites this paper.

PhiZero: A World Model Built Around Physical Language Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-07-31T01:50:30.716875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T01:50:30.716875Z digest=sha256:ae7e4d610387abc0dd4beb1c67b44fa688b9f2025368687e83735877591104b5

Observation 27df76a9-3985-4e09-8eea-2d6d47d6a9e6 · inbound

WaiT for the Signal: Simple Frequency-Aware Flow-Matching cites this paper.

WaiT for the Signal: Simple Frequency-Aware Flow-Matching Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T00:34:35.560384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:34:35.560384Z digest=sha256:36647bd47768bfc662b0488fae799564b7fa1251b7f4b6f1a890fd24d2c5b2e4

Observation e53f494d-78b6-4fdf-bbac-a1574ff4f02a · inbound

Beyond Token-Level Cross-Entropy: Fr\'echet Distributional Post-Training for Autoregressive Image Generation cites this paper.

Beyond Token-Level Cross-Entropy: Fr\'echet Distributional Post-Training for Autoregressive Image Generation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T00:44:16.956126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:44:16.956126Z digest=sha256:bd3e2eb0a5f0fcac87661f9c77b74f80cc2b74d7525b3f1273f343e897750df2