Pith. sign in

Paper Citation Record · LEDGER

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation

As of 23 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 2 inbound Pith citation observations for arXiv:2505.13439.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.13439 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:16:36.724968Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:27:00.510598Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T02:13:30.560857Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation db410918-91ed-498b-880d-92c47a057137 · outbound

This paper cites Ntire 2017 challenge on single image super-resolution: Dataset and study.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation Ntire 2017 challenge on single image super-resolution: Dataset and study

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:38.283097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:16:36.300834Z digest=sha256:b5e2b2f782a1ff8a093dddc20ad98031c1f571424e558d901dc4ef14ad940cdd

Observation 8904b23b-cdc1-4eb4-896f-f34b4f875235 · outbound

This paper cites Clifton, Yuxiong He, Dacheng Tao, and Shuaiwen Leon Song.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation Clifton, Yuxiong He, Dacheng Tao, and Shuaiwen Leon Song

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:38.261068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:16:36.312359Z digest=sha256:b953feccaec080730d3686b4ef6d7f807b5bdf2df9e76336e3f0c5a4655c7a6c

Observation 26325fd5-f7da-4ff6-990e-c4bb15145517 · outbound

This paper cites an unresolved cited work.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:16:38.232363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:16:36.319246Z digest=sha256:b13f015a0fdd0de73aa3c82954835a0d86c0e1018edc4d8c73a154ade90bfe35

Observation fb59921e-e5ca-4b5d-8f23-a7fac145d4f2 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:36.342757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:36.342757Z digest=sha256:513d7187d10769f1350b5ce3bd7f963bed704ea7ba3cf6fc9df3b1d2ea45a1b8

Observation a5b652ff-d965-4adf-988f-0bda4ea9a5a2 · outbound

This paper cites ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:36.354700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:36.354700Z digest=sha256:3c088169bcdda1be7847e902035399b1cb8a22471cc805d5c1fdbc8cce3ed894

Observation b4eb62af-8856-46b6-93e7-991a6bca1028 · outbound

This paper cites Diffusion models in vision: A survey.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation Diffusion models in vision: A survey

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:38.214092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:16:36.364010Z digest=sha256:24cab0b010b91b467b728710131a063a54c55e299051d7b8ba8d1041a04c1c28

Observation 538b30ab-20d2-416e-99bd-ef3bb6787e56 · outbound

This paper cites DeepSeek-V3 Technical Report.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation DeepSeek-V3 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:36.370349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:36.370349Z digest=sha256:4d8049c2ef7ff395e8daeac86c940482a02aa375b847ce5a6aa9081530a115f7

Observation 70d69441-ae04-4793-988e-ca469d293cab · outbound

This paper cites Taming transformers for high-resolution image synthesis.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation Taming transformers for high-resolution image synthesis

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:38.184099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:16:36.377341Z digest=sha256:78309eacc9847342ba7c66c55347be7c7605f92086fd87d5036b4505df43d11e

Observation f45c5ba9-0e6c-4fb9-8f9e-859ad9aad8af · outbound

This paper cites Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:36.384466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:36.384466Z digest=sha256:c8a9aef30dbc79bec396f3323e665d41c69cf4cd37839a05b3539950129ca0e4

Observation 1b0cb649-3dc8-4e6e-87d6-6df936d3e7e7 · outbound

This paper cites Unified Autoregressive Visual Generation and Understanding with Continuous Tokens.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation Unified Autoregressive Visual Generation and Understanding with Continuous Tokens

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:36.394421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:36.394421Z digest=sha256:bf3c71e38d7e7338411aadde2c9af6ef4a243b206a25173d96ce5c0af085cc11

Observation e7820c31-5e22-45cb-a92c-e411abf5de04 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:36.401027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:36.401027Z digest=sha256:b566c18ff41395722c59abe0d080519fff118899b4c547ad3540ef94488e4ca3

Observation affc3f69-1212-4306-b807-b4b73f4a4a68 · outbound

This paper cites Geneval: An object-focused frame- work for evaluating text-to-image alignment.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation Geneval: An object-focused frame- work for evaluating text-to-image alignment

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:38.163394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:16:36.409214Z digest=sha256:41a98913d84eda3c86d55674bd993fbf58e9d2b3563eff788c47901a65abb6b8

Observation 702514c5-fc94-4f91-a9ab-61fe4b3adfe7 · outbound

This paper cites Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:36.417004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:36.417004Z digest=sha256:526f2d768212456d651402e7013c97da31b529bca88964ab4be7bfe6173fe91b

Observation d9cbcffc-f756-45e3-ac1b-ea60dcd642a6 · outbound

This paper cites Denoising diffusion probabilistic models.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation Denoising diffusion probabilistic models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:38.127803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:16:36.422755Z digest=sha256:64cd3ddd8d06f244b478f3091cf8d234388e38ffadeaa624448e0f9606e199ef

Observation e8d033b2-051f-4aca-9fb9-ed536b844fdd · outbound

This paper cites T2i-compbench: A compre- hensive benchmark for open-world compositional text-to-image generation.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation T2i-compbench: A compre- hensive benchmark for open-world compositional text-to-image generation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:38.091876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:16:36.428618Z digest=sha256:79ea75b9d23753dabeff9166067805c31f9a9334166b56dfc33ece0c24dd2600

Observation e6337c26-64a4-4e4d-8f0a-52d9de4633a1 · outbound

This paper cites T2i- compbench++: An enhanced and comprehensive benchmark for compositional text-to-image generation.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation T2i- compbench++: An enhanced and comprehensive benchmark for compositional text-to-image generation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:38.070585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:16:36.433188Z digest=sha256:0c61e3d5eee08c272df880bc88722b40278ca21418e58aa44643361b50612527

Observation b6ea2a2f-7d13-4495-917d-5352094c9e3a · outbound

This paper cites Mixtral of Experts.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation Mixtral of Experts

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:36.437691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:36.437691Z digest=sha256:cb01dcc841cef24db8bd3a98e06b85f9c912dd3c8f116976a3eef90ba5b114ae

Observation 617f1c81-8d5e-422f-81ca-1804817f5d68 · outbound

This paper cites Gemma 3 Technical Report.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation Gemma 3 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:36.443765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:36.443765Z digest=sha256:974553d6f2fdfbc6087866df810afcf85dcb3020d35b0bf3a7359fa92ec0f6be

Observation ce6db9e8-273a-42be-8483-3aa33e23bce9 · outbound

This paper cites an unresolved cited work.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:36.452476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:36.452476Z digest=sha256:b0340c0cfc6720559479a5efc6f6147afb177ed6ed57e0efd1d81ed7ebfe5743

Observation 3dca018f-35e6-4697-80dd-f60765f49149 · outbound

This paper cites Have we unified image generation and understanding yet? An empirical study of GPT-4o's image generation ability.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation Have we unified image generation and understanding yet? An empirical study of GPT-4o's image generation ability

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:36.458745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:36.458745Z digest=sha256:8c9942c3113b104a2f3845a86a69538e7c1b44f62ea8117d29db0698a2051e01

Observation a1b5da0e-e809-4a95-810e-2af0c068cc6b · outbound

This paper cites Autoregressive image generation without vector quantization.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation Autoregressive image generation without vector quantization

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:38.038928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:16:36.466766Z digest=sha256:268c4694145eb57d8da9c7bc4d1f050e6535e4b29d88a6e619286aa5d9e9ddda

Observation 135a407e-4e75-4019-902a-a41f79aba2e2 · outbound

This paper cites DMin: Scalable Training Data Influence Estimation for Diffusion Models.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation DMin: Scalable Training Data Influence Estimation for Diffusion Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:36.473362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:36.473362Z digest=sha256:13bd16d7b7cf74f91182d0f40a63b80dbcb68fc2904aa82c750c582015d64524

Observation ea3a709d-62b7-4745-b640-8c06eef6a010 · outbound

This paper cites Token-wise influential training data retrieval for large language models.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation Token-wise influential training data retrieval for large language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:38.019293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:16:36.480095Z digest=sha256:5a23484c328a532ce04298c4ebeeb7c81da698bc1a2bc814121c9b8443b487b7

Observation f1ab93dd-d49d-4ce0-89b6-a50182fb56cf · outbound

This paper cites UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in Large Language Models.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:36.487502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:36.487502Z digest=sha256:ecc3bf0accf52c3310d776fac46c7240cda8bafb951cba67023bd9abbded2c3d

Observation 05103f85-38f4-4ae8-ac6c-8141a6e11884 · outbound

This paper cites Unified-io 2: Scaling autoregressive multimodal models with vision, language, audio, and action.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation Unified-io 2: Scaling autoregressive multimodal models with vision, language, audio, and action

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:37.994942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:16:36.493103Z digest=sha256:734207c63448af3348d3b444fd1fda4626bd8d1c292049491a89c3877e8ff5b6

Observation 5d2824b9-b490-4ea5-9cd5-26d9bd75015a · outbound

This paper cites Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:36.500770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:36.500770Z digest=sha256:c10ddbeb1c0c7fbc0764f53b98c3dde5bc0ba99a61713be541a354fce0130ead

Observation f83de64b-06dd-478d-bd12-388f9ba8b9ae · outbound

This paper cites Token-Shuffle: Towards High-Resolution Image Generation with Autoregressive Models.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation Token-Shuffle: Towards High-Resolution Image Generation with Autoregressive Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:36.507549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:36.507549Z digest=sha256:3b68fffc0c36133d4920ff325ba7d303b695fe501a83b3a58599354fda535c16

Observation d13a32ad-66c6-4e74-a2c8-081202648097 · outbound

This paper cites Blaschko, Guohao Dai, Huazhong Yang, and Yu Wang.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation Blaschko, Guohao Dai, Huazhong Yang, and Yu Wang

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:37.971465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:16:36.513880Z digest=sha256:bd0c4b82849768048009fe0940b77d1ce4e4596d98e7d72a0f80da752b659324

Observation b30d8c8b-1591-4933-95b9-e2207f1105d7 · outbound

This paper cites Introducing 4o image generation, 2025.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation Introducing 4o image generation, 2025

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:36.524284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:36.524284Z digest=sha256:bb433610b0bd83bb3a31a0c1c85550c92879ecd0a2f793d3ea8515d62f794196

Observation 9967247f-93ea-4f43-9020-89cb3ac4482b · outbound

This paper cites ALinFiK: Learning to Approximate Linearized Future Influence Kernel for Scalable Third-Party LLM Data Valuation.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation ALinFiK: Learning to Approximate Linearized Future Influence Kernel for Scalable Third-Party LLM Data Valuation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:36.531312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:36.531312Z digest=sha256:92fa00810a8f7d8110aad455e6c1d493e77fa042d76e7fbcbfb982c832254fd3

Observation 5ee96cbb-b416-476c-9567-5bb9f58cc224 · outbound

This paper cites Scalable diffusion models with transformers.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation Scalable diffusion models with transformers

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:37.929121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:16:36.536456Z digest=sha256:bbc0049627b874e23beec7eb4f1f10166f042d692fde2cc65c6af9453fe5873c

Observation 8279ae23-be39-4867-b902-05795387f8f3 · outbound

This paper cites Reasoning with large language models, a survey.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation Reasoning with large language models, a survey

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:36.542561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:36.542561Z digest=sha256:b357de31d05fff3f0d140c18719294ad44fa31cc2a3267245c466b99d274ba32

Observation a677e66f-0014-4f3e-aaaf-a910fedb4a03 · outbound

This paper cites High-Resolution Image Synthesis with Latent Diffusion Models.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation High-Resolution Image Synthesis with Latent Diffusion Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:36.548934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:36.548934Z digest=sha256:f5bb23ca187a7d829fdb0e0d4d485b0123dbafc2316fbbbc985a8318b322a1aa

Observation f678be5b-2c8b-47f3-a3c0-6b5e9c09c1c8 · outbound

This paper cites High- resolution image synthesis with latent diffusion models.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation High- resolution image synthesis with latent diffusion models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:37.902423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:16:36.557363Z digest=sha256:171a271c5f3db61ceabe10d26a3601f04a50d579545ce357603d14d8f3661807

Observation dcccecc1-3ad7-4d4c-becb-eac94a414194 · outbound

This paper cites Bernstein, Alexander C.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation Bernstein, Alexander C

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:37.878116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:16:36.563815Z digest=sha256:c3a51d0d21bbf7cc27e410d58dd903b93753b75f1ec86ec6132e1eaeb4ef7ccb

Observation 075afcea-125e-4910-ac84-9813d4aa96c9 · outbound

This paper cites Flow to the mode: Mode- seeking diffusion autoencoders for state-of-the-art image tokenization.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation Flow to the mode: Mode- seeking diffusion autoencoders for state-of-the-art image tokenization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:36.570432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:36.570432Z digest=sha256:5f4a778770aeaa58d65bf184be39804b3ca21099ed130038e67555fe1a6c0c85

Observation dedf71b8-daa2-4b2b-b43b-7ae23469a3f9 · outbound

This paper cites Drivelm: Driving with graph visual question answering.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation Drivelm: Driving with graph visual question answering

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:37.857053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:16:36.580593Z digest=sha256:39c3ab7da3533db394851a499eba2ac47b4ed976881eb9e90532d64324bc7233

Observation 71b0f393-a913-400e-854b-6d81b7579d0c · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:36.588401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:36.588401Z digest=sha256:bbfa916df5e6a5ef0fe798822df793727596f0cd5f82b125f09a7abaeff345fa

Observation a23b861a-d57b-4505-9476-2a0a2abbb635 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:36.594006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:36.594006Z digest=sha256:17fa9a948926ad870a09fbd49ebb3e52ec3ba7d646c6a7853dbd65083fdc6703

Observation 75d46c6d-b531-4021-8991-219e7328869e · outbound

This paper cites Visual autoregressive modeling: Scalable image generation via next-scale prediction.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation Visual autoregressive modeling: Scalable image generation via next-scale prediction

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:37.836930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:16:36.608247Z digest=sha256:0630388305fcb3e2a6b19da68a0a62269a52bf5ce5819d46a4105a37cef802a0

Observation dd99bfe0-6c01-4926-90c7-7c4efe76a536 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation LLaMA: Open and Efficient Foundation Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:36.614723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:36.614723Z digest=sha256:a5fc49bb7bd6b83e3176f84331ffb7810fe8cf4d9586bd00f7ac25726925bef1

Observation 7d22ed68-bec4-4b7a-a314-a11af275c057 · outbound

This paper cites Neural discrete representation learning.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation Neural discrete representation learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:37.807855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:16:36.622958Z digest=sha256:abcfbc9e25b3958063aa58166e1aa97c2ff9846dd18c61b2abe0392624287b42

Observation ce548559-a3d4-4656-aa8f-3527fdad1767 · outbound

This paper cites What Makes for Good Visual Tokenizers for Large Language Models?.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation What Makes for Good Visual Tokenizers for Large Language Models?

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:36.633448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:36.633448Z digest=sha256:8b38274082f125e5378cd85f82de3e661deee83abc6204b4ec5edab998f38168

Observation 29281ab3-185d-4244-a02c-96e4c07fcc0d · outbound

This paper cites MaskBit: Embedding-free Image Generation via Bit Tokens.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation MaskBit: Embedding-free Image Generation via Bit Tokens

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:36.642689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:36.642689Z digest=sha256:10e2e8731300918fd87da905ca8ebd58112267439a7004639d949f7a4c1b1067

Observation 9cf2c012-4a29-4f47-b383-f2762e1cc16e · outbound

This paper cites Liquid: Language Models are Scalable and Unified Multi-modal Generators.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:36.649273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:36.649273Z digest=sha256:ed018e5c800b596b33f6ac42a7299aa1726f01f04c253eca644d8015cf3caa05

Observation 340e14b2-33d6-4a46-861e-ad764a3fbe8a · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:36.655133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:36.655133Z digest=sha256:3501181d6b7292cf912887ccc9924686c5cc46e5a663b4b1e5734b4e72503c91

Observation 30cfff84-9fc7-4114-af0a-bfae6c15be7e · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:36.662678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:36.662678Z digest=sha256:b1f5bd224c9b7cf2f25862f55c232757ebc3ea99cebc4880c8677a98c59bb366

Observation 76efe441-a51c-424d-bced-dc848b720fc6 · outbound

This paper cites GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:36.676564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:36.676564Z digest=sha256:236883609892820857ca375763861a386fafb20b39ed807eb330d77c99bc58fe

Observation 8c692157-d8d9-4124-9191-adb6d8deaf03 · outbound

This paper cites Diffusion models: A comprehensive survey of methods and applications.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation Diffusion models: A comprehensive survey of methods and applications

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:37.780887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:16:36.685791Z digest=sha256:0ce2ac53f7bbae92d9cddba3be55045e712226b1848010fb11a2967ac2d43eae

Observation c372fb36-84b4-4a74-b57a-07ccf93ae185 · outbound

This paper cites Hauptmann, Ming-Hsuan Yang, Yuan Hao, Irfan Essa, and Lu Jiang.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation Hauptmann, Ming-Hsuan Yang, Yuan Hao, Irfan Essa, and Lu Jiang

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:37.759971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:16:36.692217Z digest=sha256:33a1dd05a51c1e348acbc8f0434bc31a3cd6df0bea80c5f30900c22a5869eb80

Observation be95616b-2bdc-4a0d-a803-151685542525 · outbound

This paper cites Hauptmann, Boqing Gong, Ming-Hsuan Yang, Irfan Essa, David A.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation Hauptmann, Boqing Gong, Ming-Hsuan Yang, Irfan Essa, David A

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:37.740883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:16:36.697988Z digest=sha256:f1648a98dfea8584093633327ff013780fd9c446d60bac597336d13d396e2ff0

Observation 69feac44-94b4-4641-be7f-0ce8ed4c4eb2 · outbound

This paper cites An image is worth 32 tokens for reconstruction and generation.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation An image is worth 32 tokens for reconstruction and generation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:37.715747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:16:36.704093Z digest=sha256:53e75b5f7ac33f83e4dac496ae9d35982778bca2cce172825fa18174dbca3c41

Observation 16d4eb7e-1952-43e1-8c7b-b7807d352c6f · outbound

This paper cites A simple LLM framework for long-range video question-answering.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation A simple LLM framework for long-range video question-answering

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:37.684490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:16:36.711866Z digest=sha256:f9551043dd2cf7f61dd304b92f67ef396d471f9e5ac227842aadd19dd194f998

Observation 5c3081dc-376c-4be7-b953-4568a12883b7 · outbound

This paper cites Efros, Eli Shechtman, and Oliver Wang.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation Efros, Eli Shechtman, and Oliver Wang

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:16:37.656742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:16:36.718820Z digest=sha256:c95dec43d0786b710e104eeb8356d5624415b6345558e93fb08ebcba811ebfb6

Observation 51f42578-6c45-4910-a670-8fdc1494ba87 · outbound

This paper cites Image and Video Tokenization with Binary Spherical Quantization.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation Image and Video Tokenization with Binary Spherical Quantization

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:36.724968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:36.724968Z digest=sha256:55df8af65154478d1b98d5fdb3a9fcd5bcd415b47ce8125e775693fca9829091

Pith citing papers

Observation f64afaca-72dc-4b45-a466-fe99c0a4e46a · inbound

InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation cites this paper.

InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:13:30.562810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T02:10:18.004724Z digest=sha256:3f0b4d90a02376dde4601b82d0e88c99ccb69793422d7cc3ebd1a2a24f2ea480

Observation 21b780be-50b9-42c5-a4de-f83e7427bed3 · inbound

Tokenizer Generator Coupling in Medical Image Generation cites this paper.

Tokenizer Generator Coupling in Medical Image Generation VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T00:27:00.510598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:27:00.510598Z digest=sha256:1d1eb0feac15b5b7b882f0ef62652a7c6629d52f949260f4b9bfc1ccc9a7a781