Pith. sign in

Paper Citation Record · LEDGER

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations

As of 10 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2604.24885.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.24885 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-08T04:12:58.100610Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact26
  • verified fuzzy26
  • unresolved1
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 57347b56-229a-4a40-b72c-9ea9799659a4 · outbound

This paper cites Flextok: Resam- pling images into 1d token sequences of flexible length.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Flextok: Resam- pling images into 1d token sequences of flexible length

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:18:02.862659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:8d37ff2944b02ed18cbaa9db07a73173f32a6337b9c874c44e009bb897388983

Observation 64e49770-3c38-4960-b8b4-a9da760d4746 · outbound

This paper cites Flexivit: One model for all patch sizes.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Flexivit: One model for all patch sizes

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:18:02.846842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:d66002e6cc87d3f5d2f053372e9ff8f39593d4734c800bae94582d63d6ca3f44

Observation 434c5299-0fb1-43f7-a2e5-5a1b93e1cc7d · outbound

This paper cites Maskgit: Masked generative image transformer.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Maskgit: Masked generative image transformer

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:18:02.887147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:44737c26dfeddd6134b73e0522b7905994b327d26e0cef811708b90413806a2e

Observation c09e4615-7060-4e87-b943-575ace00a510 · outbound

This paper cites Masked Autoencoders Are Effective Tokenizers for Diffusion Models.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Masked Autoencoders Are Effective Tokenizers for Diffusion Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:51:11.648909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:cfe71378457d6f2e913415aa6e937c3fe048d80b36272f4f104f78206ccd89e1

Observation 659429a7-2b41-48b9-a9e1-da3fcf2e4216 · outbound

This paper cites Softvq-vae: Efficient 1-dimensional contin- uous tokenizer.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Softvq-vae: Efficient 1-dimensional contin- uous tokenizer

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:18:02.821232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:95e68b2dd763ece9882fcf079ca655bb4b730970b8fc0491ecb9ac732368fb88

Observation c6113849-709c-40f5-aadd-d97579991a1d · outbound

This paper cites Patch n’pack: Navit, a vision transformer for any aspect ratio and resolution.Advances in Neural Information Processing Systems, 36:2252–2274.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Patch n’pack: Navit, a vision transformer for any aspect ratio and resolution.Advances in Neural Information Processing Systems, 36:2252–2274

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:18:02.859544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:411b4f67954625be5a761a73600666d71119b63b1a8431c8b52d5f10e86a9f59

Observation 561635ae-df8b-4026-8aea-c45a87639141 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Imagenet: A large-scale hierarchical image database

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:18:02.843827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:c1cba55c7c81f198f839b7cefa535b2ccb990d74b8e67efb3ec8bd7869b53aa6

Observation 27f8afd2-b80d-4439-8b2f-124070b2c538 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:51:11.659546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:a6d74a22b1551c2120588fab1f255ce4e0e38b824082cb1878661f67d3737508

Observation 934dc4bb-aa79-408d-91e1-f46463aa5567 · outbound

This paper cites Adaptive length image tokenization via recurrent allocation.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Adaptive length image tokenization via recurrent allocation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:18:02.838006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:95483cb8681e7ead290f4991012651c6485df8298675985b425bd3152bafde18

Observation 6259f07e-1825-492c-8e6d-fe80cb2adc4a · outbound

This paper cites Taming transformers for high-resolution image synthesis.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Taming transformers for high-resolution image synthesis

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:18:02.882166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:e7b6d87427dd2f3f687ce251af45fdb87f7aac460b9c1d9ffdb22463e76a1e92

Observation 213e37ba-8ca2-44b4-8ef9-8e319332b181 · outbound

This paper cites Scaling rectified flow trans- formers for high-resolution image synthesis.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Scaling rectified flow trans- formers for high-resolution image synthesis

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:18:02.811586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:ddef2eb4f80d6bdefa99b55bcfc4e2b4230bc6d9bea9340eccf579ae91c93435

Observation bdf6c9be-9284-4ece-b053-17471d34d9bc · outbound

This paper cites ViTAR: Vision Transformer with Any Resolution.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations ViTAR: Vision Transformer with Any Resolution

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:51:11.513644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:a822fe74f80115130e37fee76c5edff20fb4cec3574fc85a1552a54b44cf8e91

Observation a4f2c341-6d84-41c2-bf2a-b451ffeaa4b5 · outbound

This paper cites Rotary position embedding for vision transformer.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Rotary position embedding for vision transformer

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:18:02.879255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:ef7f029b8c65d181d93fba403e3b8cafd7919c1dd919632026107d6c00513c93

Observation c1f10377-ebe0-4e8d-a748-bc90f0bb0943 · outbound

This paper cites Classifier-Free Diffusion Guidance.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Classifier-Free Diffusion Guidance

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:51:11.547896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:870807f654d7e762a3eab4451b7393f1511bc00f369682d550c5bad05b70fb60

Observation 5997b0d6-9962-4ab1-a76b-d585c55a2342 · outbound

This paper cites Analyzing and improving the training dynamics of diffusion models.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Analyzing and improving the training dynamics of diffusion models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:18:02.834401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:5d2505a21497e7b8e5913051a46c1bd93c59bcb5fb9f6dc6123de50f28b119c2

Observation c2b46e44-ba77-4fcb-ad08-b8836ac1689f · outbound

This paper cites Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:51:11.644749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:925b13079a23c5b8d8b77ddb74cf913ba1a2dae1996c2ae4db74dee979cde7ec

Observation e03a9532-2e62-4241-b42a-e83b8bf3cc6b · outbound

This paper cites Autoregressive image generation using resid- ual quantization.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Autoregressive image generation using resid- ual quantization

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:18:02.889946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:3dd0c118c43216d048769313448eeb791fc6cdab58e73a1039c61f052c47e97a

Observation 52344866-9057-450e-b89d-e8fba5130407 · outbound

This paper cites Autoregressive image generation without vector quantization.Advances in Neural Information Processing Systems, 37:56424–56445.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Autoregressive image generation without vector quantization.Advances in Neural Information Processing Systems, 37:56424–56445

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:18:02.865501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:e6bfc420f287832a579224c299f7557853a58a0801739cbfd9ed89afa660dfd7

Observation 8da6bca7-b891-4169-a2cc-b839f7494d03 · outbound

This paper cites ImageFolder: Autoregressive Image Generation with Folded Tokens.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations ImageFolder: Autoregressive Image Generation with Folded Tokens

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:51:11.526347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:0fb9113d30644b5cbad058e81173968d1e231588c3da4e4d10bee33a69c88237

Observation 45066fc8-bc02-4a1d-9990-fe5678208045 · outbound

This paper cites an unresolved cited work.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-05-26T21:18:02.814581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:785d9488c76e181675134f37237e62ca26a3c11b02a494f8664394865c766a1f

Observation 1af6f34e-2367-4104-955a-afa89bf6201c · outbound

This paper cites Detailflow: 1d coarse-to-fine autoregressive image generation via next-detail prediction.arXiv preprint arXiv:2505.21473, 2025b.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Detailflow: 1d coarse-to-fine autoregressive image generation via next-detail prediction.arXiv preprint arXiv:2505.21473, 2025b

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:51:11.682061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:8a3e2ad70c8300c9174128198e1628aed320446cf9c90923fe3953dfe36b0fdd

Observation 75c7c2ba-257d-49f1-b597-737da24701eb · outbound

This paper cites Zeyu Lu, Zidong Wang, Di Huang, Chengyue Wu, Xihui Liu, Wanli Ouyang, and Lei Bai.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Zeyu Lu, Zidong Wang, Di Huang, Chengyue Wu, Xihui Liu, Wanli Ouyang, and Lei Bai

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:51:11.577778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:f2bd430850dbde6302f06b443e7e2b7a2d421a9a79cc306bdf781c025bfcb01b

Observation dce0c06f-883f-4ace-84dc-3b155bf75d1c · outbound

This paper cites FiT: Flexible Vision Transformer for Diffusion Model.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations FiT: Flexible Vision Transformer for Diffusion Model

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:51:11.626899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:2c3566a80a3afe036b3e37c7d7ade59e86cb935bcfb29033232bbf209963ce43

Observation 985ae8cd-ea87-4ad6-885d-83e24cc3f028 · outbound

This paper cites Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:51:11.569171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:fad8fc105c96bff6f1da96e35e3118f5d60ead57d23bc80d29f23d7f9d20153d

Observation db4a9c4e-867f-4f4c-bb80-d25088bc27b3 · outbound

This paper cites Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321, 2025a.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321, 2025a

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:51:11.556479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:b5b7f0551d8f2a500ba1fcc05b2592bb4c08fb7ff90045811bb24ce01e372ff9

Observation 09a20df6-e5ce-4cdb-9e50-3408b6da4035 · outbound

This paper cites One-D-Piece: Image Tokenizer Meets Quality-Controllable Compression.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations One-D-Piece: Image Tokenizer Meets Quality-Controllable Compression

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:51:11.676182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:e4107b4de0dec3273cc874d0fb05fc353595ba5ecd5d80b9189c369c5c480f20

Observation e5fa4440-0be9-4e2f-8f37-04131caef4a9 · outbound

This paper cites GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:51:11.653133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:edcb7c84eecec1d96787d694b155df025f82b99c7ac3e4908b72eb5aea3bbb22

Observation 10f0f233-e9fb-48dd-99e5-ce09c6c36be6 · outbound

This paper cites Eclipse: A resource-efficient text-to-image prior for image generations.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Eclipse: A resource-efficient text-to-image prior for image generations

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:18:02.849773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:86f57daeb893464f0b2d81922d0947b4df8cf035830c7914515eb73122748d3b

Observation 7b4dbec5-9102-471e-b51a-7f9aa3d17363 · outbound

This paper cites Scalable diffusion models with transformers.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Scalable diffusion models with transformers

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:18:02.818000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:11f8a0026ae0a6864037d5d50322e77cff66b18d23e67ee642299d48fa291433

Observation 67e3e1fc-c498-4535-938b-cc2b9e4da1be · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:51:11.666897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:40c184ee49fc24105a50ae2e348d167eb0139f33f215d454df7d5f270c24072a

Observation ba183e9b-27b9-4f47-9016-ddbe666ec7d5 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Movie Gen: A Cast of Media Foundation Models

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:51:11.530780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:5646f32a3720120be854bad3a1d5dbbbc68e16d37cdb538ace0d8fc7ffe784b0

Observation 16b99de0-9e75-4799-bf7c-3791e76b9b43 · outbound

This paper cites Image tokenizer needs post-training.arXiv preprint arXiv:2509.12474.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Image tokenizer needs post-training.arXiv preprint arXiv:2509.12474

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:51:11.615558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:04fec63d876a4d9ed9abfc82054a21f7de86c38ba38d09c194c775b91e1b4fb0

Observation b4cf67fa-625d-4082-987a-217275364fa2 · outbound

This paper cites Pho- torealistic text-to-image diffusion models with deep language understanding.Advances in neural information processing systems, 35:36479–36494.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Pho- torealistic text-to-image diffusion models with deep language understanding.Advances in neural information processing systems, 35:36479–36494

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:18:02.856540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:6d9df633737ee877721479765903d8e8359cab98405dcf0fc9cf2a5aebf8dad8

Observation 58413aa5-063e-46be-86ea-21d35acf4b24 · outbound

This paper cites Stretching each dollar: Diffu- sion training from scratch on a micro-budget.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Stretching each dollar: Diffu- sion training from scratch on a micro-budget

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:18:02.876754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:e5e7f37efa749734938fc7f13c366749a8bd023ada585d161739219148940089

Observation 978342ce-977a-4c8b-b183-ffe6c2a54487 · outbound

This paper cites Scalable image tokenization with index backpropagation quantization.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Scalable image tokenization with index backpropagation quantization

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:18:02.840624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:91b8f24cad28fca0bf28f94b191929e942062ef77d5947becad3c0ddaab52b3a

Observation e15859d8-54ca-4ffa-818f-e371fed166ff · outbound

This paper cites Denoising diffusion implicit models.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Denoising diffusion implicit models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:18:02.873945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:7d2a33936297b674eaf319a09f77f4329fa0d15644fb0542c64886701ecb05a3

Observation e83dd030-e40f-4a8e-baed-ae2fc66cd853 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:09:17.130563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:45c541ea928332c69611c60253eb01e5375adc3a39793d199acc911d935368a0

Observation 26a7f1ea-5c29-474e-8958-77e998b00eb3 · outbound

This paper cites NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:51:11.560529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:db6ba4f81490afe2e5589ab18d7e6ed89917e1e87aa02f6e0b2730d3379140b8

Observation afe7c311-37f6-42ab-9352-afff9c4e83c7 · outbound

This paper cites Visual autoregressive modeling: Scalable image generation via next-scale prediction.Advances in neural in- formation processing systems, 37:84839–84865.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Visual autoregressive modeling: Scalable image generation via next-scale prediction.Advances in neural in- formation processing systems, 37:84839–84865

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:18:02.868467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:8c1522493a306978209e436f3e18be468ebfdeadf4f24ca29f0d6728e6205eda

Observation 67d99c61-5f57-4f97-be52-00d550e36925 · outbound

This paper cites Resformer: Scaling vits with multi-resolution training.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Resformer: Scaling vits with multi-resolution training

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:18:02.871426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:1986f7b4d0fe54b309dca57be43b67d60cdfcf972719814a464e192210d1b774

Observation 2fd854af-b2e2-4a09-80a4-f305a4264836 · outbound

This paper cites Training data-efficient image transformers & distillation through atten- tion.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Training data-efficient image transformers & distillation through atten- tion

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:18:02.831749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:3f1cd1ff36efd75004daf0325c1c2eb49174c84ef6ee6874ed7ad1ec8476c980

Observation a9b0a434-6be3-4b1c-a30e-a19c43e36994 · outbound

This paper cites FiTv2: Scalable and Improved Flexible Vision Transformer for Diffusion Model.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations FiTv2: Scalable and Improved Flexible Vision Transformer for Diffusion Model

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:51:11.538946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:dda9af56a26bd634e0b6624575b52d6ad7727110e48c4daf6abc3bf7038a13d3

Observation cdf67ea2-5bff-46e8-b1f9-b09fb03a980a · outbound

This paper cites Native-Resolution Image Synthesis.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Native-Resolution Image Synthesis

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:51:11.672181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:8d0145aabe1337fde0246b5960f59974853892359507b1dc9b993eb0d0bee69e

Observation 6502a31f-1309-4cda-a96c-c72a29e05738 · outbound

This paper cites Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:51:11.622758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:c8bfdb4929d0dc4d935b72a26a38a3f3e18e2fb294e5b644efea32ed8ae38257

Observation c82e461e-cb88-49ad-a1d1-fb78bd30bce8 · outbound

This paper cites MaskBit: Embedding-free Image Generation via Bit Tokens.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations MaskBit: Embedding-free Image Generation via Bit Tokens

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:51:11.606278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:65d81ca3c84a251b2fb069c7dcaf1958708f96a8d6e26c9e6c2216c94e6c214e

Observation 0aca68f9-74b9-4aa2-b3ae-d54d87b9bbfd · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-16T00:26:21.484895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:5351b2972e24ab8db7d825c8a2574611e1e9b3cd3f470200fdaa439cb9bddf51

Observation df99eef3-7c35-45a5-9c5c-0ac8d3b5eb0d · outbound

This paper cites Dc-ar: Efficient masked autoregressive image generation with deep compression hybrid tokenizer.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Dc-ar: Efficient masked autoregressive image generation with deep compression hybrid tokenizer

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:18:02.892713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:e524956d793b759ff92458f7f22a31c89c81e251d43b544d8084c6e35f71ee93

Observation bb64f730-1fe1-4378-8016-f68000c1d019 · outbound

This paper cites Layton: Latent Consistency Tokenizer for 1024-pixel Image Reconstruction and Generation by 256 Tokens.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Layton: Latent Consistency Tokenizer for 1024-pixel Image Reconstruction and Generation by 256 Tokens

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:51:11.587409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:8a4e5915a24dd43d2e0fc28fe5c6a783a13ebf5a5c1979c3c27cc05f236c4320

Observation 1973197e-3d92-42a3-9c2e-8e3bbc67569a · outbound

This paper cites Vector-quantized Image Modeling with Improved VQGAN.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Vector-quantized Image Modeling with Improved VQGAN

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-16T18:40:37.571629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:dcda8bb87331d1e8fac429406d1a700650f656b66a2e88ca1402c5eb6c483327

Observation 7ca49328-e71e-409c-b5c0-1491f78a76af · outbound

This paper cites Scaling Autoregressive Models for Content-Rich Text-to-Image Generation.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:49:31.747550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:8b3d91c91475ab751da8ef3246b2625fa2ffed3b159176f4817a27312cf0c903

Observation df0314d0-fa8c-49ba-ab81-30ee5c88deba · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:06:45.090605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:76bbf810451f7996234790228185009041763b10463c918e3488452e0cc4a917

Observation 3b0d655a-6315-4eb1-94b6-6ca6eb6954fd · outbound

This paper cites An image is worth 32 tokens for reconstruction and generation.Advances in Neural Information Processing Systems, 37:128940–128966.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations An image is worth 32 tokens for reconstruction and generation.Advances in Neural Information Processing Systems, 37:128940–128966

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:18:02.824410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:c3316ac206ad9efb045d5d1e856c969b72691a452f557ba4f407b39e0c0fb884

Observation ecd18e81-4a99-4ad6-8231-846e66123614 · outbound

This paper cites Randomized autoregressive visual generation.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Randomized autoregressive visual generation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:18:02.852550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:e4c682106cfdf4eb9c2554f30befe2b9132d6e0a4f057b8f9eea4f9f73be8ae5

Observation 03374ab6-e24c-4841-97a8-80110ecef532 · outbound

This paper cites Ar- gus: A compact and versatile foundation model for vision.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations Ar- gus: A compact and versatile foundation model for vision

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:18:02.884762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:63e401aa6d89d23370a7ea05ed5867d9ecf95910169864190de321c3433b49e5

Observation 91b99e6b-dc58-435f-bb19-159511225476 · outbound

This paper cites 256x256": 0.5.

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations 256x256": 0.5

Reference 55

Resolution
malformed identifier
raw_fallback, observed 2026-05-26T21:18:02.827657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:58.100610Z digest=sha256:6e0159059cc241f6a8f89b466eb592b832a9f88f687f5b4fb0503caea0e17b8e

Pith citing papers

No inbound Pith citation observations are available.