Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-20T11:01:24.738195Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 100 of 104 outbound references and 1 inbound Pith citation observation for arXiv:2605.18390.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-20T11:01:24.738195Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-10T07:31:26.225257Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T07:36:58.009410Z
100 of 104 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4e9bae38-76a3-4f26-b64d-6a8f9dc687b5 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Building Normalizing Flows with Stochastic Interpolants
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9ae82e62-88a1-428b-a51c-052472268aac · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation FlexTok: Resampling Images into 1D Token Sequences of Flexible Length
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c661347a-3350-49e0-91cf-d1b0c1dadefb · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Autoencoders
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2e79dfd9-c309-4ec6-a2c8-d2073b2c8f41 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 64ab607e-1f3c-4490-87e5-35bf5614a5aa · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0e330b87-a4c2-4292-8f06-e354d8e348be · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Understanding disentangling in $\beta$-VAE
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6531e696-8726-40cd-bc48-12cb9c65919e · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Emerging proper- ties in self-supervised vision transformers
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 577fd82e-2723-4dcc-9376-5c734c9ba9f1 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Maskgit: Masked generative image transformer
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7f62fa5b-83c5-4d07-abe2-462a66b270e9 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation arXiv preprint arXiv:2509.25162 (2025) 4
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 86b6e111-fbb5-4e10-820e-31a75b6bc23b · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Generative pretraining from pixels
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ff4ac094-3b5e-49b8-8662-4d07118b6c23 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Improved Baselines with Momentum Contrastive Learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3d8dbb5b-cc5f-4de4-8bf1-fdf35db38174 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Detection in crowded scenes: One proposal, multiple predictions
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e26eeeb9-6e02-437e-ac6d-78c6afd63313 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Deformable convolutional networks
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b9bfbb4f-1db2-49bd-9bbc-7ca5cd2704d4 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Vision Transformers Need Registers
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation aeedf7ae-d146-4186-9207-db5d53c7a1ab · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Imagenet: A large-scale hierarchical image database
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cf1cc1b3-746e-486f-a51d-46860a78ac60 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Bert: Pre-training of deep bidirectional transform- ers for language understanding
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fa51abb2-360b-4f8e-88ab-5d997fa4d3e5 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Diffusion models beat gans on image synthesis
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4fb7f249-21f0-464f-882c-6328b9301c73 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation An introduction to variational autoencoders
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 26638286-bf8c-4de3-9fe1-611ef3610a89 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2b4e0f94-ff25-42cb-a88a-fa90513818f2 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Scaling rectified flow transformers for high- resolution image synthesis
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6abe7458-1044-4e66-bb07-dacdc7e82c62 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Taming transformers for high-resolution image synthesis
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1f77c73d-7e69-4ecc-83ed-81eda16fc45b · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation One layer is enough: Adapting pretrained visual encoders for image generation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8b4c9a63-dc23-436c-b93a-47a3c8d8824c · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Generative adversarial networks
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7ef378e6-3f7d-4a45-9540-4a2dc0c8e1a4 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Bootstrap your own latent-a new approach to self-supervised learning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a0bbab4b-173f-4e72-9dca-99de3929f7dc · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Learnings from Scaling Visual Tokenizers for Reconstruction and Generation
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a16a26aa-333c-49e9-854b-f1be08784c49 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Masked autoencoders are scalable vision learn- ers
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e56fbd35-ca4f-4507-9bc9-783a3df32da7 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Momentum Contrast for Unsupervised Visual Representation Learning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 019a7724-c873-49db-8112-7d6f1cf05ccf · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Deep residual learning for image recognition
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d90035ae-f1a2-4a67-92e1-67019c71184f · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Gans trained by a two time-scale update rule converge to a local nash equilibrium
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6e09baef-298c-4181-b4e9-5b6504b1ba1d · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Burgess, Xavier Glorot, Matthew M
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0b37b325-454d-44e1-a563-a8013fd3d1d8 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Denoising diffusion probabilistic models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f0585fe7-5767-4c77-b329-bd31a01ef568 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Image-to-image translation with conditional adversarial networks
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4f8e6445-fe22-44e2-bc70-80fa8f60ab3d · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Image-to-image translation with conditional adversarial networks
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6045b28b-f457-4b1b-8ed7-45394e3763fd · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Scaling up visual and vision-language representation learning with noisy text supervision
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e6f2b59b-2ad7-4f60-b5fe-c988e04c6c37 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Guiding a diffusion model with a bad version of itself
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 54e3dc9e-5f26-49ef-a986-f4649d371537 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation A style-based generator architecture for generative adversarial networks
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4a9803c4-5bfe-4788-8ace-3196d619838b · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Auto-Encoding Variational Bayes
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c029ead1-642e-473d-8cc9-581b4d1971e7 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation EQ-VAE: Equivariance Regularized Latent Space for Improved Generative Image Modeling
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation daf38ee2-127c-4aac-a920-8d83afc4b710 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation arXiv preprint arXiv:2504.16064 , year=
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 21c00f1e-307c-4e1f-900c-12ff3fa1da0e · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7df956ba-741f-44d6-9e2e-3a05b3bc4ab1 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Improved precision and recall metric for assessing generative models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 54a38ccd-6ccb-43d5-b38a-bca635aaeadd · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Autoregressive image generation using residual quantization
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8943d534-2bfb-4ca8-a019-d727ef7a622c · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Repa-e: Unlocking vae for end-to-end tuning with latent diffusion transformers
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation dad91f52-f94d-41cf-9d64-0f721ae82904 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Autoregressive Image Generation without Vector Quantization
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 35a3c3f3-f2dc-45c9-a751-5897e4c373f8 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation ImageFolder: Autoregressive Image Generation with Folded Tokens
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5acccd36-1bd9-4940-949d-0244ff1b5051 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Feature pyramid networks for object detection
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d5aae5c8-23ec-46d4-92e2-49c1d50c9d0a · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Focal loss for dense object detection
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2a2f07b2-511f-4463-9d9b-f517be4a0f6a · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Flow Matching for Generative Modeling
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 22c21af5-a7c8-4c39-86d7-17f6d34aa6f1 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 442793b9-5d8d-4c0b-a763-7028d8ec488a · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Decoupled Weight Decay Regularization
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3ec1f911-62c2-4e81-9af7-34807fb6efb5 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ac649c53-4f79-42db-9414-97c75499300d · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Sit: Exploring flow and diffusion-based generative models with scalable interpolant transformers
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ea9248bc-495a-4754-86c0-9445a46f9fd5 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Finite Scalar Quantization: VQ-VAE Made Simple
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c1d8f6c1-2c4e-44db-917f-b19d81e6b7ca · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Conditional Generative Adversarial Nets
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 28ed9a64-345d-4177-b51e-004c01d54576 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation One-D-Piece: Image Tokenizer Meets Quality-Controllable Compression
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 679b6030-91bb-4466-b6d5-635efde81828 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Improved denois- ing diffusion probabilistic models
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 482d5ab3-81ca-4b87-871b-1a90bc3555b4 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation DINOv2: Learning Robust Visual Features without Supervision
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 03df1eec-62d9-4851-8293-a9205af759c1 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Scalable diffusion models with transformers
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ab1bc623-546d-41bc-9581-ddde19cc7ab6 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 706f594e-e853-4a1b-bbb9-1029e9398825 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Learning transferable visual models from natural language super- vision
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5f3a9e1e-5e47-41bc-9b62-e5e1fca80fb7 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a295e7ff-2403-452d-b365-6f7fe6d4121c · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Improving language understanding by generative pre-training
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 41b546d7-5b6d-44bd-b9c1-a05bee75204c · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Generating diverse high-fidelity images with vq-vae-2
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3839fa64-2f4e-4618-85ef-7ee48050a721 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation High-resolution image synthesis with latent diffusion models
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0148ab18-6fe6-454d-bc10-ca9f1d64189d · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Photorealistic text-to-image diffusion models with deep language understanding
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5b69732e-982d-40db-9847-e8e7ce214bfd · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Improved techniques for training gans
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c79eb551-b8f4-470f-893d-aedf7ba3b0ca · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Improved techniques for training gans
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 38afef1b-136f-48a9-b668-02f7675e4f8d · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Denoising Diffusion Implicit Models
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d7ce9777-833b-4ecf-bce9-29c32b809eb1 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8a6cc765-4bf0-43bd-99e5-aabdf14ca6b0 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Roformer: Enhanced transformer with rotary position embedding
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b5d0087d-6a23-48b5-8005-31327366d18d · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cea157d7-0fdc-4739-b4a2-12f7404084e4 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Rethinking the inception architecture for computer vision
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 78470744-c66c-40b4-af5c-a17b0d24c63b · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Unilip: Adapting clip for unified multimodal understanding, generation and editing
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6a3834f3-bfc7-4c5b-bde3-a61da64e68a9 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e31dac73-9b9c-4d41-ad24-aa928e3d3772 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation LLaMA: Open and Efficient Foundation Language Models
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 67f30945-4866-4694-97f0-29fd7f691ae1 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 11e29742-d5f7-4b76-9def-091f416dd664 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Conditional 15 image generation with pixelcnn decoders
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0eac972c-3aed-4f66-9646-7ab550a8fd74 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Neural discrete representation learning
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 44d42f76-37e2-4ac2-8f12-91e3dba13773 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9687e62e-6039-45ed-8208-77a743a3694c · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation DDT: Decoupled Diffusion Transformer
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 002edb18-0621-4af6-8e8b-d724b4c6b2f1 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Image quality assessment: from error visibility to structural similarity
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4425373c-e46b-4ac9-95ca-178df2ce92c3 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation "Principal Components" Enable A New Language of Images
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d9679bc2-2704-4bca-ae8e-297e3a91043f · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 883240f2-704e-47e5-a3a7-90a814ab2cec · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Gigatok: Scaling visual tokenizers to 3 billion param- eters for autoregressive image generation
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6e6f928f-0a8b-4c01-bfd3-f99235afa6cb · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Fasterdit: Towards faster diffusion transformers training with- out architecture modification
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 578ff45d-f513-48ad-aa24-688b1b80a3d8 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Reconstruction vs
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9f53fd8b-a68e-41da-ab0b-ec1eccde035d · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Vector-quantized Image Modeling with Improved VQGAN
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 96c3b084-3c8b-484b-8aa4-572b660b0c83 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Scaling autoregressive models for content-rich text-to-image generation
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9915826e-ee5a-43c4-91d0-4356c29ed97c · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Hauptmann, Ming-Hsuan Yang, Yuan Hao, Irfan Essa, and Lu Jiang
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e695229d-c8f4-48b9-a5e4-59aa7f00e9d9 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Randomized autoregressive visual generation
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation df817617-febd-48ed-9a38-f9ca51fe0792 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation An Image is Worth 32 Tokens for Reconstruction and Generation
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f16ad447-eab2-429d-b721-9d57f3d2e685 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b6592722-2279-4413-98d5-201a1398e249 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Sigmoid loss for language image pre-training
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d91df3cb-cce2-4775-9bed-e52ddb214be8 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation The unreasonable effectiveness of deep features as a perceptual metric
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c2180349-ac9d-45af-88c5-48f7968560c0 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Holistic tokenizer for autoregressive image generation
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5f27f037-a6bf-41ad-ae26-5f5442484552 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation arXiv preprint arXiv:2507.08441 , year=
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f40168af-01ef-4c09-9325-62c5c106458f · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Diffusion Transformers with Representation Autoencoders
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 764c9b29-8e43-4f63-b9c6-4cd0fe1d0125 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Movq: Modulating quantized vectors for high-fidelity image generation
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 07539e4e-6a3f-4ae2-ae0c-00fbec1cf69b · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation Fast Training of Diffusion Models with Masked Transformers
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 985670ae-37ab-402a-a5c9-6117c06516b1 · outbound
Vision Foundation Models as Generalist Tokenizers for Image Generation ibot: Image bert pre-training with online tokenizer
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cc960685-3011-4161-b367-c7aa67f85770 · inbound
DeltaV: Thinking with Visual State Updates in Unified Large Multimodal Models Vision Foundation Models as Generalist Tokenizers for Image Generation
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.