Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T22:54:20.710373Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 47 inbound Pith citation observations for arXiv:2412.03069.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T22:54:20.710373Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T04:41:47.730706Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
77 of 77 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 6d0b9588-1f0b-4fde-98c5-e1d5b1e072db · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a9c9a12-e813-42db-b699-35d41b04f9fe · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Qwen Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a600daf4-fa30-4998-98c4-aaf2a5cf9cc6 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e42f739-a626-4adb-9fdb-6eac2aac593f · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Improving image generation with better captions
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0126eda4-bacf-4cd4-8c8e-490363c9c0a4 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Coyo-700m: Image-text pair dataset
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83b9e60b-9a7b-4007-8f90-4783bc6ac15f · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Muse: Text-To-Image Generation via Masked Generative Transformers
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bde1a2d9-d9a1-4920-9081-92afa09a5756 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee66adad-dd0e-4eae-b10f-14d17ca91814 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Vitamin: Designing scalable vision models in the vision-language era
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3bcc1fd3-13ea-4a9a-ab99-7c06c4a9641e · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62f6e974-6380-4f05-8464-83a3484c7266 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0bcf96c-b243-4657-8491-d244bdf98ed7 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Scaling vision transformers to 22 billion pa- rameters
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 277ced2b-40c5-4f62-8949-5edbf2e6179c · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Imagenet: A large-scale hierarchical image database
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b41a819-0890-419e-aa81-45f85d443dd6 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Taming transformers for high-resolution image synthesis
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8442a6c1-1f43-4240-a106-4becb8e2f815 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b935a5e-e2e8-49b5-a594-cceb16895586 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Geneval: An object-focused framework for evaluating text- to-image alignment
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 34ba906f-63fd-43aa-9dc1-7255b69946b8 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77e81f5a-c599-44c3-8940-1b40d1538d1a · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Classifier-Free Diffusion Guidance
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1eb8ae94-2fc6-4c62-8163-8cd34638aa07 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 275b0fea-09db-4020-94dc-f24fb0517144 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9465b5a9-afb4-4bd5-9d1d-47c20d3b6441 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation A diagram is worth a dozen images
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc895db4-5e7e-4816-9e78-2b335767cbc0 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Autoregressive image generation using residual quantization
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1b2f7917-e8a3-4552-8f38-da51b150f588 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c17fdfb-fa90-4736-8ba1-43dc7d9771e7 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Seed-bench: Bench- marking multimodal large language models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation baa69641-420a-4977-aaee-7f0f29323e22 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc9ee8dd-5b24-4b43-a77a-ddfbe1152e28 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation ImageFolder: Autoregressive Image Generation with Folded Tokens
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 882d1ad6-b025-4264-988f-d5b5ab787d8d · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Evaluating Object Hallucination in Large Vision-Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8190837a-a921-4522-8ec4-960382bb187b · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd2555aa-4a88-491b-aa82-b381cddce663 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd19d9d4-709d-4af7-885b-579623b7efa9 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Improved baselines with visual instruction tuning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 81e0e3a7-ba44-455b-9d72-1a14cfebe7db · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Visual instruction tuning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ecdae2f-e6b6-43b7-b059-12bb71766442 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation World Model on Million-Length Video And Language With Blockwise RingAttention
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91137bbb-6da4-4879-b7ed-4efac947e0ca · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 25bbee17-e5d8-4dfa-a61d-269935bd978d · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24522319-7885-4295-8579-54bddb232fc1 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation STAR: Scale-wise Text-conditioned AutoRegressive image generation
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8332b65-6797-4a6c-b077-13e575320895 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a49e759-944f-4cc3-be32-39fcc5dc92f6 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 236c09c7-6794-46f9-b308-3c4730825e3f · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Learning transferable visual models from natural language supervi- sion
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d8f784aa-ffff-4faa-a70f-414a7e0a1104 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Zero-shot text-to-image generation
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d85a53c-c0cc-4ca7-a145-53fe4e229c54 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Hierarchical Text-Conditional Image Generation with CLIP Latents
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e145cf4c-c565-448b-9a35-40ca5035c484 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Gener- ating diverse high-fidelity images with vq-vae-2
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38b9f5a3-b13a-49ec-97d3-3cc4ef2844cd · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation High-resolution image syn- thesis with latent diffusion models, 2021
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation adf34022-b3d7-42cc-8bec-9feb71bf7ddf · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Laion-5b: An open large-scale dataset for training next generation image-text models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a54e9c2c-08d4-49d7-a53f-49f48bad442b · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Towards vqa models that can read
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 32dc96f1-7c70-44c4-ad5d-7dedb75769a3 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94e41039-f9a3-4a65-aece-a7f03e0d5363 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39475a2a-77f4-4587-8cf4-16dbd376f497 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Emu: Generative Pretraining in Multimodality
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c578ad44-d825-4931-a306-606d9cd9bbdc · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation HART: Efficient Visual Generation with Hybrid Autoregressive Transformer
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9661d7ca-622d-44f1-af91-f94d215ba70b · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d99523a9-a43e-4a55-9857-3de3ba680703 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Gemini: A Family of Highly Capable Multimodal Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4162a1dc-3979-49d4-ae3a-06d2d8f3a122 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Qwen2.5: A party of foundation models, 2024
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 830a81bd-5396-4b1b-a97e-418a0757aa38 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e749ca7-9d6e-4057-9833-238eb6aa9bc6 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efa95757-7033-4ac6-bdf8-8aad11a1e1bd · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbc79bd2-d88d-4710-b43e-82da69f3eec0 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Neural discrete representation learning
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e529a2b-a17c-4374-91d6-c2067f818c68 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Emu3: Next-Token Prediction is All You Need
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35d89243-fe77-40d6-9d27-edea4a0642c6 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Small-scale proxies for large-scale Transformer training instabilities
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e178fc9a-8e52-494c-a1f5-4f06585b27a2 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afb78ad8-e84f-47e9-ba30-fc20ec80c3e9 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation NExT-GPT: Any-to-Any Multimodal LLM
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87a6c221-2386-4a65-871c-cc21054f86ae · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation GPT-4o: Visual perception performance of multimodal large language models in piglet activity understanding
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f2b19e7-e8e9-4aa4-85e0-30d1f732fc58 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f44a8ff-e9ce-412b-8206-ffe46e8cf1d5 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Realworldqa, 2024
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation fcbb2340-86c3-4e17-8b36-e5a415ef2cec · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6e2e3f1-cc77-4f06-86c4-fc44a43933be · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Imagere- ward: Learning and evaluating human preferences for text- to-image generation
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1be0404-742e-4ae5-a5f0-6ac69a505cde · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Vector-quantized Image Modeling with Improved VQGAN
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ae76a07-7f94-4cd5-b142-23eecc37f21b · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Scaling Autoregressive Models for Content-Rich Text-to-Image Generation
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbb3738f-1681-489b-ba3c-9f5bb0bf6e5d · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7433235-cd95-4bab-90c6-6caff9664116 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97924604-63c1-4e25-a780-deb3a4ece033 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 790683d6-8e4c-4ea1-903b-2ed8dc57ebe6 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Sigmoid loss for language image pre-training
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation af8f51cb-c556-4eed-a38d-a92ff6ef9f5b · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Regularized vector quantization for tokenized im- age synthesis
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 619eabbd-f148-493d-8c4e-a0bb68dfbf66 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Lmms- eval: Reality check on the evaluation of large multimodal models, 2024
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6bc54eb1-c27a-4c9a-aca0-85f77db61c34 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Tinyllama: An open-source small language model, 2024
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2fa11b19-356d-4ae3-bce8-de512028d510 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Movq: Modulating quantized vectors for high- fidelity image generation
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation fc46919f-a677-458f-a2bd-1bc7a87618d5 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be76c2cd-affe-4c68-9975-795ac727ba99 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52a578d8-715e-4592-bc2e-8762c087d474 · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83bfaea3-6f60-4619-b6bd-e5e90a504cde · outbound
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation Unresolved cited work
Reference 251
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9eae1cb8-a1ef-4948-a502-f8e92d36566b · inbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88f6eeff-8513-4a51-849a-7144f2448515 · inbound
Scalable Image Tokenization with Index Backpropagation Quantization TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0e22818-1e16-4dc4-ae1e-b17e0b3e6e2b · inbound
Liquid: Language Models are Scalable and Unified Multi-modal Generators TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5df510ed-3dba-488d-a51c-b6cb1d70606b · inbound
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43d0d774-8cf4-4c06-babb-9e4c11199368 · inbound
Next Patch Prediction for Autoregressive Visual Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b9e76d7-d684-4499-8d69-a1284bc0d394 · inbound
Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 996164ff-57dc-495c-91de-c36211141636 · inbound
On Fairness of Unified Multimodal Large Language Model for Image Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfbc628b-1402-4466-848a-af856d58595a · inbound
Masked Autoencoders Are Effective Tokenizers for Diffusion Models TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d12143ba-0a32-42a9-a00a-e6f1a22c4ae1 · inbound
Multimodal Medical Code Tokenizer TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88e622de-3e0f-4eca-bf68-b219530fdf91 · inbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cca3d66c-37d2-482a-8e41-0ba9f8c5a2c1 · inbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f55f4173-db02-4d93-a5e4-96030d315cea · inbound
DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e6a9ae03-2e7b-4916-846d-3c9c7d390273 · inbound
T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7807693a-b4d5-4c74-8b19-ebf9a6b6560e · inbound
Ming-Lite-Uni: Advancements in Unified Architecture for Natural Multimodal Interaction TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3a2bb0e-85f8-4263-95f9-52aa70fd6cad · inbound
Position: Foundation Models Need Digital Twin Representations TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2fffa2c-5aa1-4434-9c32-e0510cea9d2e · inbound
TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9e53c24-575a-4fc3-b371-025ed7011778 · inbound
Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2356c3ea-3ece-449c-b702-cc2e7d541be6 · inbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6869a934-ada0-4a3d-a060-0bf823273652 · inbound
BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ec67038c-8d8a-4cdd-b29e-5fc2732a4d94 · inbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 078059fc-f144-40af-b551-fc62dd3a4209 · inbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98fed8a3-a44b-49fa-8069-5a61c92b004b · inbound
Emerging Properties in Unified Multimodal Pretraining TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 49711da3-9b0c-44c5-8b6b-ba35915ed162 · inbound
TokBench: Evaluating Your Visual Tokenizer before Visual Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9455675-b83a-4d95-beea-188e4e763b43 · inbound
FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ed1468c-7cf9-4a0c-b1c7-8dcfafffb433 · inbound
UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29dcb26a-57e5-44e0-8b57-e43c8271223a · inbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation deba6181-a615-42f0-841b-d0ff73d2e289 · inbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f69c249c-b6b3-444c-90c8-e4e86784bdca · inbound
Ming-Omni: A Unified Multimodal Model for Perception and Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb7f0dab-7b70-4b83-b7d9-fbfcbae00a30 · inbound
Show-o2: Improved Native Unified Multimodal Models TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ce07186b-2fa2-4449-9253-24ff0f49ae5d · inbound
UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78609e91-2f10-4de1-8f9d-266779a24f78 · inbound
Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1aa9569-6165-4191-ae36-9f863a158d8a · inbound
Rethinking Discrete Tokens: Treating Them as Conditions for Continuous Autoregressive Image Synthesis TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e77a85cc-3c45-4774-b6cd-e6648a3feeb8 · inbound
DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid Tokenizer TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07056f24-8547-4f77-acc2-e499b83f6a3f · inbound
A Unified Low-level Foundation Model for Enhancing Pathology Image Quality TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6a80fd0-984a-4d3a-b633-fccfc4c216a0 · inbound
Interleaving Reasoning for Better Text-to-Image Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8d924e2-91f0-486d-85fd-6eaf58b666e0 · inbound
Reconstruction Alignment Improves Unified Multimodal Models TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b88afb4d-7579-4fad-ae74-0661592f29e0 · inbound
A Unified and Controllable Framework for Layered Image Generation with Visual Effects TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 609a9b36-99f1-4287-a95c-8be0c11ae5cb · inbound
FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f7bec7b5-3764-4502-9b7e-e1511547bb45 · inbound
FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 211c994b-f3f4-45f2-a9ff-72e6f011aac3 · inbound
Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 85b178fd-4ad0-47f0-974e-080507de5924 · inbound
Free Lunch for Unified Multimodal Models: Enhancing Generation via Reflective Rectification with Inherent Understanding TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation fec4d6af-7265-4a0e-a833-957939712b01 · inbound
Meta-CoT: Enhancing Granularity and Generalization in Image Editing TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 226722f2-0983-4fef-883d-c2374909187d · inbound
Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b7e1ded3-b160-45dd-bb73-6d9cfdeb56cc · inbound
What Matters for Diffusion-Friendly Latent Manifold? Prior-Aligned Autoencoders for Latent Diffusion TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ab1bc623-546d-41bc-9581-ddde19cc7ab6 · inbound
Vision Foundation Models as Generalist Tokenizers for Image Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 179e42ab-e52d-411d-96ca-81da6fff4ca9 · inbound
Histogram-constrained Image Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c07c4d84-88bb-4bb0-bf4c-91baa3e94539 · inbound
OSVE: One Step Video Editing with One Step Diffusion Models TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.