Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T17:00:09.097663Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 100 of 113 outbound references and 9 inbound Pith citation observations for arXiv:2412.09604.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T17:00:09.097663Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T21:35:40.295495Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-25T05:50:24.280970Z
100 of 113 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 64717b3e-5766-41ef-bc33-24fb9ff5f58d · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Qwen Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52df0d8e-5e4b-4c91-ad81-67c963194a47 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Introducing our multimodal models, 2023
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c06b6b5b-7756-4c95-8d47-81284370bb54 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Improving image generation with better captions
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74690afe-ae2d-4c07-a240-ce62d8a052d7 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding PaliGemma: A versatile 3B VLM for transfer
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b31bf2f-9ede-4edf-889e-4b218696c7d6 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Scene text visual question answering
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 149636dc-4047-4c18-baf0-d8e3ccd562e8 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Coyo-700m: Image-text pair dataset
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 469a3509-912e-4cef-ad0f-e7775866a0cc · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding InternLM2 Technical Report
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c8501eb-da92-42bd-9a93-275274bfe0d2 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fa3bc14-ee38-46f5-805c-451a5aa2d819 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0c3328d-b256-43ab-91ea-11c8c42329ab · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding SOLO: A Single Transformer for Scalable Vision-Language Modeling
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6eaa2e8e-9a92-4ffc-a583-6cfb99c13621 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff0bc6a6-0d81-449a-b057-7d09068f27b0 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Gonzalez, Ion Stoica, and Eric P
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ae26fac-b719-49c0-bf2e-634442dfe42f · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Icdar2019 robust read- ing challenge on arbitrary-shaped text-rrc-art
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcd04dcb-f70d-4cd8-85c7-3ee7375efbf0 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38b1d0b2-bff2-44c0-9818-d6b0adaaead3 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Simple and effec- tive multi-paragraph reading comprehension
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 335d9490-860c-4909-9367-94d530bf6ff9 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Opencompass: A universal evaluation plat- form for foundation models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d232235-38d6-4cd3-8067-9e0d3da90210 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Funnel-transformer: Filtering out sequential redundancy for efficient language processing
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2dc95c6-0a66-4cec-90c5-5d0ba44fe7e3 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Imagenet: A large-scale hierarchical im- age database
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16ee3840-c0b6-4cfc-b6cd-a81b11080e4d · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Unveiling Encoder-Free Vision-Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75c0af2d-1227-46a2-ac31-40fded98835c · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Dreamllm: Synergistic multimodal com- prehension and creation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 911769e8-a758-46ad-a81a-229f218f1da3 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Vlmevalkit: An open-source toolkit for evaluating large multi-modality models, 2024
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1a80be6-4fd1-4f70-b0ce-c21512933a8e · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Taming transformers for high-resolution image synthesis
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 013fd86f-f901-4e86-9735-ee1bfb0fec5d · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68bbb49b-9e7c-446d-8a3d-0c96b5ba8c6d · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04779d43-a551-4012-8363-8c79d7b22d1b · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Geneval: An object-focused framework for evaluating text- to-image alignment
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e38d3f0f-f276-4cc4-be6f-7748058118c1 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9be28103-c8dc-44c3-b7b3-8d19d339168e · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Gans trained by a two time-scale update rule converge to a local nash equilib- rium
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 106f2a8d-9f05-48fc-8eb2-2614b2393266 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Block Transformer: Global-to-Local Language Modeling for Fast Inference
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dd45ea8-1df6-4c6e-bd3b-0d8003951a46 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Hudson and Christopher D
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b620bf40-8876-43f4-9941-598a7bd4da2d · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Unified language-vision pre- training in llm with dynamic discrete visual tokenization
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 272e9666-5abd-41ee-ad29-07d525c51c4a · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding A diagram is worth a dozen images
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 498a6a59-1ea5-4cb2-aea0-a32e5f9a3bfd · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Ocr- free document understanding transformer
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b3bda50-48e1-46a2-83de-6c43b090a86a · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Segment Anything
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c8051d8-123f-4bc3-8092-360ddb98fc38 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24f69ec6-b6c5-49d2-ae41-769e4211d4db · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a31a52af-6750-4c01-9158-4dd633caea2b · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Evaluating object hallucination in large vision-language models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 214cf737-bda8-40cc-b2f3-d8c40d832ef4 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94cc0d62-b2ba-4968-84f9-fe064100517e · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 953ab72b-0ed0-4c7f-9633-134915fe8266 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a3ea85b-a8eb-462c-95be-3027a66e2f37 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fa7e85b-d979-4fc3-8779-87d05185bff0 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Improved Baselines with Visual Instruction Tuning
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6542c7df-d06f-48c2-ad87-151e4a6aab20 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Visual instruction tuning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18066164-3919-4d33-9fb1-6e4b2a0ec29d · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding World Model on Million-Length Video And Language With Blockwise RingAttention
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 823748f5-ebdf-47d1-9410-ea0dc575633b · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding MMBench: Is Your Multi-modal Model an All-around Player?
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b231a867-28da-4937-a338-28fd809b4f31 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c6fa849-dae6-4939-9a3b-948912f29079 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding DeepSeek-VL: Towards Real-World Vision-Language Understanding
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95ceb76c-ab1b-4a18-b139-e4de7dfd9faf · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Learn to explain: Multimodal reason- ing via thought chains for science question answering
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 759827bb-5dd1-4dfb-b87c-9cc116ef2d52 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff1527d0-8cf4-40b9-89e2-24b1568aa1ac · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ab74eb3-4106-4e02-9452-c0710906d1a0 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Megalith-huggingface
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb06c487-3230-4df5-bc62-809a502e7238 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Chartqa: A benchmark for question answering about charts with visual and logical reasoning
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05299ded-ce4b-4456-ac0c-314d96aa490c · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Infographicvqa
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4532c57a-e7b4-43fe-865e-cfe7eed7a93a · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Plotqa: Reasoning over scientific plots
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47ba5385-a623-4650-af1f-1098c0ff8435 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Hierarchical Attention Encoder Decoder
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e27deaa8-8d15-40dd-b367-e53e8c1d7eeb · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Datamux: Data multiplexing for neu- ral networks
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99e1cc0f-5b05-44ea-8e9b-3ed3b673eaa4 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c80ed520-c997-4125-ba85-ebcd734c8541 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding GPT-4 Technical Report
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3f4e258-fe17-4fc3-9245-f566b3cb439d · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dbf7f94-e3c6-46e9-809a-30f5351066a6 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fbfb2d1-850a-4ddd-921c-8a66a008e2d0 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Learning transferable visual models from natural language supervision
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91b37229-ae72-4618-938b-b506c60cb5a2 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Zero-shot text-to-image generation
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4a80815-c067-4f9d-b21b-d067927c4c59 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Hierarchical Text-Conditional Image Generation with CLIP Latents
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ef2d716-f786-4666-82af-c461d8c1c00b · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding High-resolution image synthesis with latent diffusion models
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e68d3105-79b6-4e27-b487-bfd73f660c2f · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Photorealistic text-to-image diffusion models with deep language understanding
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2332029b-12ff-4e04-bf90-da2ddc051afe · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Laion-5b: An open large-scale dataset for training next generation image-text models
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c8e5273a-56d5-46f8-a807-040844f3d113 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Laion coco: 600m synthetic captions from laion2b-en
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 65db8cd2-ac47-4fec-a726-1397959be663 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Objects365: A large-scale, high-quality dataset for object detection
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6cbf65a4-a649-4f50-94c6-09e1b6a3910e · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Icdar2017 competition on reading chinese text in the wild (rctw-17)
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 959718d6-7bdb-469a-8019-85e987e29bc0 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Textcaps: A dataset for image caption- ing with reading comprehension
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 768ed25d-a7bb-4c4f-a38a-6fbe8c187ec9 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Towards VQA models that can read
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bd3c2e8-7aa1-4505-9165-c7ef25287978 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Textocr: Towards large- scale end-to-end reasoning for arbitrary-shaped scene text
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96cb2e30-59aa-4257-a0c6-f3a388c19d7c · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Journeydb: A benchmark for genera- tive image understanding
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 079cfd98-7c10-4185-bd09-ef0be4b338b8 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 402872a8-4df9-4a19-a347-359994e98ea9 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Generative Multimodal Models are In-Context Learners
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f18f29e1-cc4d-4038-9989-a633c6a96516 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Emu: Generative Pretraining in Multimodality
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0a777c7-9113-4297-9fed-e80617cdc55f · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Generative pretraining in mul- timodality
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f8b244ba-eaed-4688-b8a0-9e6bee3a2ec6 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Icdar 2019 competition on large-scale street view text with partial labeling-rrc-lsvt
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 51119b44-e400-4c8a-a388-6b047d62fb0d · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Gemini: A Family of Highly Capable Multimodal Models
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe5c4185-b6e7-4225-be54-3b6011fb1680 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d37afe2-c69f-487d-a723-3f19ffceecc5 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07d95d67-17bd-4b36-9dbb-aa0a50a21013 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding LLaMA: Open and Efficient Foundation Language Models
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 156dff87-32bd-4c4b-8dfa-5b636ff75aab · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Unsplash Dataset
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b3ce1c53-7158-47fc-8175-0cead03127b1 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Gomez, Lukasz Kaiser, and Illia Polosukhin
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7a9bf386-814f-4c73-8d83-2b4bf614e94c · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 387791ba-366c-465e-bc13-02b5434d7dd9 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0eaa8d3-4c83-47aa-8133-c6cca7ca74b2 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding CogVLM: Visual Expert for Pretrained Language Models
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4869e8a8-b6a5-4e97-a538-d665ac4851d4 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding The all-seeing project: Towards panoptic visual recognition and understanding of the open world
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a2ae4707-6c25-409e-9d45-af722ddd40c2 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Emu3: Next-Token Prediction is All You Need
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65f2ae82-5b65-47ab-aca6-7a247cb6a623 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 147f3c0c-fba2-47d6-abdf-32112d9451b2 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding NExT-GPT: Any-to-Any Multimodal LLM
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ae3b315-0142-423a-aba0-ef07ecaded38 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9b940be-4569-464c-bd13-4b9f31ba7cea · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding OmniGen: Unified Image Generation
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccf93bb2-e6cb-4609-9403-c9770a4acae3 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0048fa3f-a283-4fc2-a2ce-afe0b736b402 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Raphael: Text-to- image generation via large mixture of diffusion paths
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 81d84120-61a7-4c2d-9495-3b0b83f072bf · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Qwen2 Technical Report
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ab5983e-50a0-43fa-8299-6bbcaed61542 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8538b517-b6c4-42c0-9ae9-3b67a9b93a03 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a03121b6-dd7a-4630-b6a9-73e3fb28ea33 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9112f0d1-801f-416b-ab8e-2b4b4d0bad98 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Megabyte: Predict- ing million-byte sequences with multiscale transformers
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8e052247-65f5-4fff-be4d-7aac0e180649 · outbound
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 103
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c94749cc-882d-4423-99fc-664959a4ca99 · inbound
Liquid: Language Models are Scalable and Unified Multi-modal Generators SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 853fab78-30c7-40f5-af36-9aec3406c865 · inbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bec555d-1264-4718-b863-d5dba6560ffc · inbound
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding
Reference 324
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7bfb031-8595-4721-b910-a1f59cedd831 · inbound
WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d4fa1d2f-af8a-4a60-aea4-08ba69542773 · inbound
UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e06dda0f-75f6-4bd1-8c8d-770a35ed0620 · inbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c8e9261-e7f5-452b-b842-2ef5ab478aef · inbound
Show-o2: Improved Native Unified Multimodal Models SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 718b4b77-42d6-4f16-baa8-b702afb4b3c8 · inbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1547e029-5114-4d1b-8b0e-5954bf0be3ca · inbound
Discrimination Is Generation: Unifying Ranking and Retrieval from a Tokenizer Perspective SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.