Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T12:37:38.651307Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 10 inbound Pith citation observations for arXiv:2411.17762.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T12:37:38.651307Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:09:11.003848Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-22T23:52:16.787762Z
85 of 85 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8507a12f-5c66-4239-89cf-7aa50e5e1745 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Flamingo: a visual language model for few-shot learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f2835c2-94b2-4ee0-ab42-298aa5981538 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44f81dd7-ef0e-4cdd-be92-ace2a921a6d1 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Qwen2.5-VL Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 181c9d4c-55e4-4cdc-8f93-1ab384f99582 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Improving image gener- ation with better captions, 2023
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8188ff07-bcc3-468b-84a8-b273b9f110a8 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding PaliGemma: A versatile 3B VLM for transfer
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e8cdbab-6cd0-4e2c-a1bc-87123ab9d4ef · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Coyo-700m: Image-text pair dataset
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0c30bf2-068c-4040-86d3-624da3fda694 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Maskgit: Masked generative image transformer
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46fac681-e18d-4352-92cf-20f8622924e7 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Conceptual 12m: Pushing web-scale image-text pre- training to recognize long-tail visual concepts
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ff15864-576e-448f-ae88-1bcfc6b4608f · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Pixart-α: Fast training of dif- fusion transformer for photorealistic text-to-image synthesis,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7321540-fc52-4c0b-ba29-49db3a5fbe34 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0755b7e6-9260-4b5e-89ef-dad026e78597 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a23dbd96-a73b-4706-940e-b375fe56e811 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding A single transformer for scalable vision-language modeling
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation be27bda7-e457-42a3-a31d-906c4220e703 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86e29768-1cb0-49c5-9148-e8c2709fb709 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Instructblip: Towards general-purpose vision-language models with instruction tuning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1cc40e7b-e48a-4215-9dd6-e63e13be356f · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Imagenet: A large-scale hierarchical image database
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8269385-5186-4029-8240-d1db20093b41 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Unveiling encoder-free vision-language models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c5cbe90c-0b31-4a3f-8209-6196182da9c5 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding DreamLLM: Synergistic multimodal com- prehension and creation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation edf6367b-6928-470e-b0fa-ba984f0b2371 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Vlmevalkit: An open- source toolkit for evaluating large multi-modality models,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a774c49e-351f-40f0-8295-e88260ba7339 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Taming transformers for high-resolution image synthesis
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation de316236-edad-4bf0-86f1-c444ee35cb11 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Planting a SEED of Vision in Large Language Model
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 050a9970-6069-4910-b918-cb01401a1834 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Making LLaMA SEE and Draw with SEED Tokenizer
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf65880e-e82c-49fe-aec9-8b456fab5f61 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a112c621-df74-473f-b192-e6cf7232c89c · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Geneval: an object-focused framework for evaluating text- to-image alignment
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation fac078f3-e6fa-49ce-a13d-16264a3e9ba9 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Gans trained by a two time-scale update rule converge to a local nash equilib- 9 rium
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7db6e583-f438-4db4-8bd1-775cdb21f8f7 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 217f4771-531d-4c0f-99fe-5513bdc0e364 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Image-to-image translation with conditional adver- sarial networks
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d475cae5-3ffe-47f8-80f8-e0a0bdfe072b · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Mixtral of Experts
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e057b765-c471-4fe0-aa22-12fda3024726 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Video-lavit: Unified video- language pre-training with decoupled visual-motional tok- enization
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 790dfede-b3b5-476c-b0a8-bedcb1b2ef77 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Unified language-vision pre- training in llm with dynamic discrete visual tokenization
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6772638-6e3c-4429-af9d-54e4a4c5b509 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding A diagram is worth a dozen images
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9a99ed6-b38b-42d5-a69a-f6c083a85c77 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Autoregressive image generation using residual quantization
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f0fe1532-543f-4944-ad45-82718fc6ed8f · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1eacb37a-ecee-4c44-a67b-8072f8f2ec02 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding LLaVA-OneVision: Easy Visual Task Transfer
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dfdeb88-cdc4-4ef8-8970-8575ea2e15f0 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0311b86d-9704-4735-a305-1c20988d5530 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Synergen-vl: Towards synergistic image understanding and generation with vision experts and token folding
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9b2a2e49-95c0-4620-8352-fc5b7f1b97be · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d4a519d-d3d2-4184-bcae-c643493ecc3c · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Vila: On pre-training for vi- sual language models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de4edad4-a053-4b22-851c-d703198a7bfa · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Visual instruction tuning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 938db05a-41dc-4f17-8d7e-64589cf4e74c · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Improved baselines with visual instruction tuning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9004d45e-0226-4e01-9044-520c3cfd5552 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 944bd66b-aba3-42a3-b092-2efade876215 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Visual instruction tuning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8594f6a-d89a-420f-b453-2f476b68f1ed · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding World Model on Million-Length Video And Language With Blockwise RingAttention
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02f5b4a2-410a-4108-8c2f-d0e905b784a0 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Mmbench: Is your multi-modal model an all-around player? In Computer Vi- sion – ECCV 2024 , pages 216–233, Cham, 2025
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 0986c505-5ecd-4108-82a4-fff2f474cff2 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation df18e532-8713-45bf-80fa-fe9bf69b021c · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 81a69498-4e29-4f7f-9aec-e49a6e501431 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3b5dee9-b53a-41b6-abf1-3ecd4e22ee1e · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Mono-internvl: Pushing the boundaries of monolithic multimodal large lan- guage models with endogenous visual pre-training
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c1ddd618-72a0-467c-8869-84e6ff4c89e5 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d1eac30-e13e-4032-b697-bb9eacb1b3b3 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c2d4188-3e4a-4e14-a5b1-040d9fe4f0b6 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 928a4784-305b-4549-80a3-935d52c53ffc · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding SDXL: Improving latent diffusion models for high-resolution image synthesis
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9eae1cb8-a1ef-4948-a502-f8e92d36566b · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7127dd30-8e47-4d6a-9fbf-e28497e8f858 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Learning transferable visual models from natural language supervi- sion
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e1cae9b-0dc3-4351-85a2-1e5914cc7257 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Hierarchical Text-Conditional Image Generation with CLIP Latents
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 327bd742-16c7-438d-8cd0-2bab38c88a84 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding High-Resolution Image Synthesis with Latent Diffusion Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7971c1ec-6ff4-482b-bd02-244292c86b5a · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding High-resolution image synthesis with latent diffusion models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 643dccb0-703a-4256-9e77-09adc72a5e99 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Towards vqa models that can read
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7a29f674-6f75-4afb-99a1-785b8e20a81a · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d07b46b5-dd7a-494b-944e-ad869bbacdfa · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Emu: Generative pretraining in multimodality
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e390ce0e-62de-4df9-b20d-8e9a25512d8d · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Generative multimodal mod- els are in-context learners
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 17afdea5-2751-4b84-993c-2bf661cd1cea · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a795a371-5486-4f6d-b69a-2719d19c29d6 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Qwen2.5: A party of foundation models, 2024
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 125df43d-e16a-428d-9d30-a3fc7c6dd891 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Visual autoregressive modeling: Scalable image generation via next-scale prediction
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b9062860-e059-4d24-925b-d323978f88f4 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbf62361-0b2f-4888-b8b1-dd486c432920 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1adda57b-1123-4341-b5fb-2cb8baeee63f · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Neural discrete representation learning
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 64c8d11b-0dcf-4452-a12e-3cb9f04259a4 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83738f8f-fc14-4959-9974-06ca2a0890a9 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25151ae7-8a02-4bc5-b9d4-1688a741b03b · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Emu3: Next-Token Prediction is All You Need
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55987781-c816-4a6f-b6c3-ddc7535f1212 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2e8fbbb-5d8b-4ae4-b676-f7759ffee295 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 811fa692-3184-4ab3-837e-0366cc8f571e · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5530d9a1-8ccf-42e7-a1b0-044c693ce596 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4df3d05-64ac-476c-a2fe-5250e3d50689 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Baichuan 2: Open Large-scale Language Models
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f176319-5c37-4baf-88d2-6caa060fe8b3 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Qwen2 Technical Report
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 856dc8cb-2354-46ec-8969-d882413e17cc · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Yi: Open Foundation Models by 01.AI
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf8b35a9-6665-4d69-8d61-24d12580a170 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Vector-quantized Image Modeling with Improved VQGAN
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4ac04c4-cfc4-48e0-b0ae-01e2ee87f543 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81882873-12df-4f4e-b209-06307429a998 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be546edd-4fb1-4cf5-94db-7263ef56d2f7 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Sigmoid loss for language image pre-training
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ed30e047-a24b-4686-873d-3621c9257364 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding AnyGPT: Unified multimodal LLM with discrete sequence modeling
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d0396117-f4f7-40d2-9e9e-d1c8ee1bedca · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding The unreasonable effectiveness of deep features as a perceptual metric
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fef925f1-c1d6-4d07-8ddb-2f522cf4f1ff · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Transfusion: Pre- dict the next token and diffuse images with one multi- modal model
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 6507bd3b-9380-4942-8210-a501790590e2 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding The results show that our model exhibits better perfor- mance than other unified models such as Chameleon [61] and SEED-X [22]
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d97facfc-3ef9-4483-b47f-dc610b56de74 · outbound
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Unresolved cited work
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c23031d0-36ec-4096-8cf9-467741caeef5 · inbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91c82944-264e-4de9-a118-6323c34ab188 · inbound
DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b28a401b-fad3-4f8b-a797-6119d726a32e · inbound
TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9c097b8-2027-4c6b-aa58-b6df8e08860f · inbound
Emerging Properties in Unified Multimodal Pretraining MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d5249f30-1d36-4c24-8522-d533979582ff · inbound
Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 48b86967-00b7-4792-bd88-fd439dc7beef · inbound
Show-o2: Improved Native Unified Multimodal Models MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding
Reference 129
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation eca16e4f-e042-4a5e-9470-7a5a5991c70b · inbound
UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bac26ac1-be08-415c-9dd1-0396a829c3e6 · inbound
InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f2053d36-e18c-4b43-9566-42a2ccfc4c53 · inbound
ChatUMM: Robust Context Tracking for Conversational Interleaved Generation MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30e90ac8-9442-4990-8e84-20318512ac85 · inbound
Twins: Learn to Predict Unified Representations with Focal Loss MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.