Pith. sign in

Paper Citation Record · LEDGER

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding

As of 20 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 10 inbound Pith citation observations for arXiv:2411.17762.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.17762 v4

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:37:38.651307Z

measured 95 of 95 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:09:11.003848Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T23:52:16.787762Z

Reference resolution

85 of 85 outbound references displayed

  • verified exact0
  • verified fuzzy26
  • unresolved59
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8507a12f-5c66-4239-89cf-7aa50e5e1745 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.279275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.279275Z digest=sha256:4ca8af5e6b05c0bff4302b761a794500ab0de63e9d33eb193e9e7f67fd842319

Observation 6f2835c2-94b2-4ee0-ab42-298aa5981538 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.284316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.284316Z digest=sha256:82df894c45243580ba65dd4059c854646f80b63dc6f9c23a85b7e97b3987892e

Observation 44f81dd7-ef0e-4cdd-be92-ace2a921a6d1 · outbound

This paper cites Qwen2.5-VL Technical Report.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.289200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.289200Z digest=sha256:1ec7beb38b99c73b1c9a9ea2347c1abe6fea5983a7cf0303eb9ec1a4f4a6adc3

Observation 181c9d4c-55e4-4cdc-8f93-1ab384f99582 · outbound

This paper cites Improving image gener- ation with better captions, 2023.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Improving image gener- ation with better captions, 2023

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.294238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.294238Z digest=sha256:36216e6e46af121446c72ecb692201f139c02c3a993e8052498111766e35ec86

Observation 8188ff07-bcc3-468b-84a8-b273b9f110a8 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding PaliGemma: A versatile 3B VLM for transfer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.298315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.298315Z digest=sha256:90b82d8b3cdffc6206160af409ca3d557baa204fe942be049d2617c7c2112973

Observation 4e8cdbab-6cd0-4e2c-a1bc-87123ab9d4ef · outbound

This paper cites Coyo-700m: Image-text pair dataset.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Coyo-700m: Image-text pair dataset

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.302967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.302967Z digest=sha256:4357b9bf87890eeba660f9cb56e6e24444b89f82a8bbcfad7aaae6bacdef8db5

Observation e0c30bf2-068c-4040-86d3-624da3fda694 · outbound

This paper cites Maskgit: Masked generative image transformer.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Maskgit: Masked generative image transformer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.307157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.307157Z digest=sha256:0444b33bb6ca2116ec6be9135448bfa0e95ab57925a9147a04f4e4ae0a018d30

Observation 46fac681-e18d-4352-92cf-20f8622924e7 · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre- training to recognize long-tail visual concepts.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Conceptual 12m: Pushing web-scale image-text pre- training to recognize long-tail visual concepts

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.311505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.311505Z digest=sha256:b6ddf0dab705a3e343f51fb77570135ed39ce551125c0bf5a689c594f1cf6d8e

Observation 1ff15864-576e-448f-ae88-1bcfc6b4608f · outbound

This paper cites Pixart-α: Fast training of dif- fusion transformer for photorealistic text-to-image synthesis,.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Pixart-α: Fast training of dif- fusion transformer for photorealistic text-to-image synthesis,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.315696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.315696Z digest=sha256:0c7c063bbef6a0a97877b56d3b1dfa70552b6496192b78c7e535db57511661ff

Observation e7321540-fc52-4c0b-ba29-49db3a5fbe34 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.319953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.319953Z digest=sha256:052141c1c3f2c2a4f4da2bf4e2e9994075a702ad3fa172e00fea80acc5eb6189

Observation 0755b7e6-9260-4b5e-89ef-dad026e78597 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.324818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.324818Z digest=sha256:f6b48e85a436fef543513059e702fba4b92d5cd3c759d1ba606b7a9311446b2d

Observation a23dbd96-a73b-4706-940e-b375fe56e811 · outbound

This paper cites A single transformer for scalable vision-language modeling.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding A single transformer for scalable vision-language modeling

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:37:39.700185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:37:38.329088Z digest=sha256:927bfe57d4bf2b5eb5859e0b54d126313a44dc1bbd37e85b920cabed6c7e4a09

Observation be27bda7-e457-42a3-a31d-906c4220e703 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.333478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.333478Z digest=sha256:98db8263b422f3ecc409888483bc3fe6c0e80a93ce8b4b8bd451d8f482c44f74

Observation 86e29768-1cb0-49c5-9148-e8c2709fb709 · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Instructblip: Towards general-purpose vision-language models with instruction tuning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:37:39.675776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:37:38.337771Z digest=sha256:8c0f6a84ef22d5a9c4c7c45e946f370a788be22cf5da2a9a15bb7856544b8f32

Observation 1cc40e7b-e48a-4215-9dd6-e63e13be356f · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Imagenet: A large-scale hierarchical image database

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.342024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.342024Z digest=sha256:570d3667397a5af44f93d8434aaf5cd5deea5cb995c48cd31eb48db66a5f001c

Observation b8269385-5186-4029-8240-d1db20093b41 · outbound

This paper cites Unveiling encoder-free vision-language models.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Unveiling encoder-free vision-language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:37:39.652584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:37:38.345752Z digest=sha256:505ce686de79c4a60a5b61674fa6b9edf7c70e120e317a3e1e80280f04be3c9d

Observation c5cbe90c-0b31-4a3f-8209-6196182da9c5 · outbound

This paper cites DreamLLM: Synergistic multimodal com- prehension and creation.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding DreamLLM: Synergistic multimodal com- prehension and creation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:37:39.639085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:37:38.349828Z digest=sha256:43be438c362d8846025cb72ca2a64aa9d14332d3bfb900e1d88d92205341424a

Observation edf6367b-6928-470e-b0fa-ba984f0b2371 · outbound

This paper cites Vlmevalkit: An open- source toolkit for evaluating large multi-modality models,.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Vlmevalkit: An open- source toolkit for evaluating large multi-modality models,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.353767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.353767Z digest=sha256:4a41c26edde710b2a5bbf59de07b33fd9c1cd628cc7601258fe340f8d87bbcc2

Observation a774c49e-351f-40f0-8295-e88260ba7339 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Taming transformers for high-resolution image synthesis

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:37:39.616979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:37:38.358590Z digest=sha256:3ef5f955254151c1d380aa361f2bf888274308230d32ac1f49d0ee4904a23194

Observation de316236-edad-4bf0-86f1-c444ee35cb11 · outbound

This paper cites Planting a SEED of Vision in Large Language Model.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Planting a SEED of Vision in Large Language Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.362436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.362436Z digest=sha256:dc7b58381b3296bde8abaa091370817f6bb085ada271097b326065a88ac4a90e

Observation 050a9970-6069-4910-b918-cb01401a1834 · outbound

This paper cites Making LLaMA SEE and Draw with SEED Tokenizer.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Making LLaMA SEE and Draw with SEED Tokenizer

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.366549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.366549Z digest=sha256:cd67050b77db30b667a0e8e3c8b04d6f65ed4fd97fa50146ae47fe18087ebeed

Observation cf65880e-e82c-49fe-aec9-8b456fab5f61 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.371001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.371001Z digest=sha256:8e8278fd403bc7488fcf53c8af69d1eb68639c2c234cc979f72ebc39214bba21

Observation a112c621-df74-473f-b192-e6cf7232c89c · outbound

This paper cites Geneval: an object-focused framework for evaluating text- to-image alignment.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Geneval: an object-focused framework for evaluating text- to-image alignment

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:37:39.601938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:37:38.375023Z digest=sha256:7facdd98bc4d81e4e2635c1e9784967a2863a3aedf477eee0c229ac7481fecd3

Observation fac078f3-e6fa-49ce-a13d-16264a3e9ba9 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- 9 rium.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Gans trained by a two time-scale update rule converge to a local nash equilib- 9 rium

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:37:39.586564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:37:38.378895Z digest=sha256:b097edf4e2380eed06610ffe6b758829f32731abd6e3245185259d09dbab8791

Observation 7db6e583-f438-4db4-8bd1-775cdb21f8f7 · outbound

This paper cites ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.382700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.382700Z digest=sha256:736c76b9aaf06afc94c00f1cc81026582f4535097092f41b959b755cd3b82f41

Observation 217f4771-531d-4c0f-99fe-5513bdc0e364 · outbound

This paper cites Image-to-image translation with conditional adver- sarial networks.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Image-to-image translation with conditional adver- sarial networks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.387207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.387207Z digest=sha256:2de6738da2eff8f7dee5729710633a0bc626b725eb7815d047557b86dfdc6fb3

Observation d475cae5-3ffe-47f8-80f8-e0a0bdfe072b · outbound

This paper cites Mixtral of Experts.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Mixtral of Experts

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.391634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.391634Z digest=sha256:33f8bf24779345ccdc5a7c379e332fec77edfe9f93eefbdee5b6139064f23a82

Observation e057b765-c471-4fe0-aa22-12fda3024726 · outbound

This paper cites Video-lavit: Unified video- language pre-training with decoupled visual-motional tok- enization.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Video-lavit: Unified video- language pre-training with decoupled visual-motional tok- enization

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:37:39.563958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:37:38.396086Z digest=sha256:d2ccfd37cdea9a6f1c8e91393179332d8d01320074d7f35be7fb791f74edc7fc

Observation 790dfede-b3b5-476c-b0a8-bedcb1b2ef77 · outbound

This paper cites Unified language-vision pre- training in llm with dynamic discrete visual tokenization.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Unified language-vision pre- training in llm with dynamic discrete visual tokenization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.400402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.400402Z digest=sha256:8061d3e93e2c4a846f1abe1f32eed3691073e6f9abd33c95c8936aa63bb2e4d5

Observation d6772638-6e3c-4429-af9d-54e4a4c5b509 · outbound

This paper cites A diagram is worth a dozen images.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding A diagram is worth a dozen images

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.404909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.404909Z digest=sha256:ede648df699abcf057ed72ed2f495baf18c0183001d23be4b4e5360f8a7034cf

Observation b9a99ed6-b38b-42d5-a69a-f6c083a85c77 · outbound

This paper cites Autoregressive image generation using residual quantization.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Autoregressive image generation using residual quantization

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:37:39.533759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:37:38.409135Z digest=sha256:51123b0f1422e86d93940e02952025ea91796212c12fff2c4a6bdbfc29515b0c

Observation f0fe1532-543f-4944-ad45-82718fc6ed8f · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.412969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.412969Z digest=sha256:36fac50f3a6816ee8c5759e414144200f154ffe17aefecc6b54be10413087ef1

Observation 1eacb37a-ecee-4c44-a67b-8072f8f2ec02 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding LLaVA-OneVision: Easy Visual Task Transfer

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.417137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.417137Z digest=sha256:597d5657966263ea3b02feac2e747f573d621874db21eb10a0f6b4a65f2c7c87

Observation 2dfdeb88-cdc4-4ef8-8970-8575ea2e15f0 · outbound

This paper cites Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.421542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.421542Z digest=sha256:04f11e62a469de3aef9079c444ce688d52808dbe24dbfd412b61624daa7e6eea

Observation 0311b86d-9704-4735-a305-1c20988d5530 · outbound

This paper cites Synergen-vl: Towards synergistic image understanding and generation with vision experts and token folding.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Synergen-vl: Towards synergistic image understanding and generation with vision experts and token folding

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:37:39.519743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:37:38.426056Z digest=sha256:9b1e487aa4e0bc80b2d3affb27bffbd26991240f282e659a4f7b75039418751f

Observation 9b2a2e49-95c0-4620-8352-fc5b7f1b97be · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.430261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.430261Z digest=sha256:3b3e616319d2945b948444b52bef8ac4df9eaba550b35f9efbdd44827ec1e255

Observation 4d4a519d-d3d2-4184-bcae-c643493ecc3c · outbound

This paper cites Vila: On pre-training for vi- sual language models.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Vila: On pre-training for vi- sual language models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.434945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.434945Z digest=sha256:44e020e94b43f21ab80563b34bd5d8e59ae505624d88193a98e78a59b3fce0ca

Observation de4edad4-a053-4b22-851c-d703198a7bfa · outbound

This paper cites Visual instruction tuning.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Visual instruction tuning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.438834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.438834Z digest=sha256:0f232f2810088925e6b59cdf94f6c31f4b8f3f699e8e3595f7857639a2a8f552

Observation 938db05a-41dc-4f17-8d7e-64589cf4e74c · outbound

This paper cites Improved baselines with visual instruction tuning.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Improved baselines with visual instruction tuning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:37:39.477251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:37:38.442431Z digest=sha256:300d7e8ae98f89f82afd2ddae40c2f1c51e0f49a26b8361abbebc893d353af45

Observation 9004d45e-0226-4e01-9044-520c3cfd5552 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.446553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.446553Z digest=sha256:0bea0d9776a7bd2a9d94bb078d5b9daf9c35b5e693d8ac275e4c1f8207f38de4

Observation 944bd66b-aba3-42a3-b092-2efade876215 · outbound

This paper cites Visual instruction tuning.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Visual instruction tuning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.450970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.450970Z digest=sha256:af8b2745189fc64e55c8c6dbe6168315f311a8567b21be909f9c25d84164b40f

Observation a8594f6a-d89a-420f-b453-2f476b68f1ed · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.455270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.455270Z digest=sha256:9f8a31b881b5ac957b7dfcd49289575dd2884f1e6fcb9cfef66d0ba1b4c1252a

Observation 02f5b4a2-410a-4108-8c2f-d0e905b784a0 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In Computer Vi- sion – ECCV 2024 , pages 216–233, Cham, 2025.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Mmbench: Is your multi-modal model an all-around player? In Computer Vi- sion – ECCV 2024 , pages 216–233, Cham, 2025

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:37:39.447694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:37:38.459700Z digest=sha256:e6b4274019f42bc9a07f8ade5578ab45da3db4f70dd3468979bd71f3b5249beb

Observation 0986c505-5ecd-4108-82a4-fff2f474cff2 · outbound

This paper cites Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:37:39.434319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:37:38.464944Z digest=sha256:4387b15377712f29698172bf5b718d26abba9e13b99f0c02a8ec9cd08f0b818f

Observation df18e532-8713-45bf-80fa-fe9bf69b021c · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:37:39.419541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:37:38.469364Z digest=sha256:74dc874b8135830eac84a48f322b010b238eb8e8a7f2d71fe4fef328e523a065

Observation 81a69498-4e29-4f7f-9aec-e49a6e501431 · outbound

This paper cites Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.473693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.473693Z digest=sha256:61bdc1b0e592ddd72aa9bb5dbfd4bd8bdb5b836d35734dde366f6a5c5b76f69b

Observation d3b5dee9-b53a-41b6-abf1-3ecd4e22ee1e · outbound

This paper cites Mono-internvl: Pushing the boundaries of monolithic multimodal large lan- guage models with endogenous visual pre-training.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Mono-internvl: Pushing the boundaries of monolithic multimodal large lan- guage models with endogenous visual pre-training

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:37:39.398240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:37:38.478374Z digest=sha256:0ff2553f1605a4e01ad5bc3e3b9484a4f2d312e68d9df85900b7c122c518dc85

Observation c1ddd618-72a0-467c-8869-84e6ff4c89e5 · outbound

This paper cites UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.482118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.482118Z digest=sha256:84eeb01cb9af0059d7d1f05688d5d6939bd2b016748fe34d4b9109ac0613478a

Observation 0d1eac30-e13e-4032-b697-bb9eacb1b3b3 · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.486456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.486456Z digest=sha256:7e8150797d0e61980adf848c82e4d04b1806749d9eb2a1448b538812afb42051

Observation 2c2d4188-3e4a-4e14-a5b1-040d9fe4f0b6 · outbound

This paper cites BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.491552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.491552Z digest=sha256:1852f732b6e55c1375e90bdad9f614b9e47e244a1521c4974936ec5d698693a6

Observation 928a4784-305b-4549-80a3-935d52c53ffc · outbound

This paper cites SDXL: Improving latent diffusion models for high-resolution image synthesis.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding SDXL: Improving latent diffusion models for high-resolution image synthesis

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:37:39.383464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:37:38.496124Z digest=sha256:3bef8259b0112997b2946cf6705836e86c3df01e39d87eacc33760a6595dc644

Observation 9eae1cb8-a1ef-4948-a502-f8e92d36566b · outbound

This paper cites TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.500293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.500293Z digest=sha256:72f4756e48c86d2a3ba4e4fe0971ab48fe4d9748892e13859f8bef023aa2fc38

Observation 7127dd30-8e47-4d6a-9fbf-e28497e8f858 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Learning transferable visual models from natural language supervi- sion

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.504735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.504735Z digest=sha256:19b4ed1c8f130c795f86cfecd89e5a1ae25bd7e749830c1fc2dd06f4b6797f67

Observation 7e1cae9b-0dc3-4351-85a2-1e5914cc7257 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.508880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.508880Z digest=sha256:bd39f9caf23cbe1fef56d82c1cbbc93bbda3126b299fc11a740e0453b5696231

Observation 327bd742-16c7-438d-8cd0-2bab38c88a84 · outbound

This paper cites High-Resolution Image Synthesis with Latent Diffusion Models.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding High-Resolution Image Synthesis with Latent Diffusion Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.512915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.512915Z digest=sha256:c8218c3184d5bb962b230b3c7fffa8ba8dd111c5715149561713ee8d56ddeb3e

Observation 7971c1ec-6ff4-482b-bd02-244292c86b5a · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding High-resolution image synthesis with latent diffusion models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.522175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.522175Z digest=sha256:240d5f7abae804270f11c46cc380d6238abb5728802762b88796d5023e435cd4

Observation 643dccb0-703a-4256-9e77-09adc72a5e99 · outbound

This paper cites Towards vqa models that can read.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Towards vqa models that can read

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:37:39.331482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:37:38.526601Z digest=sha256:c3de158a65b291f0860cf74acd692258f00c9444dc5b1bb9ca0700d8ca29abe4

Observation 7a29f674-6f75-4afb-99a1-785b8e20a81a · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.530485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.530485Z digest=sha256:d010f85e20f07db9c95aac5c2ab61da8b9d6dfe0d206ed8c826022a8fe2692c7

Observation d07b46b5-dd7a-494b-944e-ad869bbacdfa · outbound

This paper cites Emu: Generative pretraining in multimodality.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Emu: Generative pretraining in multimodality

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:37:39.316561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:37:38.535006Z digest=sha256:8ab11a8a3d97379cef3f657640dbcad1ab51987855488f56bcb00b7bd40b632e

Observation e390ce0e-62de-4df9-b20d-8e9a25512d8d · outbound

This paper cites Generative multimodal mod- els are in-context learners.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Generative multimodal mod- els are in-context learners

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:37:39.302125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:37:38.539406Z digest=sha256:3636e99c0955af2ded8c12b7ee2e2c155c5b8c11c712bc9841b1c6bdd98dca74

Observation 17afdea5-2751-4b84-993c-2bf661cd1cea · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.543832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.543832Z digest=sha256:0c01f31964bacd320e11496357cec27608e93eba43e27b43b0ad168102d0dfc0

Observation a795a371-5486-4f6d-b69a-2719d19c29d6 · outbound

This paper cites Qwen2.5: A party of foundation models, 2024.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Qwen2.5: A party of foundation models, 2024

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:37:39.287793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:37:38.548050Z digest=sha256:58cfd4a1639444f716be08703f9ad87f9fc653c655dee54df080a9ebf2821ece

Observation 125df43d-e16a-428d-9d30-a3fc7c6dd891 · outbound

This paper cites Visual autoregressive modeling: Scalable image generation via next-scale prediction.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Visual autoregressive modeling: Scalable image generation via next-scale prediction

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:37:39.273477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:37:38.551822Z digest=sha256:af7ffff023fb48fc8425c0506da4bc87d382d7d1a9b22f2fb948a0da9cddcb5f

Observation b9062860-e059-4d24-925b-d323978f88f4 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.555571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.555571Z digest=sha256:e663b806d2e6703455d7b92e77f83505797f2c99f08ec9ec2b847ce763733585

Observation bbf62361-0b2f-4888-b8b1-dd486c432920 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.560512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.560512Z digest=sha256:0db7a06ad251b2f13154732b7282f30b77f0c62a5b8184e41077c5159353af73

Observation 1adda57b-1123-4341-b5fb-2cb8baeee63f · outbound

This paper cites Neural discrete representation learning.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Neural discrete representation learning

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:37:39.259205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:37:38.564816Z digest=sha256:8f89b50116eba8bc82a9a137b0f8815b9f70b376c62da16dbd05ca8fd2d91128

Observation 64c8d11b-0dcf-4452-a12e-3cb9f04259a4 · outbound

This paper cites ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.568950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.568950Z digest=sha256:0aac9222797fbd59fdff9f2b801ee77f9f44b260e84ad7ea3b5aae03b0632147

Observation 83738f8f-fc14-4959-9974-06ca2a0890a9 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.573476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.573476Z digest=sha256:fe4242e6f0e8e0439d5a064b58196b31825e2f19b3614fb29b46b8100cfa52fb

Observation 25151ae7-8a02-4bc5-b9d4-1688a741b03b · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Emu3: Next-Token Prediction is All You Need

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.578138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.578138Z digest=sha256:fbbeef5ebf79313d2ca4e93f6ca28a4a72f1e9ca46747a078daa5bf438031f74

Observation 55987781-c816-4a6f-b6c3-ddc7535f1212 · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.583581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.583581Z digest=sha256:7a27444bb423a33aadf89efa04cd6daed7e8131b05795e0dfffca0404b2d92d4

Observation f2e8fbbb-5d8b-4ae4-b676-f7759ffee295 · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.588446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.588446Z digest=sha256:97da21d767fe7bbb1fc434e2bcf59c324fc258ca612532fc0e0f0a5b7819709e

Observation 811fa692-3184-4ab3-837e-0366cc8f571e · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.593731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.593731Z digest=sha256:2339a24e157a97744ce04ff9324e563800cc32b4a78ccd3c8da9692cb5a3d119

Observation 5530d9a1-8ccf-42e7-a1b0-044c693ce596 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.598924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.598924Z digest=sha256:cd355c849a0ea38a3bfa4b8c696096e8c3d27185116344a97dcd409a472b0c3b

Observation a4df3d05-64ac-476c-a2fe-5250e3d50689 · outbound

This paper cites Baichuan 2: Open Large-scale Language Models.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Baichuan 2: Open Large-scale Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.603113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.603113Z digest=sha256:4d699d50a4dba6bfed034604281c7154812382fd9ac57c5b7b32be205a79a897

Observation 0f176319-5c37-4baf-88d2-6caa060fe8b3 · outbound

This paper cites Qwen2 Technical Report.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Qwen2 Technical Report

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.612286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.612286Z digest=sha256:99c6fc5a073e85b60f409907ccae6fcc59da8979dd2115d898cc00d77ff576cb

Observation 856dc8cb-2354-46ec-8969-d882413e17cc · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Yi: Open Foundation Models by 01.AI

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.616627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.616627Z digest=sha256:050afd6ac1b382c5426892411f7e7131570fc2dbeb7ca2cf36d0c5024c1795f9

Observation bf8b35a9-6665-4d69-8d61-24d12580a170 · outbound

This paper cites Vector-quantized Image Modeling with Improved VQGAN.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Vector-quantized Image Modeling with Improved VQGAN

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.621145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.621145Z digest=sha256:af6eec6ee820939b3ec2e8d9043d94f86b1ef8bb7f9ea3eb64ff1470bb8e2e1e

Observation d4ac04c4-cfc4-48e0-b0ae-01e2ee87f543 · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.625660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.625660Z digest=sha256:b81699b3bd7577e244bf87e82d92f78a4322e837ef93ca43c022f61df84fdf6e

Observation 81882873-12df-4f4e-b209-06307429a998 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.629961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.629961Z digest=sha256:c39f623116d4c5342e3c958ba020c3346df543cd5c92373cfc7d34ce83ea0093

Observation be546edd-4fb1-4cf5-94db-7263ef56d2f7 · outbound

This paper cites Sigmoid loss for language image pre-training.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Sigmoid loss for language image pre-training

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:37:39.237302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:37:38.634171Z digest=sha256:574af80e4ac7ad1c03b0d04ac377a1fe6697383742c18070d5efb175db3f17bb

Observation ed30e047-a24b-4686-873d-3621c9257364 · outbound

This paper cites AnyGPT: Unified multimodal LLM with discrete sequence modeling.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding AnyGPT: Unified multimodal LLM with discrete sequence modeling

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:37:39.223312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:37:38.638029Z digest=sha256:ef0aa418da81e4ab4ed2e83b02aa7eac82c9f51c60d135bf39be5693fd14ce4f

Observation d0396117-f4f7-40d2-9e9e-d1c8ee1bedca · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual metric.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding The unreasonable effectiveness of deep features as a perceptual metric

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-12T12:37:38.642021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:37:38.642021Z digest=sha256:ccf3539238a59365b7c9b6b44f1e854b2d18e992ac84c7831bb8c537fedbc9cd

Observation fef925f1-c1d6-4d07-8ddb-2f522cf4f1ff · outbound

This paper cites Transfusion: Pre- dict the next token and diffuse images with one multi- modal model.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Transfusion: Pre- dict the next token and diffuse images with one multi- modal model

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:37:39.200574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:37:38.646433Z digest=sha256:365eb3c32dc29bd318f82b18c689fc0d53dae6740b5fdbd18d79954fc25d15e2

Observation 6507bd3b-9380-4942-8210-a501790590e2 · outbound

This paper cites The results show that our model exhibits better perfor- mance than other unified models such as Chameleon [61] and SEED-X [22].

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding The results show that our model exhibits better perfor- mance than other unified models such as Chameleon [61] and SEED-X [22]

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:37:39.185893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:37:38.651307Z digest=sha256:d83c0bbce394318a10a43c2cdfbe27f6ab03acd4b7bf21077b12caaa7a2ca655

Observation d97facfc-3ef9-4483-b47f-dc610b56de74 · outbound

This paper cites an unresolved cited work.

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Unresolved cited work

Reference 2022

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:37:39.353540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:37:38.517494Z digest=sha256:f9a6049fb9e8fa4a1cd362cae5470bb4279bb78b8b6abbf3dadcc3cf023883a9

Pith citing papers

Observation c23031d0-36ec-4096-8cf9-467741caeef5 · inbound

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models cites this paper.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.957533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.957533Z digest=sha256:5ad827238d46ce204a21309c64309a4eaa7c5fe79b3d0cea6a4d3fc2136d3eda

Observation 91c82944-264e-4de9-a118-6323c34ab188 · inbound

DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies cites this paper.

DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:52:16.792292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T23:51:43.934329Z digest=sha256:19dc16e272b4d5bc015c107d72415d55b928ef264565fe4bdd7c95f6a683e0ef

Observation b28a401b-fad3-4f8b-a797-6119d726a32e · inbound

TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation cites this paper.

TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-15T23:09:11.003848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:09:11.003848Z digest=sha256:289494ea07a2de4347fc0406a266c6b7b34b7eab99f9a5986201326ab3a93895

Observation b9c097b8-2027-4c6b-aa58-b6df8e08860f · inbound

Emerging Properties in Unified Multimodal Pretraining cites this paper.

Emerging Properties in Unified Multimodal Pretraining MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-10T16:23:42.072208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T16:23:41.854132Z digest=sha256:09c7d85fbb049f1299fe3902d7dcbe174bd95fd1595ea8ebd6237524edd1482b

Observation d5249f30-1d36-4c24-8522-d533979582ff · inbound

Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM cites this paper.

Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-22T02:10:56.214807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T02:06:35.204166Z digest=sha256:b697981d33a8d5afab1587660ca94912769a89b81bfa439136293642fa2783b5

Observation 48b86967-00b7-4792-bd88-fd439dc7beef · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding

Reference 129

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:51:15.866825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:e02cd6a00f1b2f7cf6cad7b2a4a549462f894ff93d0e4f58b52b83b47d5ebe56

Observation eca16e4f-e042-4a5e-9470-7a5a5991c70b · inbound

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation cites this paper.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.883147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.883147Z digest=sha256:06b3ffeaf957a82c74674714f0861904f106bba9b4a7644b21978704ccd187cc

Observation bac26ac1-be08-415c-9dd1-0396a829c3e6 · inbound

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs cites this paper.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:10:45.435592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:5d1f740c99c70ea52afa970e05c12354ce30525bca23a21063710f2ce5e25370

Observation f2053d36-e18c-4b43-9566-42a2ccfc4c53 · inbound

ChatUMM: Robust Context Tracking for Conversational Interleaved Generation cites this paper.

ChatUMM: Robust Context Tracking for Conversational Interleaved Generation MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T03:57:32.313692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:57:32.313692Z digest=sha256:1e9e18f5143f73575ec8ad118be82d60f8efb729b87020fbcc71a5344927743a

Observation 30e90ac8-9442-4990-8e84-20318512ac85 · inbound

Twins: Learn to Predict Unified Representations with Focal Loss cites this paper.

Twins: Learn to Predict Unified Representations with Focal Loss MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-01T04:29:49.896752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:29:49.896752Z digest=sha256:12814708343b7fd741205a1f572988416397a983fcff0ef7b39696ef8b298656