Pith. sign in

Paper Citation Record · LEDGER

MMaDA: Multimodal Large Diffusion Language Models

As of 3 August 2026, this Paper Citation Record lists 100 of 126 outbound references and 69 inbound Pith citation observations for arXiv:2505.15809.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15809 v2

Coverage vector

measured 100 of 126 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-15T14:50:59.661153Z

measured 169 of 169 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 69 of 69 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T16:21:38.232907Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T03:07:51.413936Z

Reference resolution

100 of 126 outbound references displayed

  • verified exact40
  • verified fuzzy60
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fb949630-e33e-4d9c-b529-7fe530ed1b86 · outbound

This paper cites Improving language understanding by generative pre-training.

MMaDA: Multimodal Large Diffusion Language Models Improving language understanding by generative pre-training

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.044693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:5ea468a0370c5b41069100454e8120ff33c7fa3d0c981d56e0f0557ac8de872c

Observation a2fb3350-ce98-49dc-8bf6-8228dc4af0f7 · outbound

This paper cites Language models are few-shot learners.

MMaDA: Multimodal Large Diffusion Language Models Language models are few-shot learners

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.052136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:d6ae787d98c9b6efee736f050f170c12b7ee29a5da0a2b3e21de50332f703340

Observation 921f9849-59d5-4d60-a968-b371ad70fa52 · outbound

This paper cites OpenAI o1 System Card.

MMaDA: Multimodal Large Diffusion Language Models OpenAI o1 System Card

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:50:59.729150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:45175a0f7b8bbff7b5c39b73832337770ad20e7508d9580d9b98953e2bec2cf9

Observation 0d456aa7-05ec-41fe-a08f-a3a0500c832c · outbound

This paper cites VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation.

MMaDA: Multimodal Large Diffusion Language Models VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:50:59.848027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:fccc781dd0425864ea42fa2cb4211e103b951c187904e87047f49202e255aeea

Observation 24ae7dfa-e2f3-4a95-857e-278095069e6f · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

MMaDA: Multimodal Large Diffusion Language Models Emu: Generative Pretraining in Multimodality

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:22:11.462164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:c18e4b40e2c2b2821e91ad1c654737da3545beb7fdacb17e3309dc8e9784dc42

Observation ec11eb0b-6777-46b3-b61b-780cc315fee6 · outbound

This paper cites Generative Multimodal Models are In-Context Learners.

MMaDA: Multimodal Large Diffusion Language Models Generative Multimodal Models are In-Context Learners

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:50:59.902495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:46859ab1714b39c0a19d0a289346997ea2457ff527e8b7b860156382fcc77603

Observation d9716231-2902-45fa-84b5-fae5bfbace46 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

MMaDA: Multimodal Large Diffusion Language Models Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:50:59.718265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:da2fcc7c1970b25807a932a0380cf6cd7bf0f3bd5d333fe527638c41ef9b6a3c

Observation 0e418431-77d6-4234-87bb-f6c2b035c798 · outbound

This paper cites World model on million-length video and language with ringattention.

MMaDA: Multimodal Large Diffusion Language Models World model on million-length video and language with ringattention

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.056481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:a088947baea6c22b4603cb8dfce051662f5b2b58a17023c6389a8a828e7fc14f

Observation 571b9114-4fb5-4417-8ff5-6fef32f79f08 · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

MMaDA: Multimodal Large Diffusion Language Models VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T00:26:21.484895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:d500ed57ef2eea70a6a123a0c4e08964dcb1f4431d759bfe752d8e4473184605

Observation ec1089e9-f5f8-4811-9685-b0ca27da6215 · outbound

This paper cites Emu: Generative pretraining in multimodality.

MMaDA: Multimodal Large Diffusion Language Models Emu: Generative pretraining in multimodality

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.060611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:fc8e49ffe501ff4adb8f351d3193e3fef9ccffb1022cacbe97dedc002ea7e4ed

Observation 90dc9801-715f-437f-8f3a-fda1bd743430 · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

MMaDA: Multimodal Large Diffusion Language Models Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:09:16.660059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:680a8fbfc4bdf75e07e4326ae53852f24b83d87aacf49625991568c27579ba64

Observation b996e979-8360-48f4-b03c-9fd914e15219 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

MMaDA: Multimodal Large Diffusion Language Models Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:50:59.891401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:87dacae6b1c35da1ad63999529254838918c461218fb6c1662b5442c240f13d3

Observation 7c44082a-23b0-4d85-9b1e-dffd36223999 · outbound

This paper cites Gpt-4 technical report.

MMaDA: Multimodal Large Diffusion Language Models Gpt-4 technical report

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.063791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:d99f9ca48a000ce09be5f1b21a0b2684f28f46f63c98a8f3f6493cca097d284a

Observation 19b8fbdd-2b1d-4fe5-9a6b-71628249ff35 · outbound

This paper cites DreamLLM: Synergistic multimodal comprehension and creation.

MMaDA: Multimodal Large Diffusion Language Models DreamLLM: Synergistic multimodal comprehension and creation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.068371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:e9908bb06f0187621c527dfdad6f5f48e38df99a76a3c42ae76194ddabdd8030

Observation 6053d725-9f06-4d06-8744-80528a8cd981 · outbound

This paper cites Generating images with multimodal language models.

MMaDA: Multimodal Large Diffusion Language Models Generating images with multimodal language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.071892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:0ed11eafd654e9ebe9bcce400d5b2e2f749331054a5ff61caafe4f3bc042108c

Observation 1131beae-9237-465c-aa6b-c535e9745890 · outbound

This paper cites Llava-plus: Learning to use tools for creating multimodal agents.

MMaDA: Multimodal Large Diffusion Language Models Llava-plus: Learning to use tools for creating multimodal agents

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.074711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:fc3e5af6497bc657d2b6f3d0b5f08a8ffa25822c11a19675c43100f476fad8f8

Observation 9e61143c-38eb-40b6-b20d-e5576defed9b · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

MMaDA: Multimodal Large Diffusion Language Models SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:48:36.519178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:c6bc09c7621d070f9a2e2522ef897bc08266c78af4e27c7f893c07b3002e521e

Observation 0b0ba415-f8e8-4a23-bfa3-0bf9125946ef · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

MMaDA: Multimodal Large Diffusion Language Models Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:50:59.764951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:437888038deee69d19e164a6f0778991cc7a6df4a07c5981cb6627b60a56d495

Observation 867ee91a-bbfd-41da-8309-17a8fbfb801b · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

MMaDA: Multimodal Large Diffusion Language Models Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:50:59.778692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:49ebdbb45ceba1651387b4b632ad34f300ff6127ace49b738d9dfdad13729a20

Observation f335ff4b-d157-4cd3-8d00-4fcb2a3e26d2 · outbound

This paper cites Large Language Diffusion Models.

MMaDA: Multimodal Large Diffusion Language Models Large Language Diffusion Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:50:59.793008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:a929d09e00ff709dfa6ea4ee68cb509226fc77546617fa74cec0bfcb43d492ed

Observation 62406cb4-bb3f-4aa6-a51d-0720ff0926b3 · outbound

This paper cites Denoising diffusion probabilistic models.

MMaDA: Multimodal Large Diffusion Language Models Denoising diffusion probabilistic models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.078103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:2c474c620def0203435fcde3cf006cc9d5fd390ea9354c81c1745d8c035c28d8

Observation 47ddc277-904f-43b0-a2d7-987c74eeb0ae · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

MMaDA: Multimodal Large Diffusion Language Models Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:50:59.838917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:a2479f1584df8eb0a93b357108d6dc985a27640a7b54a9007c9c2b5065e28ccf

Observation 182cc9b2-c698-4796-a379-09b667703afc · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

MMaDA: Multimodal Large Diffusion Language Models Emu3: Next-Token Prediction is All You Need

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:50:59.843135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:b1102b5105c51cead3d87485be6feb2910cfe2aee899d485c8042c6808bb97fe

Observation 4731bda0-a302-4137-ad12-e883f6e02947 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837.

MMaDA: Multimodal Large Diffusion Language Models Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.081828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:5eda54c7941fdb034e26cf497ab0c2beed2d1fb08736cc3f2585d50f7f373482

Observation 702b57a8-b5f9-40ad-983f-3436b1ea454b · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

MMaDA: Multimodal Large Diffusion Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:50:59.874569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:d391678860eb63b6472e50ae2138c21a6677e46a1a9da3f2f842eec6af81a8d7

Observation d1574be5-1323-4b1a-9bb0-075f3eb242ba · outbound

This paper cites d1: Scaling reasoning in diffusion large language models via reinforcement learning.

MMaDA: Multimodal Large Diffusion Language Models d1: Scaling reasoning in diffusion large language models via reinforcement learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.085232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:65035c39df040877cb3e58875abc8c62bf7b07525e4ceef3ec745510880d21e4

Observation ed9aa9d7-c1bd-423b-9044-166a6e6b1be7 · outbound

This paper cites Xing, and Liang Lin.

MMaDA: Multimodal Large Diffusion Language Models Xing, and Liang Lin

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.088423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:7beccbb2c441b452bcd16052090598f20852d9ad208391393713095b84ee0c1a

Observation 1fe952ff-2cde-489a-a6b1-8c4983afaec8 · outbound

This paper cites Lawrence Zitnick, and Ross Girshick.

MMaDA: Multimodal Large Diffusion Language Models Lawrence Zitnick, and Ross Girshick

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.091767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:7218a4748c59f372f40f84c59c25e1ae9887dadb68a382ee5ee1abe32b1cca63

Observation 3e29b659-8c9e-4825-b0c3-48b94a10db48 · outbound

This paper cites Improved baselines with visual instruction tuning.

MMaDA: Multimodal Large Diffusion Language Models Improved baselines with visual instruction tuning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.095137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:587812e86ae925e37f00f4b68d3bd31b70393b7974745b624a45eb8d798c65be

Observation ff1def59-772d-428e-a4ef-0569566a1f0d · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.

MMaDA: Multimodal Large Diffusion Language Models Instructblip: Towards general-purpose vision-language models with instruction tuning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.098554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:93c88f8085a872078916c990bc00e22ec406014f9989e5a3b5ce439d58d2c4ed

Observation 759a2b7b-9112-4ba1-bd89-794b230386cf · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

MMaDA: Multimodal Large Diffusion Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:50:59.724596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:1b6d403d60c0db93ac392a14e2899e2a66dc44e996c9361a6916acb2edf96c2d

Observation 8a4a4225-c644-40ef-85da-e8588c7d0ea0 · outbound

This paper cites mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration.

MMaDA: Multimodal Large Diffusion Language Models mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.101889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:d8151a15ff31468fb60752d5ffd8a6e6fe3fd1caf93673f1edb9e2c71af3c568

Observation 51dbc47a-cd85-4cfb-9901-3c1590ec3fea · outbound

This paper cites Llava-phi: Efficient multi-modal assistant with small language model.

MMaDA: Multimodal Large Diffusion Language Models Llava-phi: Efficient multi-modal assistant with small language model

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.105582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:a25db4e816c079df9389c16e403230dde5eafc72d02dbf425e228f4a758ea4db

Observation 31acadb4-d211-49c0-8d7b-30780d80408c · outbound

This paper cites The refinedweb dataset for falcon LLM: outperforming curated corpora with web data only.

MMaDA: Multimodal Large Diffusion Language Models The refinedweb dataset for falcon LLM: outperforming curated corpora with web data only

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.109242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:3120e490d17e2079b1a655fe8c17573a0918929cc0205f4687491894875f8dbe

Observation 23f19fd7-322a-4a97-a093-5d810bbad181 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

MMaDA: Multimodal Large Diffusion Language Models Imagenet: A large-scale hierarchical image database

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.113396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:44553c10201f23da28868a1de8ce5b48ed64e6b5aa1c47f88910bc457ba8c1e2

Observation e838c433-1215-44f3-9ce2-9bf5ff23d73e · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

MMaDA: Multimodal Large Diffusion Language Models Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.117308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:f71be106365ab3116e5e13910cf1f93cc1eba412474f6bb2086d5bbd7c0978eb

Observation 722114fa-ddfd-415f-8ec7-eacb352f28a5 · outbound

This paper cites Segment anything.

MMaDA: Multimodal Large Diffusion Language Models Segment anything

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.120433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:f9e5b2672ebaf6984c80dcc8cb8d5ea285fd024d0ccfd11b9949bc164152e434

Observation f25b407c-9c2b-4b1b-8e29-f0758e096c08 · outbound

This paper cites laion-aesthetics-12m-umap.

MMaDA: Multimodal Large Diffusion Language Models laion-aesthetics-12m-umap

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.123336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:ecc78ebf3c33e936a93bdaebcf7a4ec2790105b176e914552ebf4a2e91d2c83c

Observation cda28788-7f4e-42f9-a41b-6a09ceb63cde · outbound

This paper cites Journeydb: A benchmark for generative image understanding.

MMaDA: Multimodal Large Diffusion Language Models Journeydb: A benchmark for generative image understanding

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.126203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:7d08291a9ea3495e4e510697d47cb2d2656db3229dcdcb906dbafa9ed1244e54

Observation 98246a7a-4802-440b-951e-9e615c656d70 · outbound

This paper cites Hashimoto.

MMaDA: Multimodal Large Diffusion Language Models Hashimoto

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.129053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:f564419e0a415d04ff213f758f8224e38617d1f7121dbc560d60dde316868b6a

Observation a0c2d6af-8636-467e-9d76-cdcf6fdaf682 · outbound

This paper cites ReasonFlux: Hierarchical LLM Reasoning via Scaling Thought Templates.

MMaDA: Multimodal Large Diffusion Language Models ReasonFlux: Hierarchical LLM Reasoning via Scaling Thought Templates

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:50:59.864337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:367aef3c4ee8da5cd46afc59fcecaaefb52be3adc3cad574b31b3e6ecfac6e8a

Observation 78a86927-f67a-4792-b7c6-3df24d26a94a · outbound

This paper cites Limo: Less is more for reasoning.

MMaDA: Multimodal Large Diffusion Language Models Limo: Less is more for reasoning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.132173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:d96e7713a19b205f01d61b3d6333139a4eb54326a2e93a7791104f9e3e08baae

Observation 4ef0f7f0-3045-41d0-87c7-051708ccf443 · outbound

This paper cites s1: Simple test-time scaling.

MMaDA: Multimodal Large Diffusion Language Models s1: Simple test-time scaling

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.135337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:5497a60de34a95ace5083b95e9275574506acabfd063cd5bd58449676e6aec40

Observation 212ea4a7-d9a5-4017-9faa-0ffc325155ef · outbound

This paper cites Open Thoughts.

MMaDA: Multimodal Large Diffusion Language Models Open Thoughts

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.138270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:72d3a53a3f36e5c7b5fc8d2294368517d203be11d8c3c635cf741734789fd85f

Observation f601dd0e-860b-49b8-9bc0-d5bc2e3c78c5 · outbound

This paper cites Acemath: Advancing frontier math reasoning with post-training and reward modeling.

MMaDA: Multimodal Large Diffusion Language Models Acemath: Advancing frontier math reasoning with post-training and reward modeling

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.141465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:17a39ca65fcca52333e3ea6a4d3f1eea8217f6bbd25035b16aec8565aa8035e2

Observation 5ef9c87f-ce9e-4e93-b53f-8e5202aede9b · outbound

This paper cites Lmm-r1: Empowering 3b lmms with strong reasoning abilities through two-stage rule-based rl.

MMaDA: Multimodal Large Diffusion Language Models Lmm-r1: Empowering 3b lmms with strong reasoning abilities through two-stage rule-based rl

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.144198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:250e9804302613781d285cd89e54189bc1811d1df709e5a28cfc21b8a1fae17b

Observation 33dcdefe-e59e-430d-8966-2633d29c5969 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

MMaDA: Multimodal Large Diffusion Language Models Training Verifiers to Solve Math Word Problems

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:50:59.907266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:f2d93fd161534a24d7ae99ef6155a34adcc06fbad1122d0fddfa516a185be347

Observation 5311a623-c235-45d2-bc4f-373c347057c0 · outbound

This paper cites Learning transferable visual models from natural language supervision.

MMaDA: Multimodal Large Diffusion Language Models Learning transferable visual models from natural language supervision

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.147138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:92695eae66aeb49f215f050006437957a8794adf41e1612f4ad1fbbdca556c2c

Observation 07cd823d-cbc9-4016-a31d-a06def6de80e · outbound

This paper cites Imagereward: Learning and evaluating human preferences for text-to-image generation.Advances in Neural Information Processing Systems, 36.

MMaDA: Multimodal Large Diffusion Language Models Imagereward: Learning and evaluating human preferences for text-to-image generation.Advances in Neural Information Processing Systems, 36

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.150146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:2439b8d57ff75f1efee07f18d44bd4a2e47614160b8056036fee64d4df6597ac

Observation fa6e826d-f70f-4b22-99ca-3c632c35d5f4 · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.

MMaDA: Multimodal Large Diffusion Language Models Geneval: An object-focused framework for evaluating text-to-image alignment

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.153363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:d5e0fab50f0e5ce3d8c79c4fb8d966b7fe95daafa98021b2cfd5feb8aec8a1c0

Observation 5d67eccc-d9d6-45be-92dc-f8b664f910f4 · outbound

This paper cites WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation.

MMaDA: Multimodal Large Diffusion Language Models WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:24:27.819650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:5eff388f1ab0c388b4e6736b301a18181ca58966d7af79e710398c84db2ae3eb

Observation 6127edc3-50d1-463e-8830-0832603bb267 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

MMaDA: Multimodal Large Diffusion Language Models High-resolution image synthesis with latent diffusion models

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.156716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:6fb7f23a72ff48df7e63bf2f64985cc218ff823d6e39b92dbda3f2205a906e39

Observation b57b3c1b-f38c-43cc-84c9-7533e36819c4 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

MMaDA: Multimodal Large Diffusion Language Models SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:50:59.745826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:3be4b284d5090a429008a7444a451c49b34e0357838d9165e69ef30381498f02

Observation 510bc526-9323-432f-a5c1-cf8a86e89652 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

MMaDA: Multimodal Large Diffusion Language Models Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:50:59.749411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:c57ee0e33249d94c43b275226f6253157d4cd522f84aa5e22f5d000b36e5a45a

Observation e6b68b22-0eb4-4371-9e7a-5af03bd4b398 · outbound

This paper cites VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model.

MMaDA: Multimodal Large Diffusion Language Models VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:50:59.753839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:98885f05b4212868903b7df820e3c0501ecfe04e942653ca59ec06068798f064

Observation 7cf3fa26-7b8a-44aa-b7fb-8a79f285668d · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

MMaDA: Multimodal Large Diffusion Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:50:59.757783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:5bb5a1841eede8072ab3cba32042e17c0c02f809d535cceb542ad2dabe235c37

Observation 9c078740-7d9f-4a6b-8861-875e33ca3925 · outbound

This paper cites Openai o1 system card.preprint.

MMaDA: Multimodal Large Diffusion Language Models Openai o1 system card.preprint

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.159878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:a1d14b23f14f53d5d6ff1e9176dcb8b87a9532fa60ea995aeb8d0c3c8c5685ab

Observation b4fae0d7-6a48-4de8-8ee4-c4d75be11488 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

MMaDA: Multimodal Large Diffusion Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:50:59.771898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:17a31e81c6b31fea2713db549abb0386a0e7fea8186cacda10c40a3ed3b1face

Observation 92383866-01a2-4c19-959c-3a824be15364 · outbound

This paper cites Multimodal foundation models: From specialists to general-purpose assistants.Foundations and Trends® in Computer Graphics and Vision, 16(1-2):1–214.

MMaDA: Multimodal Large Diffusion Language Models Multimodal foundation models: From specialists to general-purpose assistants.Foundations and Trends® in Computer Graphics and Vision, 16(1-2):1–214

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.163124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:ebe9e3421c5e9faef66fddc3c2898ca7100bb0a2ee29940b200e06fb0cdd187c

Observation 253b4759-b89e-45ff-bce7-2c1833eac5e3 · outbound

This paper cites A Survey on Multimodal Large Language Models.

MMaDA: Multimodal Large Diffusion Language Models A Survey on Multimodal Large Language Models

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:42.979926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:042f9b332be02bd2c14df139be6ceed2b6b6626cac51562b5a2e2b91c9ac8f0a

Observation 37c64dfe-60db-43e3-b996-6bb00ec520d7 · outbound

This paper cites Hallucination of Multimodal Large Language Models: A Survey.

MMaDA: Multimodal Large Diffusion Language Models Hallucination of Multimodal Large Language Models: A Survey

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:50:59.786071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:401d281bf8897c9eb4a7b1334cc166927d19cc356c17197f99fa23a8a13df2f9

Observation 71799f63-ff7e-4461-8c58-c67cb357dcb4 · outbound

This paper cites Visual instruction tuning.NeurIPS, 36.

MMaDA: Multimodal Large Diffusion Language Models Visual instruction tuning.NeurIPS, 36

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.166508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:36af96097491e49159cc42c4201b75f938469ca35265a366b5c10959c7e817e6

Observation eb458215-0d14-4806-904f-b4a4fbf4b0c8 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

MMaDA: Multimodal Large Diffusion Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:50:59.796628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:61094d1e49de87c14d9599e336fac55f9bdb3a99858a7f8599689bc49dada4ef

Observation 2c1fac96-d5eb-49f7-879a-bc2c985ea8ee · outbound

This paper cites Qwen Technical Report.

MMaDA: Multimodal Large Diffusion Language Models Qwen Technical Report

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:50:59.808440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:a70ddd02a062e2a296e2ca47731545a610f18b94b0e3ed125ed9f1b5cff118aa

Observation 0c45f6b8-6c9c-4d34-beed-57b63926474f · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

MMaDA: Multimodal Large Diffusion Language Models MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:09:36.761640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:dc9e3ef8049aa217ef0d77634e8c2e0dc62976e367a5ad40ee7bc6db9d3c7105

Observation bbd1b1fa-0b9c-4025-be4d-057ddaff9198 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MMaDA: Multimodal Large Diffusion Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:50:59.818881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:56fbb8819d1be29a6ea141d55a6c401a4c4ba3c340596845d449b9491d66f9f6

Observation d9dfa50b-319f-4700-b3da-9d70218584c3 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

MMaDA: Multimodal Large Diffusion Language Models Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:50:59.824463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:b90c89e47ae83378c03968c420564672ec7b767ede8fca1ab8c665daf3e1259e

Observation 1794a132-4200-4fb3-bc8e-bf60d18828b9 · outbound

This paper cites Kwok, Ping Luo, Huchuan Lu, and Zhenguo Li.

MMaDA: Multimodal Large Diffusion Language Models Kwok, Ping Luo, Huchuan Lu, and Zhenguo Li

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.169646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:ebbdf9096380337b38d2ebe4bcc4a08c23df86876d99ce7de956154c3a92fd59

Observation 248141da-dd18-4816-a150-b61bbafd69f5 · outbound

This paper cites GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models.

MMaDA: Multimodal Large Diffusion Language Models GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:50:59.834549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:4ef728e87d54367095114522c59f0501efcb4d9cc3b6a08d33bd84890f5bfefa

Observation 8e93e53e-8629-4271-8caf-c93d509f8562 · outbound

This paper cites Raphael: Text-to- image generation via large mixture of diffusion paths.NeurIPS, 36.

MMaDA: Multimodal Large Diffusion Language Models Raphael: Text-to- image generation via large mixture of diffusion paths.NeurIPS, 36

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.172627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:472d79f27da54077f6d62a952e9c2b5e772acbe6d38a5eda375abd0416d4ca1f

Observation 949f6199-9132-4b85-a745-fcc2a450ad08 · outbound

This paper cites Mastering text-to-image diffusion: Recaptioning, planning, and generating with multimodal llms.

MMaDA: Multimodal Large Diffusion Language Models Mastering text-to-image diffusion: Recaptioning, planning, and generating with multimodal llms

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.175302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:bfa4c939c9b3bbaade10851fb6c4dfb2eb5375557bd1b86dc4e9e4093432771a

Observation 0d88f640-9c5a-4b81-9635-46bea4750d2e · outbound

This paper cites Itercomp: Iterative composition-aware feedback learning from model gallery for text-to-image generation.

MMaDA: Multimodal Large Diffusion Language Models Itercomp: Iterative composition-aware feedback learning from model gallery for text-to-image generation

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.178079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:aabe9fc625ef4313a8b67db61404d190748cb9a23999762c9008713dac46d5af

Observation f7995592-57e2-4dfd-b7cc-6b9e0d61a067 · outbound

This paper cites An Overview of Diffusion Models: Applications, Guided Generation, Statistical Rates and Optimization.

MMaDA: Multimodal Large Diffusion Language Models An Overview of Diffusion Models: Applications, Guided Generation, Statistical Rates and Optimization

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:50:59.853373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:d5c78f8e2c291a4c5765ec22cf48ad594a377c889a4b729a345cfd0c2c028994

Observation 3df2197d-09f1-4cb0-9dbd-59dcb47bd885 · outbound

This paper cites Diffusion-Sharpening: Fine-tuning Diffusion Models with Denoising Trajectory Sharpening.

MMaDA: Multimodal Large Diffusion Language Models Diffusion-Sharpening: Fine-tuning Diffusion Models with Denoising Trajectory Sharpening

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:50:59.858515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:6b47aa9130196b65f1e752ca487887d7ff1f29eaa40762c54520673907c91cb1

Observation d89783d8-ee86-447f-82de-4b03473bc5e9 · outbound

This paper cites Reward-directed conditional diffusion: Provable distribution estimation and reward improvement.Advancesin Neural Information Processing Systems, 36:60599–60635.

MMaDA: Multimodal Large Diffusion Language Models Reward-directed conditional diffusion: Provable distribution estimation and reward improvement.Advancesin Neural Information Processing Systems, 36:60599–60635

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.180452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:a6a44e9899be711deae425306565a6cfef4c83bb17263532d73772d91289d9bd

Observation 21ad3f4a-8f13-4fb1-9598-168fdd07bd19 · outbound

This paper cites Gradient Guidance for Diffusion Models: An Optimization Perspective.

MMaDA: Multimodal Large Diffusion Language Models Gradient Guidance for Diffusion Models: An Optimization Perspective

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:50:59.868901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:ffb35c3648973008ac7226c918da9ceaffe2ba797bbbdf485bf9806e02b3d6d9

Observation b4237290-8605-422c-91d7-6b17f0d12439 · outbound

This paper cites Score approximation, estimation and distribution recovery of diffusion models on low-dimensional data.

MMaDA: Multimodal Large Diffusion Language Models Score approximation, estimation and distribution recovery of diffusion models on low-dimensional data

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.182881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:1b364c577d7dc5c7da468cba393b4d15d4e54da83883e1f8c0a94a5db717ef15

Observation c5430022-d2a2-46f6-8bb5-2dcde028a193 · outbound

This paper cites Rectified Diffusion: Straightness Is Not Your Need in Rectified Flow.

MMaDA: Multimodal Large Diffusion Language Models Rectified Diffusion: Straightness Is Not Your Need in Rectified Flow

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:50:59.881022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:bc71aa0e0f765a91ea6df3c8453c1350fdc7384ed5ef50ee0917c92bedff34a9

Observation a6f5c7d5-a4d4-4d66-857c-e97fe9dae415 · outbound

This paper cites Improving diffusion-based image synthesis with context prediction.Advances in Neural Information Processing Systems, 36:37636–37656.

MMaDA: Multimodal Large Diffusion Language Models Improving diffusion-based image synthesis with context prediction.Advances in Neural Information Processing Systems, 36:37636–37656

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:51:00.185358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:9b48fa6feca536ab9ade01768ebc645e48720732c2e62a101dc254b2ca43c0f4

Observation 543b501d-f5d6-4375-bfba-703bac231bfa · outbound

This paper cites Structure-guided adversarial training of diffusion models.

MMaDA: Multimodal Large Diffusion Language Models Structure-guided adversarial training of diffusion models

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:50:59.910876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:e625a141646af4624869c1b73584fc682cfa9795055b2fefe30a7d0d4e085b49

Observation 76aaedea-6f6d-4064-a725-783986063885 · outbound

This paper cites Videotetris: Towards compositional text-to-video generation.Advancesin Neural Information Processing Systems, 37:29489–29513.

MMaDA: Multimodal Large Diffusion Language Models Videotetris: Towards compositional text-to-video generation.Advancesin Neural Information Processing Systems, 37:29489–29513

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:50:59.914403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:670e3d2ba87e74bd277e2f5d9092f448489764b8407af84a7c196e7cf5417b3f

Observation 049a8ebd-8223-4cb0-afd1-22fd19b05dfc · outbound

This paper cites Structured denoising diffusion models in discrete state-spaces.NeurIPS, pages 17981–17993.

MMaDA: Multimodal Large Diffusion Language Models Structured denoising diffusion models in discrete state-spaces.NeurIPS, pages 17981–17993

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:50:59.917872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:4fa04f57a20eec429be0abaa4c751be37ea48af16019fa7bba846e44a87556c7

Observation 790d0e44-e598-49d6-b8ea-bdfc16c93625 · outbound

This paper cites Vector quantized diffusion model for text-to-image synthesis.

MMaDA: Multimodal Large Diffusion Language Models Vector quantized diffusion model for text-to-image synthesis

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:50:59.922026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:ca05b6d8c4f721d83d795ec648a09975d042a070bd10a953b2fdc3d28687cd91

Observation 2362884e-f3b6-4043-a4d3-5e4239975d83 · outbound

This paper cites Gritsenko, Jasmijn Bastings, Ben Poole, Rianne van den Berg, and Tim Salimans.

MMaDA: Multimodal Large Diffusion Language Models Gritsenko, Jasmijn Bastings, Ben Poole, Rianne van den Berg, and Tim Salimans

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:50:59.925236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:79852fd22d0f7e651e6c364875ee5a21e08438531cc4b794a89bc9499e5d774f

Observation 59ab54db-1ca8-4f3a-8657-8526c0e9bb87 · outbound

This paper cites Murphy.Probabilistic Machine Learning: AdvancedTopics.

MMaDA: Multimodal Large Diffusion Language Models Murphy.Probabilistic Machine Learning: AdvancedTopics

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:50:59.929018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:72173528e3d6443134bcc1d68be96123637db7994488837d92fd12fcfe24abda

Observation fdd5bb84-3abf-41d9-9342-1657e0830363 · outbound

This paper cites Rethinking the objectives of vector- quantized tokenizers for image synthesis.

MMaDA: Multimodal Large Diffusion Language Models Rethinking the objectives of vector- quantized tokenizers for image synthesis

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:50:59.932577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:3cfbd9aa5a326fd9cf0b8ef619a2b89bb4687dd0275d023bec07c2c89b612066

Observation 19c4bfa5-7d2b-4891-84d7-e719c1307237 · outbound

This paper cites Attention is all you need.NeurIPS, 30.

MMaDA: Multimodal Large Diffusion Language Models Attention is all you need.NeurIPS, 30

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:50:59.935764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:35405f64179361833ab3df524af197cf4e66f6c5add6453196ebc2dc4bb084b3

Observation 09b78d8a-4575-4647-a26d-de6db18337e0 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67.

MMaDA: Multimodal Large Diffusion Language Models Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:50:59.940031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:2d753722f6d368944f5de0ea3d19eb99baf221ac243f6d6b218dab001852d431

Observation fb099ec4-5ddd-4b41-b20b-b2fbd04e324c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

MMaDA: Multimodal Large Diffusion Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 94

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:50:59.742112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:d7db242cf0ddbca3ec4812575e27e2eea18fc8890a5b8aeb89efe96ed9ea8af0

Observation 0f1742aa-d590-4dc5-87b0-f2e39c33fe95 · outbound

This paper cites Image transformer.

MMaDA: Multimodal Large Diffusion Language Models Image transformer

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:50:59.943825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:6ab4b1f3bc0099ac857cdb56f291b4dbf0a88c96e312acf066b5c52d29fa507a

Observation 25311f4d-123b-48d5-9cc9-ba4344ced037 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

MMaDA: Multimodal Large Diffusion Language Models Taming transformers for high-resolution image synthesis

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:50:59.947564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:5a6ed8089f7a15d40c0e4463c47c96ee867dc5899b119d0fdfaeb453e7d8dd3b

Observation ad00e2d3-c149-4788-9428-78d0128f8b05 · outbound

This paper cites Classification accuracy score for conditional generative models.NeurIPS, 32.

MMaDA: Multimodal Large Diffusion Language Models Classification accuracy score for conditional generative models.NeurIPS, 32

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:50:59.951298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:7ff3177841c2446f750f31b7d5d697f5c8f9fa78c07c101b433e910b8b526eb5

Observation 03201462-0c46-4cba-bc4f-0974193d3968 · outbound

This paper cites Generative pretraining from pixels.

MMaDA: Multimodal Large Diffusion Language Models Generative pretraining from pixels

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:50:59.954919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:7a278d3655cf1d26c993e90f1f04ac3a2c2f522008773d7db49e51989c4f46eb

Observation e9eef0f7-ce2b-4a08-8c80-2884929218b5 · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

MMaDA: Multimodal Large Diffusion Language Models VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:51:05.830483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:771bc894e5f9d26c4b1bb59fe7ef52b3a1e59b2871a5a6c5a97fa6b5f2f46291

Observation 85e03685-b3c2-428e-9944-0f96f0b08232 · outbound

This paper cites Visual autoregressive modeling: Scalable image generation via next-scale prediction.Advancesin neural information processing systems, 37:84839–84865.

MMaDA: Multimodal Large Diffusion Language Models Visual autoregressive modeling: Scalable image generation via next-scale prediction.Advancesin neural information processing systems, 37:84839–84865

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:50:59.958488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:1b8765c4fb7c0c1db4253906ee79d885e43629a96115c18449250092d4d1cebf

Observation bfb03666-dc73-483e-a948-ce00f2e52775 · outbound

This paper cites NExT-GPT: Any-to-Any Multimodal LLM.

MMaDA: Multimodal Large Diffusion Language Models NExT-GPT: Any-to-Any Multimodal LLM

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:50:59.768622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:a3db110c6d337457d9cdfabc00a723d118249c498c25d98e30fc7a03b539e3b4

Observation d0aea713-0a11-434f-91c3-b8bf8e3b0ae0 · outbound

This paper cites Any-to-any generation via composable diffusion.

MMaDA: Multimodal Large Diffusion Language Models Any-to-any generation via composable diffusion

Reference 102

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:50:59.961939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:3395d614471290c34a386bba36200f3846de7183b44b45d3eb14851533fc7baf

Observation abff7c5e-1703-4c7e-8ceb-c800c58692f3 · outbound

This paper cites X-VILA: Cross-Modality Alignment for Large Language Model.

MMaDA: Multimodal Large Diffusion Language Models X-VILA: Cross-Modality Alignment for Large Language Model

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:50:59.775698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:c679fb8b2cabf41677099f2e80465b6aa918eb72b2f3d4fd98d7c9ad78326a8b

Observation 427529eb-446c-43df-a223-7c2d03618e88 · outbound

This paper cites Jointly training large autoregressive multimodal models.

MMaDA: Multimodal Large Diffusion Language Models Jointly training large autoregressive multimodal models

Reference 104

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:50:59.965192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:191d639da16ebd674a691a9448c547f2e5e7b3deef9feea9aaa710bb5ff9fcfc

Observation f698e10f-5ba2-467b-8996-e285e6f2fe01 · outbound

This paper cites Vector quantized diffusion model for text-to-image synthesis.

MMaDA: Multimodal Large Diffusion Language Models Vector quantized diffusion model for text-to-image synthesis

Reference 105

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:50:59.968451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:2237041b70d93278c76a6913518d21ab4ceb37b2669a57e84aa9fc5b9eb09ada

Pith citing papers

Observation 1b7c86df-7425-4b28-8daa-40d4aa28a340 · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models MMaDA: Multimodal Large Diffusion Language Models

Reference 135

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:51:00.186449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:7c76d85a8bd611e88b94f66e7793c4a49166e60a5a0b67ce5228550ab48f2fd6

Observation ab559257-454d-43bf-8b9b-c0bcee9e46ca · inbound

GIFT: Guided Importance-Aware Fine-Tuning for Diffusion Language Models cites this paper.

GIFT: Guided Importance-Aware Fine-Tuning for Diffusion Language Models MMaDA: Multimodal Large Diffusion Language Models

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T14:51:30.260333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T14:49:08.081740Z digest=sha256:f0abfd76febdd1fdd2f395d37f017d91e7d97c18bd0509b1919785090e4ecc8b

Observation fef952f2-3797-49a6-a50f-cf81e24ba9ed · inbound

Discrete Guidance Matching: Exact Guidance for Discrete Flow Matching cites this paper.

Discrete Guidance Matching: Exact Guidance for Discrete Flow Matching MMaDA: Multimodal Large Diffusion Language Models

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-05-18T14:16:27.810932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T14:13:48.523955Z digest=sha256:5bc5f21b97531140ea0a609e1f9487243afbafde37fa7d174a61f50c39f2b368

Observation cad80625-c2a5-4ed7-ad51-3222285652f4 · inbound

Motus: A Unified Latent Action World Model cites this paper.

Motus: A Unified Latent Action World Model MMaDA: Multimodal Large Diffusion Language Models

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:51:00.186449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T18:44:36.636455Z digest=sha256:f762bf78430e679644248a7730de2bd24388d851177fad6c6379e1858485b4ad

Observation 1dd19a22-ab81-4af8-b0e9-42b4d1a2685c · inbound

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models cites this paper.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models MMaDA: Multimodal Large Diffusion Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:38.232907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:38.232907Z digest=sha256:4e727a7f8e6df8f3eceb464c391a583ff3beba409e1e524cc4489ffc03fb3dc5

Observation 32197b4c-006c-4b17-ad27-5c150d8fc6dd · inbound

Efficient-DLM: From Autoregressive to Diffusion Language Models, and Beyond in Speed cites this paper.

Efficient-DLM: From Autoregressive to Diffusion Language Models, and Beyond in Speed MMaDA: Multimodal Large Diffusion Language Models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-16T22:31:19.339185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-16T22:29:08.669964Z digest=sha256:8b8dcff4426d5de3f7d4d524c0a0dc990f0d4ab6511abc6ab1f6fd5964d1b70f

Observation 56a7a3e3-ce11-432b-839d-baaca9810c21 · inbound

ViewMask-1-to-3: Multi-View Consistent Image Generation via Multimodal Discrete Diffusion Models cites this paper.

ViewMask-1-to-3: Multi-View Consistent Image Generation via Multimodal Discrete Diffusion Models MMaDA: Multimodal Large Diffusion Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T16:19:36.475336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:19:36.475336Z digest=sha256:bf86e8eb553999e09a32ad373036931bdd6c34c92392511eee30d4e7a32a4e74

Observation df090975-d48e-4dbf-acf9-6cc62755331b · inbound

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models cites this paper.

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models MMaDA: Multimodal Large Diffusion Language Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:38:24.504106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:ba82601e94ccf07667d15fc10ad5aee29a83eb19302e24b7898e46728bcd0c38

Observation 13f313ab-603b-4cef-b83e-0c3c78dcc9b2 · inbound

The Flexibility Trap: Rethinking the Value of Arbitrary Order in Diffusion Language Models cites this paper.

The Flexibility Trap: Rethinking the Value of Arbitrary Order in Diffusion Language Models MMaDA: Multimodal Large Diffusion Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T09:03:21.281392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:03:21.281392Z digest=sha256:dd4ff525edb00806145108f0e00412320507dd2fb86be862eb1c6c03aab22f18

Observation 3c3f005f-105c-4236-a0e2-1d02d5f23d12 · inbound

ChatUMM: Robust Context Tracking for Conversational Interleaved Generation cites this paper.

ChatUMM: Robust Context Tracking for Conversational Interleaved Generation MMaDA: Multimodal Large Diffusion Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T03:57:32.391604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:57:32.391604Z digest=sha256:02ca14d9ec64e0a20e99cbb943f0f3c297cb28287bdf8d5c595c5ee39449d214

Observation 52ce73a6-7591-4de5-a891-63e8b6af6cd8 · inbound

Improving Sampling for Masked Diffusion Models via Information Gain cites this paper.

Improving Sampling for Masked Diffusion Models via Information Gain MMaDA: Multimodal Large Diffusion Language Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-25T07:35:29.597706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-25T07:30:59.831550Z digest=sha256:3501c534630a90136ef4de21a9bcec2b2651aac6722b13c879aa8a0568841ef4

Observation 11c44097-da16-4beb-934e-295162996ef0 · inbound

Improving Full Waveform Inversion in Large Model Era cites this paper.

Improving Full Waveform Inversion in Large Model Era MMaDA: Multimodal Large Diffusion Language Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-02T20:00:15.054046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:00:15.054046Z digest=sha256:1e5bd32f723b0df825b7aae96ac9369e66ff4f94e7e43fc9ce9dd194ad215e99

Observation 5160d785-f484-4f58-a685-23dfa6d5bbcd · inbound

Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion cites this paper.

Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion MMaDA: Multimodal Large Diffusion Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-15T13:43:40.241796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:43:40.241796Z digest=sha256:5b350087d07b79b723cae79c0438243a0d9af34934eed1af8d503130995554e3

Observation 43ecc968-9484-4cf2-85df-ffd568c4061b · inbound

Reinforcement Learning for Diffusion LLMs with Entropy-Guided Step Selection and Stepwise Advantages cites this paper.

Reinforcement Learning for Diffusion LLMs with Entropy-Guided Step Selection and Stepwise Advantages MMaDA: Multimodal Large Diffusion Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:51:00.186449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T12:38:02.339424Z digest=sha256:0999b2fff0be6de6fa13130629822abdf11f1e74f5f1b80c4f1a8c51f97521bd

Observation 4595ef8c-3fe9-434d-84a8-0f7ffbe37779 · inbound

Thinking Diffusion: Penalize and Guide Visual-Grounded Reasoning in Diffusion Multimodal Language Models cites this paper.

Thinking Diffusion: Penalize and Guide Visual-Grounded Reasoning in Diffusion Multimodal Language Models MMaDA: Multimodal Large Diffusion Language Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:51:00.186449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T19:06:45.417233Z digest=sha256:392544459de552aac5a4d429519656624a49a6a3e82d3872f302aa35dc8dc60a

Observation 12d2da17-a5d3-448c-9437-594b4866016d · inbound

DMax: Aggressive Parallel Decoding for dLLMs cites this paper.

DMax: Aggressive Parallel Decoding for dLLMs MMaDA: Multimodal Large Diffusion Language Models

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:51:00.186449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T17:58:17.880199Z digest=sha256:ee1ca3c4ac4b535699e74a6011d8f2c47e5f3bb93096fa00e4f36e10c31ae9fb

Observation ecbb25a0-6515-4899-ae16-743bcb6f6f51 · inbound

DMax: Aggressive Parallel Decoding for dLLMs cites this paper.

DMax: Aggressive Parallel Decoding for dLLMs MMaDA: Multimodal Large Diffusion Language Models

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:47:40.258886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T16:46:56.743268Z digest=sha256:cf68ea02d5ed6b88998b210e8fe51690b726def95ae9f9fac8544475dc364c6e

Observation 70f94ca4-0977-44d6-ae4c-89a99b862690 · inbound

Learning Preference-Based Objectives from Clinical Narratives for Dynamic Sepsis Treatment cites this paper.

Learning Preference-Based Objectives from Clinical Narratives for Dynamic Sepsis Treatment MMaDA: Multimodal Large Diffusion Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-12T22:22:06.385856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:22:06.385856Z digest=sha256:df7345b747294dcb646513af03f2c575b8f5c6198e6b186ba6f95fe297f86989

Observation e0087f75-aa6b-41d7-add5-d7a9910619cd · inbound

TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training cites this paper.

TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training MMaDA: Multimodal Large Diffusion Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:51:00.186449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T15:32:01.551829Z digest=sha256:1938e48cfb134a16748c53af0365a86e33e05914384869de324d0ed2c6049e6a

Observation 8128049b-d782-4112-a1c9-3daee9e4283c · inbound

TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training cites this paper.

TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training MMaDA: Multimodal Large Diffusion Language Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:59:55.288148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T08:59:10.877437Z digest=sha256:ab3525d2ca16e2320746ece9c7c5cf700086b25f700bb1a82e509c42ab320aec

Observation abe51ad2-db00-45d6-8aee-0aaa5fbd0530 · inbound

BARD: Bridging AutoRegressive and Diffusion Vision-Language Models Via Highly Efficient Progressive Block Merging and Stage-Wise Distillation cites this paper.

BARD: Bridging AutoRegressive and Diffusion Vision-Language Models Via Highly Efficient Progressive Block Merging and Stage-Wise Distillation MMaDA: Multimodal Large Diffusion Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:51:00.186449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T14:28:47.505755Z digest=sha256:b8578a2f1a58a73fb2fa8cc20ea64dab1ca816cdb4558971f97f2a4e65ce107f

Observation 241ce839-64e1-4fc9-bf7c-a547a7d67a7d · inbound

Stability-Weighted Decoding for Diffusion Language Models cites this paper.

Stability-Weighted Decoding for Diffusion Language Models MMaDA: Multimodal Large Diffusion Language Models

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T14:51:00.186449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T06:12:37.804816Z digest=sha256:94208b7fe9aea03c36aa995bcff3a35b778a06bd057774163f5a4dad50d8298b

Observation 9f9a43b9-bc4d-49ff-9e33-9b4302fa9024 · inbound

UniGenDet: A Unified Generative-Discriminative Framework for Co-Evolutionary Image Generation and Generated Image Detection cites this paper.

UniGenDet: A Unified Generative-Discriminative Framework for Co-Evolutionary Image Generation and Generated Image Detection MMaDA: Multimodal Large Diffusion Language Models

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:51:00.186449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-09T22:21:10.133655Z digest=sha256:d33cf3974c5360503824d54dc6165d2de27db2193af949afe826f4759225be32

Observation bf864c24-1876-4b1e-a4ce-c555a257bbdd · inbound

dWorldEval: Scalable Robotic Policy Evaluation via Discrete Diffusion World Model cites this paper.

dWorldEval: Scalable Robotic Policy Evaluation via Discrete Diffusion World Model MMaDA: Multimodal Large Diffusion Language Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:51:00.186449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T11:45:18.081248Z digest=sha256:7008c0919aa9f952dbf30392eb0c98a21af8e3b526f3151d924a70d4c45e3edf

Observation de4bc6a6-a86c-4a91-a960-b3f01ce383a9 · inbound

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation cites this paper.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation MMaDA: Multimodal Large Diffusion Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:51:00.186449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:31:26.325118Z digest=sha256:778d8305a614c235d749d3ebf914fc71cc64dccd285bb443af02cd7ab24432a3

Observation 9e3482fc-0701-46c0-8218-8693526b6cbe · inbound

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation cites this paper.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation MMaDA: Multimodal Large Diffusion Language Models

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:43:51.129679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:d1eee0125d4f126c296ac254f655f7b8b4254300f1c5c9f713b2c69fae22d63a

Observation a9b9f25b-f62e-4c85-91d7-af8e98de2a9a · inbound

Why Does Grounding Hurt Medical VQA? Benchmarking, Diagnosis, and Fine-Tuning of Vision-Language Models cites this paper.

Why Does Grounding Hurt Medical VQA? Benchmarking, Diagnosis, and Fine-Tuning of Vision-Language Models MMaDA: Multimodal Large Diffusion Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:51:00.186449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T05:25:44.844424Z digest=sha256:2771752596ec3d95532b3a2d59e14f58f988feaa8c852751d268a0e1aca7f8d8

Observation afa1081a-abf3-4e4c-8977-380af283f1dc · inbound

Why Does Grounding Hurt Medical VQA? Benchmarking, Diagnosis, and Fine-Tuning of Vision-Language Models cites this paper.

Why Does Grounding Hurt Medical VQA? Benchmarking, Diagnosis, and Fine-Tuning of Vision-Language Models MMaDA: Multimodal Large Diffusion Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-02T15:12:46.786429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:12:46.786429Z digest=sha256:04c44408a2c80e398e72bd9be4db36de2199117a94a219b41bff95dc5d0ffe19

Observation 22896690-8163-4db9-82d3-8bc8dd727ee6 · inbound

Break the Block: Dynamic-size Reasoning Blocks for Diffusion Large Language Models via Monotonic Entropy Descent with Reinforcement Learning cites this paper.

Break the Block: Dynamic-size Reasoning Blocks for Diffusion Large Language Models via Monotonic Entropy Descent with Reinforcement Learning MMaDA: Multimodal Large Diffusion Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:51:00.186449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T18:38:07.949718Z digest=sha256:368dc66843c6111e4c651fde0d2f2990b037a3fe50552cab3e5ec7167afd49b1

Observation 5ee87dd0-caaf-4517-9c38-3bbcb043155c · inbound

Break the Block: Dynamic-size Reasoning Blocks for Diffusion Large Language Models via Monotonic Entropy Descent with Reinforcement Learning cites this paper.

Break the Block: Dynamic-size Reasoning Blocks for Diffusion Large Language Models via Monotonic Entropy Descent with Reinforcement Learning MMaDA: Multimodal Large Diffusion Language Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:25:10.332978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T23:59:02.580783Z digest=sha256:05b5a4de6827e6c71e6ef80516d74b1d4a5c6b95fe9e4531fc9dc648a1f83b7d

Observation b100ee1d-9e9c-4680-8b07-6c8ec67f1723 · inbound

NoiseRater: Meta-Learned Noise Valuation for Diffusion Model Training cites this paper.

NoiseRater: Meta-Learned Noise Valuation for Diffusion Model Training MMaDA: Multimodal Large Diffusion Language Models

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:51:00.186449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T01:44:04.922621Z digest=sha256:174efa1ecc0ceb327887ad7056e5c7e60a6ed6f662a40e3627b4e5cad09148de

Observation 538becbd-efb7-4b90-8fc1-932698fa52a8 · inbound

Discrete Langevin-Inspired Posterior Sampling cites this paper.

Discrete Langevin-Inspired Posterior Sampling MMaDA: Multimodal Large Diffusion Language Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:51:00.186449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:21:21.408659Z digest=sha256:6a74d2b564025aee28bd2c6761f49153b8d5504679cb53a3292f977ebe02d989

Observation dc424ff0-e64d-469c-81c9-73bb81583ce7 · inbound

Relative Score Policy Optimization for Diffusion Language Models cites this paper.

Relative Score Policy Optimization for Diffusion Language Models MMaDA: Multimodal Large Diffusion Language Models

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:51:00.186449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-12T03:47:42.196931Z digest=sha256:54bf50d9bfd39c69d84756fcec2bf8d35b948c81124377aa380713fa53c50738

Observation b9d10b12-947a-46c1-8d4c-826d241f0725 · inbound

UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning cites this paper.

UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning MMaDA: Multimodal Large Diffusion Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:51:00.186449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T01:29:23.774196Z digest=sha256:5dcc33c67ca0a86e89e8e5fd7a58fe506b3c0ffe2d3bbdde27da82bf2fefffa0

Observation 869fb5ae-7717-4b82-9fe7-66b12b308dec · inbound

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models cites this paper.

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models MMaDA: Multimodal Large Diffusion Language Models

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:51:00.186449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T07:03:00.503644Z digest=sha256:85716643adb1244e071e6d9e0eea96ad31cb9e58c786710875a24251ef751baf

Observation 4793b85f-c456-4ab3-bd5b-0d9c5a154f2f · inbound

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models cites this paper.

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models MMaDA: Multimodal Large Diffusion Language Models

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:51:00.186449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T21:06:01.667173Z digest=sha256:e719aaed1eba238aab02d7852e33f0b707411c50ce1b044c1b52bb76c79ee24c

Observation cbf8dbdc-c4e7-4fc5-9623-f054887d5dbf · inbound

Language Generation as Optimal Control: Closed-Loop Diffusion in Latent Control Space cites this paper.

Language Generation as Optimal Control: Closed-Loop Diffusion in Latent Control Space MMaDA: Multimodal Large Diffusion Language Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:51:00.186449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-15T01:45:39.649473Z digest=sha256:83e60834f886ed370a7c403fe0b95e0dd44722354f0887adc266b78dbd9b52b7

Observation e28f85a9-b768-42a9-bc43-e4bb70a80e15 · inbound

Language Generation as Optimal Control: Closed-Loop Diffusion in Latent Control Space cites this paper.

Language Generation as Optimal Control: Closed-Loop Diffusion in Latent Control Space MMaDA: Multimodal Large Diffusion Language Models

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:32:39.592192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-19T16:27:59.293776Z digest=sha256:8db9f18265021fb57674f01ecc7cfac0c7e5ae5d7df265a91f1f3a8b0b99e6a3

Observation 1a888b34-b3c8-48f3-a060-0a77a183a5a0 · inbound

Language Generation as Optimal Control: Closed-Loop Diffusion in Latent Control Space cites this paper.

Language Generation as Optimal Control: Closed-Loop Diffusion in Latent Control Space MMaDA: Multimodal Large Diffusion Language Models

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:35:46.492322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-30T21:16:13.804718Z digest=sha256:d44a73da9f9a53a70bc771c3b1ceb4e08dfa1ea63ff2ccd656ef12e69b9cb7c5

Observation a88484b6-616b-40f8-bb85-b1c46bcb2aac · inbound

Roll Out and Roll Back: Diffusion LLMs are Their Own Efficiency Teachers cites this paper.

Roll Out and Roll Back: Diffusion LLMs are Their Own Efficiency Teachers MMaDA: Multimodal Large Diffusion Language Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-19T20:42:46.155603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T20:41:48.063871Z digest=sha256:977cefeaa68b69002b2f3da34e64a4f374f34329d0d1d055166ed9d4c33ab69d

Observation 358a728a-f7ca-42f1-abd3-bc76902c247c · inbound

Fast-dDrive: Efficient Block-Diffusion VLM for Autonomous Driving cites this paper.

Fast-dDrive: Efficient Block-Diffusion VLM for Autonomous Driving MMaDA: Multimodal Large Diffusion Language Models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:06:38.505163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-25T05:00:58.024124Z digest=sha256:b9dcfce603bab90bff67a44236b539f1b562a177f834b9d26e2406f2989c9880

Observation e667c714-65ec-4adf-a8d0-f821f5f87e9b · inbound

Fast-dDrive: Efficient Block-Diffusion VLM for Autonomous Driving cites this paper.

Fast-dDrive: Efficient Block-Diffusion VLM for Autonomous Driving MMaDA: Multimodal Large Diffusion Language Models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-06-30T16:35:12.295098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T16:34:21.422620Z digest=sha256:75c00cc9a4085ea441429fa9988a41461ccd02488cf3a745643d074d8418cf31

Observation 086e7709-73b4-463a-b55b-8fdc273f2be7 · inbound

Visual-Redundancy-Controlled Parallel Decoding for Diffusion-Based Multimodal Large Language Models cites this paper.

Visual-Redundancy-Controlled Parallel Decoding for Diffusion-Based Multimodal Large Language Models MMaDA: Multimodal Large Diffusion Language Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:14:01.647644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T23:09:32.594194Z digest=sha256:575ead6017350e90698d92695a03ff36537128c317ea123b15724fc83b936808

Observation 600fe1e3-ad13-4431-b520-e22a69821dea · inbound

Guidance Contrastive Token Credit Assignment for Discrete Policy Optimization cites this paper.

Guidance Contrastive Token Credit Assignment for Discrete Policy Optimization MMaDA: Multimodal Large Diffusion Language Models

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-06-29T09:03:15.889347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T08:59:03.375697Z digest=sha256:c77b6a548ef9884d760eaa6928f27af15b9095ab0dd23f29d248cbded1379484

Observation a827c1ca-a1b2-4d91-ada0-e3a5a414332c · inbound

GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models cites this paper.

GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models MMaDA: Multimodal Large Diffusion Language Models

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:53:16.303177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T08:44:53.969301Z digest=sha256:ecd0fcb9180b8c933638308481c0a1df13198ea63bf1e170ac1c56c9dfb5b3dd

Observation 6939d29e-488c-4439-aea2-ed4237fc773b · inbound

dMoE: dLLMs with Learnable Block Experts cites this paper.

dMoE: dLLMs with Learnable Block Experts MMaDA: Multimodal Large Diffusion Language Models

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-06-28T22:52:45.090923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-28T22:50:51.900169Z digest=sha256:b6d247bdbf9959d3954ef7f1a2584d3574a7eae91e1694828b44f70dab9f005c

Observation 7f5534d9-0081-4af6-b3be-3e7c875ee5bd · inbound

Lumos-Nexus: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models cites this paper.

Lumos-Nexus: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models MMaDA: Multimodal Large Diffusion Language Models

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-07-01T19:16:00.507408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-28T22:56:21.783415Z digest=sha256:5a7d0a04e2f2e347ec1655ea5e6c882dc6f1fba22b62f966090ddf709d49c617

Observation b0f2c5cc-9191-4ff2-8abc-c9f2a8921161 · inbound

SimSD: Simple Speculative Decoding in Diffusion Language Models cites this paper.

SimSD: Simple Speculative Decoding in Diffusion Language Models MMaDA: Multimodal Large Diffusion Language Models

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:56:19.507957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-28T14:58:00.723105Z digest=sha256:923a15918fd4b37e3b52f8621fe866732b0b1f0e53af7749b4176ca4f181a51b

Observation 88613756-b150-49be-8009-094b86058f00 · inbound

MaskForge: Structure-Aware Adaptive Attacks for Jailbreaking Diffusion Large Language Models cites this paper.

MaskForge: Structure-Aware Adaptive Attacks for Jailbreaking Diffusion Large Language Models MMaDA: Multimodal Large Diffusion Language Models

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:56:24.441096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-28T13:43:51.171443Z digest=sha256:83ca3544705eb48d6b549d7c4d8c861fe4b99e84c667f818c6d2b5ae9515f7a5

Observation aa8649db-da31-4f93-b55e-85eee930f019 · inbound

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation cites this paper.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation MMaDA: Multimodal Large Diffusion Language Models

Reference 60

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T03:06:29.573608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:a4c63a54aade9e6d8a02fa1140cf136c33628911661c6d3cfa8458371227b34b

Observation 778bbb38-14fc-4d4c-b762-eec8e2a7b276 · inbound

Dynamic Infilling Anchors for Format-Constrained Generation in Diffusion Large Language Models cites this paper.

Dynamic Infilling Anchors for Format-Constrained Generation in Diffusion Large Language Models MMaDA: Multimodal Large Diffusion Language Models

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-07-02T08:06:48.424273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-28T06:20:27.041099Z digest=sha256:b4bff43d1afff68d946dc61b4132b59852ed1239419899ea5a942ca6858b5a96

Observation 010827f9-95ea-4d51-816e-476b412c26f5 · inbound

Back on Track: Aligning Rewards and States for Reasoning in Diffusion Large Language Models cites this paper.

Back on Track: Aligning Rewards and States for Reasoning in Diffusion Large Language Models MMaDA: Multimodal Large Diffusion Language Models

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T22:57:26.605270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-27T18:31:21.493677Z digest=sha256:22a44404942ee7a9abcde13eeb0920b70682a54b491e5f9b03467e3fe283a065

Observation 60b3ec82-4b0b-4ddc-baa1-f7972b607dd4 · inbound

TimeROME-DLM: Temporal Causal Tracing and Low-Rank Inference-Time Knowledge Editing for Masked Diffusion Language Models cites this paper.

TimeROME-DLM: Temporal Causal Tracing and Low-Rank Inference-Time Knowledge Editing for Masked Diffusion Language Models MMaDA: Multimodal Large Diffusion Language Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-03T13:38:19.407383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-27T07:42:45.896118Z digest=sha256:d209b56fecc189f6a89d3354e4ca062283ab2b863a2626e754dd5acb6a1c5608

Observation a0b40dd6-1ac5-4f2c-a8c8-aa80f431ff51 · inbound

InterleaveThinker: Reinforcing Agentic Interleaved Generation cites this paper.

InterleaveThinker: Reinforcing Agentic Interleaved Generation MMaDA: Multimodal Large Diffusion Language Models

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-03T15:08:32.844173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-27T06:42:34.126336Z digest=sha256:29601daf878c54c32b4ce5e207a74532a71cf516aafa78c8fd1b3cc9675da6e7

Observation 95866b58-e1d8-4d40-9303-159209c7b898 · inbound

DiPOD: Diffusion Policy Optimization without Drifting Apart cites this paper.

DiPOD: Diffusion Policy Optimization without Drifting Apart MMaDA: Multimodal Large Diffusion Language Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:18:22.943777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-27T07:06:28.426709Z digest=sha256:a1d68abc69f6b926e8f334d1c49f3fc0d0bc16ebf21a0ad224c75f7c5e93d295

Observation 0a64f099-d9aa-4a2c-810e-8af79b3dfc97 · inbound

UniTranslator: A Unified Multi-modal Framework for End-to-end In-Image Machine Translation cites this paper.

UniTranslator: A Unified Multi-modal Framework for End-to-end In-Image Machine Translation MMaDA: Multimodal Large Diffusion Language Models

Reference 51

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T16:09:57.567477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-26T00:50:22.839005Z digest=sha256:4ab91d1e977528a8133dc5863dba74ad01f92ee4bea486d938765e23c4839205

Observation cd5ff93b-0488-4833-8a04-b6b31e6082d8 · inbound

Improved Large Language Diffusion Models cites this paper.

Improved Large Language Diffusion Models MMaDA: Multimodal Large Diffusion Language Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:20:05.679205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-25T21:34:45.620481Z digest=sha256:2cfdc4939c2facd955861d677a5ba4ce39da2218189f7e7aca2f0428ce2b6017

Observation 595926b4-c45a-4e9c-8f11-1313fd0b44cd · inbound

LithoDreamer: A Physics-Informed World Model for Multi-Stage Computational Lithography cites this paper.

LithoDreamer: A Physics-Informed World Model for Multi-Stage Computational Lithography MMaDA: Multimodal Large Diffusion Language Models

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T13:49:52.355089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-26T04:51:18.153338Z digest=sha256:296b6dd4c9fd820898c7b0a1c13440093eed7c3aa566d2c38d9667d9a3792302

Observation c2d8373b-64d8-4ad0-aadf-7063bd4454b7 · inbound

Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis cites this paper.

Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis MMaDA: Multimodal Large Diffusion Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-06-30T08:14:26.662482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T06:07:55.600338Z digest=sha256:eac21e7cd0b002a72f392f4f3f2b860d6b36372192105b1857fe1e93174bcda8

Observation 78dcc2da-fea6-4cd4-89cc-8bbba6b9a481 · inbound

Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis cites this paper.

Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis MMaDA: Multimodal Large Diffusion Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T09:39:32.630179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:39:32.630179Z digest=sha256:7cdd30137fdca983713545482cf4b7269433cdc4dbd6dfea5d7bde9e1894f14d

Observation fbc51968-a12b-4d36-b0b9-dd53948f64b0 · inbound

Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation cites this paper.

Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation MMaDA: Multimodal Large Diffusion Language Models

Reference 55

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T06:04:20.937206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T06:04:07.327934Z digest=sha256:867a4cd8d3f77e29a69e09371f5d255d5ee284171f3458712a37ab6f45c80e07

Observation a4e419d1-7ebb-46d6-bc5c-eb5a0328f9a7 · inbound

Bridging Video Understanding and Generation in a Unified Framework cites this paper.

Bridging Video Understanding and Generation in a Unified Framework MMaDA: Multimodal Large Diffusion Language Models

Reference 77

Resolution
verified exact
local_arxiv, observed 2026-07-01T10:05:40.367208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-01T05:57:54.653504Z digest=sha256:da8c35e9a54ffca1991fddd6a43557f26dfd7a71cae73fed2408381528cf4552

Observation dfd29b72-a0f2-406a-9f9b-dc0bbda9c39c · inbound

TACG: Trajectory-Aware Commit Gating for Diffusion Language Model Decoding cites this paper.

TACG: Trajectory-Aware Commit Gating for Diffusion Language Model Decoding MMaDA: Multimodal Large Diffusion Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-12T03:56:26.770729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T03:56:26.770729Z digest=sha256:72e8bc41b4e8f89900e26842da6871751a4c8a0bdfb99b156bbd3780e54a7990

Observation 9f547940-83d8-426b-85f1-49597944d867 · inbound

Transferability Between Understanding and Generation in Unified Multimodal Models cites this paper.

Transferability Between Understanding and Generation in Unified Multimodal Models MMaDA: Multimodal Large Diffusion Language Models

Reference 102

Resolution
unresolved
no resolver link, observed 2026-07-11T19:17:19.634242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:17:19.634242Z digest=sha256:e0bc3819bcd74df2d780f2328b58a2b0f1d1ea9579e663a9f08e3059ac17dffb

Observation a28d3a0d-2935-4b6f-9904-97d996fbf086 · inbound

Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding cites this paper.

Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding MMaDA: Multimodal Large Diffusion Language Models

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-11T03:07:51.443936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-11T03:04:12.500342Z digest=sha256:9b0dd897ee4215a365c8311af92da393c382b956df2a977e492dc173ce3fb0d5

Observation 04cb6e76-ea97-4900-bd0b-80bd18c98368 · inbound

Discrete Diffusion Models: A Unified Framework from Tokenization to Generation cites this paper.

Discrete Diffusion Models: A Unified Framework from Tokenization to Generation MMaDA: Multimodal Large Diffusion Language Models

Reference 190

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:48.230225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:48.230225Z digest=sha256:fa7a38adcc7efa31f24e88828cd3fbc074f9f8f9e8726cf69956794fa2c89b04

Observation 43a35c33-fa2b-4b4e-9ef5-bb071eb4c73a · inbound

Seeing the End at Step Zero: Accelerating Diffusion MLLMs via MLP Sparsity-Aware Truncation cites this paper.

Seeing the End at Step Zero: Accelerating Diffusion MLLMs via MLP Sparsity-Aware Truncation MMaDA: Multimodal Large Diffusion Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-02T01:48:56.029119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:48:56.029119Z digest=sha256:20ad0abc8c3e54667a349e35325329dbb3d975418bec1df3c54ea75205135f65

Observation 47df349c-489c-436f-aec7-59f53a5cbe36 · inbound

ST-Veto: Spatio-Temporal Token Veto for Diffusion MLLMs via Taylor Prediction and Visual Grounding cites this paper.

ST-Veto: Spatio-Temporal Token Veto for Diffusion MLLMs via Taylor Prediction and Visual Grounding MMaDA: Multimodal Large Diffusion Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T16:48:36.570741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:48:36.570741Z digest=sha256:815477d4e4e3560d847122fbd0c27d3634c0185f5abcdb48e9224649c29ced9a

Observation 231c1717-8a8d-452d-8840-1cedb5534a6d · inbound

Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation cites this paper.

Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation MMaDA: Multimodal Large Diffusion Language Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-01T02:14:09.788866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:14:09.788866Z digest=sha256:7646df4a40a665adeda0cdce523181ffb44a801225a2641507d0c3eaa40ce0a1